$fastqc filename.fastq
Started analysis of filename.fastq
Exception in thread "Thread-1" java.lang.OutOfMemoryError: Java heap space
at uk.ac.babraham.FastQC.Modules.GCModel.GCModel.<init>(GCModel.java:69)
at uk.ac.babraham.FastQC.Modules.PerSequenceGCContent.processSequence(PerSequenceGCContent.java:198)
at uk.ac.babraham.FastQC.Analysis.AnalysisRunner.run(AnalysisRunner.java:88)
at java.base/java.lang.Thread.run(Thread.java:844)
$ grep -n 'Xm' /home/prakki/sw/FastQC/fastqc
167: unshift @java_args,"-Xmx${memory}m";
170: unshift @java_args,'-Xmx250m';
$ vi /home/prakki/sw/FastQC/fastqc
Changed the 250m to 10g (depends on the RAM of your computer) in the line 170
$ fastqc filename.fastq
Started analysis of filename.fastq
Approx 5% complete for filename.fastq
Approx 10% complete for filename.fastq
Approx 15% complete for filename.fastq
Approx 20% complete for filename.fastq
Approx 25% complete for filename.fastq
Approx 30% complete for filename.fastq
Approx 35% complete for filename.fastq
Approx 40% complete for filename.fastq
Approx 45% complete for filename.fastq
Approx 50% complete for filename.fastq
Approx 55% complete for filename.fastq
Approx 60% complete for filename.fastq
Approx 65% complete for filename.fastq
Approx 70% complete for filename.fastq
Approx 75% complete for filename.fastq
Approx 80% complete for filename.fastq
Approx 85% complete for filename.fastq
Approx 90% complete for filename.fastq
Approx 95% complete for filename.fastq
Analysis complete for filename.fastq
Showing posts with label Solved. Show all posts
Showing posts with label Solved. Show all posts
Monday, April 1, 2019
Tuesday, February 13, 2018
Merged entries in the non-redundant Nucleotide (nt) database
I ran the following command for to extract a sequence from nt database:
I noticed that this is a merged entry in the 'non-redundant' Nucleotide (nt) database, where two nucleotide sequences which have the same sequence are combined under a single FASTA entry.
When I used the second identifier in the merged entry, blastdbcmd command gave me the same result in return:
$ blastdbcmd -db /mnt/Storage/nt/nt/nt -entry KJ413946.1 | grep '>' >KJ413946.1 Escherichia coli strain ECS01 plasmid pNDM-ECS01, complete sequence >KP900017.1 Raoultella ornithinolytica strain YNKP001 plasmid pYNKP001-NDM, complete sequence
I noticed that this is a merged entry in the 'non-redundant' Nucleotide (nt) database, where two nucleotide sequences which have the same sequence are combined under a single FASTA entry.
When I used the second identifier in the merged entry, blastdbcmd command gave me the same result in return:
$ blastdbcmd -db /mnt/Storage/nt/nt/nt -entry KP900017.1 | grep '>' >KJ413946.1 Escherichia coli strain ECS01 plasmid pNDM-ECS01, complete sequence >KP900017.1 Raoultella ornithinolytica strain YNKP001 plasmid pYNKP001-NDM, complete sequence
Curious to know, I checked if they are the same length. Yes they are of same length (41190).
$ blastdbcmd -db /mnt/Storage/nt/nt/nt -entry KP900017.1 | grep -v '>' | tr -d '\n' | wc 0 1 41190 $ blastdbcmd -db /mnt/Storage/nt/nt/nt -entry KJ413946.1 | grep -v '>' | tr -d '\n' | wc 0 1 41190
Again curious to know, if these two sequences are sharing an exact similarity match, I blasted both of them. And with no surprise, they share 100 % identity. See blast screenshots below:
So, how to sort this out? I need only one annotation instead of multiple for my analysis and do not want the second identifier to confuse my downstream analysis scripts.
By searching around, problem could solved by using "-target_only" option while running the command. I believe this situation might arise in nr database too but switching this on should help to achieve the purpose.
$ blastdbcmd -db /mnt/Storage/nt/nt/nt -entry KJ413946.1 -target_only | grep '>' >KJ413946.1 Escherichia coli strain ECS01 plasmid pNDM-ECS01, complete sequence
Wednesday, February 7, 2018
[bcf_sync] incorrect number of fields (0 != 5) at 0:0
The following samtools v0.1.18 mpileup command gave me error:
$ samtools mpileup -E -M0 -Q25 -q30 -m2 -D -S test_SortSam_withReadGroup.bam | bcftools view -vcg - > test.vcf
[mpileup] 1 samples in 1 input files
<mpileup> Set max per-file depth to 8000
[bcf_sync] incorrect number of fields (0 != 5) at 0:0
[afs] 0:0.000
including option "-g" helped to overcome the problem.
$ samtools mpileup -g -E -M0 -Q25 -q30 -m2 -D -S test_S8_SortSam_withReadGroup.bam | bcftools view -vcg - > test.vcf
"-g " computes genotype likelihoods and output them in the binary call format (BCF).
$ samtools mpileup -E -M0 -Q25 -q30 -m2 -D -S test_SortSam_withReadGroup.bam | bcftools view -vcg - > test.vcf
[mpileup] 1 samples in 1 input files
<mpileup> Set max per-file depth to 8000
[bcf_sync] incorrect number of fields (0 != 5) at 0:0
[afs] 0:0.000
including option "-g" helped to overcome the problem.
$ samtools mpileup -g -E -M0 -Q25 -q30 -m2 -D -S test_S8_SortSam_withReadGroup.bam | bcftools view -vcg - > test.vcf
"-g " computes genotype likelihoods and output them in the binary call format (BCF).
Subscribe to:
Posts (Atom)

