BIO331 – Lab 07: Bioinformatics and Working with Strings of Biological Data
Bioinformatics
Author
Dee Ruttenberg
Published
March 18, 2026
Today we’re going to apply our work on R to Bioinformatic data sets. The typical output of a DNA sequencing experiment is a fastq file, consisting of often millions of configs which need to be stitched together (do you remember the two ways we do that?). Fastq files are different from fasta files as they contain quality scores – this is useful to see how reliable your fastq data is. The first tool we are going to use is Fastqc, to interpret the results of our fastq files.
Quality Control
if (!require("BiocManager", quietly =TRUE))install.packages("BiocManager")
Bioconductor version '3.19' is out-of-date; the current release version '3.23'
is available with R version '4.6'; see https://bioconductor.org/install
BiocManager::install("Rqc")
Bioconductor version 3.19 (BiocManager 1.30.27), R 4.4.3 (2025-02-28)
Warning: package(s) not installed when version(s) same as or greater than current; use
`force = TRUE` to re-install: 'Rqc'
We will use a specific small dataset used specifically for teaching bioinformatics. Rqc will output an .html file (essentially a web page). For this class, I will ask that you understand specifically the results of ‘Average Quality’, ‘Cycle-Specific Average Quality’, ‘Cycle-specific Quality Distribution’, and ‘Cycle-Specific Base Call Proportion’.
library(Rqc)
Loading required package: BiocParallel
Loading required package: ShortRead
Loading required package: BiocGenerics
Attaching package: 'BiocGenerics'
The following objects are masked from 'package:stats':
IQR, mad, sd, var, xtabs
Welcome to Bioconductor
Vignettes contain introductory material; view with
'browseVignettes()'. To cite Bioconductor, see
'citation("Biobase")', and for packages 'citation("pkgname")'.
Attaching package: 'Biobase'
The following object is masked from 'package:MatrixGenerics':
rowMedians
The following objects are masked from 'package:matrixStats':
anyMissing, rowMedians
The final skill we need to learn is how to align .fasta files to a reference genome. The data we have today is aligned so we’ll have to arbitrarily “unalign” it. We will use our first sequence as the reference.
Look at the ‘Cycle-specific Base Call Proportion’ rqc file. What is the ratio of A/C/G/T in cycles 1-10. What is the ratio of A/C/G/T in cycles 65-75? Based on your knowledge of DNA-sequencing, propose a hypothesis for why the cycles have a different ratio.
Look at the ‘Cycle specific Quality Distribution’. What is the typical quality of cycles 1-10? What is the typical quality of cycles 65-75? Based on your knowledge of DNA-sequencing, propose a hypothesis for why the cycles have a different quality
Homework Question 2: Loading a File, Reverse Complements
For one of the research projects, we will be studying a specific protein coding site of hcmv associated with the transition of hcmv from latent to active. Align the following two queries to any genome in the reference file. Which one successfully aligns? What does that mean about the site of interest?