Conversion of gbk Genome Data for CoreGenes

Step 1: Extract Proteins from gbk File

The first step is to extract the proteins from the gbk file. For this I use Genome2D resulting in the following type of results for phage 4HA13 (locus tag AC4HA13):

>AC4HA13_000
MKPNYVAIRKSKEAMFHRFIEAKRKAELEGKVVIKKKKKNKKKNYNIFFFLLSR
>AC4HA13_010
MTYYDADLGLVMCESELSLEIDALDWEHELPKGEPQWGDDDYVYVAPTDEFDIPF
      

Step 2: Copy to Notepad

Step 3: Check for Spurious Characters

Step 4: Use Replace Feature

Using the Replace feature of Notepad to replace >AC4HA13 with >gp|AC4HA13|AC4HA13_% giving:

>gp|AC4HA13|AC4HA13_%000
MKPNYVAIRKSKEAMFHRFIEAKRKAELEGKVVIKKKKKNKKKNYNIFFFLLSR
>gp|AC4HA13|AC4HA13_%010
MTYYDADLGLVMCESELSLEIDALDWEHELPKGEPQWGDDDYVYVAPTDEFDIPF

Step 5: Add Organism Information

The tedious part - paste |[Escherichia phage 4HA13] at the end of each fasta row:

>gp|AC4HA13|AC4HA13_%000|[Escherichia phage 4HA13]
MKPNYVAIRKSKEAMFHRFIEAKRKAELEGKVVIKKKKKNKKKNYNIFFFLLSR
>gp|AC4HA13|AC4HA13_%010|[Escherichia phage 4HA13]
MTYYDADLGLVMCESELSLEIDALDWEHELPKGEPQWGDDDYVYVAPTDEFDIPF

Step 6: Paste into CoreGenes

Paste this formatted data into Custom Data in CoreGenes 3.5

Updated: January, 2026