Conversion of gbk Genome Data for CoreGenes
Step 1: Extract Proteins from gbk File
The first step is to extract the proteins from the gbk file. For this I use Genome2D resulting in the following type of results for phage 4HA13 (locus tag AC4HA13):
>AC4HA13_000
MKPNYVAIRKSKEAMFHRFIEAKRKAELEGKVVIKKKKKNKKKNYNIFFFLLSR
>AC4HA13_010
MTYYDADLGLVMCESELSLEIDALDWEHELPKGEPQWGDDDYVYVAPTDEFDIPF
Step 2: Copy to Notepad
Step 3: Check for Spurious Characters
Step 4: Use Replace Feature
Using the Replace feature of Notepad to replace >AC4HA13 with
>gp|AC4HA13|AC4HA13_% giving:
>gp|AC4HA13|AC4HA13_%000 MKPNYVAIRKSKEAMFHRFIEAKRKAELEGKVVIKKKKKNKKKNYNIFFFLLSR >gp|AC4HA13|AC4HA13_%010 MTYYDADLGLVMCESELSLEIDALDWEHELPKGEPQWGDDDYVYVAPTDEFDIPF
Step 5: Add Organism Information
The tedious part - paste |[Escherichia phage 4HA13] at the end of
each fasta row:
>gp|AC4HA13|AC4HA13_%000|[Escherichia phage 4HA13] MKPNYVAIRKSKEAMFHRFIEAKRKAELEGKVVIKKKKKNKKKNYNIFFFLLSR >gp|AC4HA13|AC4HA13_%010|[Escherichia phage 4HA13] MTYYDADLGLVMCESELSLEIDALDWEHELPKGEPQWGDDDYVYVAPTDEFDIPF
Step 6: Paste into CoreGenes
Paste this formatted data into Custom Data in CoreGenes 3.5
Updated: January, 2026