Job Data Converter

Our Job Data Feeds are delivered as gzip-compressed JSON Lines files on AWS S3 and AWS Data Exchange. This page shows how to turn them into CSV, XML, RSS 2.0, Atom, Parquet, Excel or NDJSON - with complete, tested source code in JavaScript, Python, Java and Bash that produces the same structure as the formats of our Job Postings API.

Table of Contents

1. From JSON Lines to Your Format

Every delivery file is named techmap_jobs_{countryCode}_{YYYY-MM-DD}.jsonl.gz - for example techmap_jobs_us_2026-09-06.jsonl.gz - and contains one job posting as a JSON object per line. JSON Lines is ideal for streaming and for data warehouses, but many tools expect a table, a feed or a columnar file. The converters below:

  1. stream the gzip file line by line, so even files with hundreds of thousands of postings never have to fit into memory,
  2. map each posting to the job structure of the Job Postings API (same field names as in the API result, see Section 2),
  3. write one or more output formats in a single pass - with the same columns, element names and types as the API's format= parameter.

Because the output matches the API, you can mix both sources: e.g. backfill a job board from the daily files and keep it current with the API, using the same importer.

FormatOutput fileSame as APITypical use
CSV.csvformat=csvExcel, Google Sheets, pandas, SQL imports
XML.xmlformat=xmlEnterprise integrations, ATS and HR systems
RSS 2.0.rss.xmlformat=rssJob boards, WordPress job plugins, feed readers
Atom.atom.xmlformat=atomFeed consumers that prefer Atom 1.0
Parquet.parquetformat=parquetSnowflake, BigQuery, Redshift, Athena, Spark, DuckDB
Excel.xlsx(CSV columns)Business users, small countries
NDJSON.ndjsonformat=json (result array)Code that already reads the API

2. Field Mapping

The files contain the full export schema documented in the data dictionary - raw text and html, the original page JSON, company.*, location.*, salary.* and more. The converters derive the compact job structure of the API from it. Arrows (→) are fallbacks used when the first value is missing or empty.

2.1 File Fields to API Fields

API field / CSV columnField in the .jsonl.gz file
titlename
companycompany.name → company.nameOrg → json.jsonLD.hiringOrganization.name
jsonLDjson.jsonLD (schema.org JobPosting incl. description) → json.schemaOrg
occupationjson.jsonLD.relevantOccupation → position.name → "N/A"
industryjson.jsonLD.industry → json.inferredTags.INDUSTRIES[0] → "N/A"
departmentjson.jsonLD.employmentUnit → json.inferredTags.DEPARTMENTS[0] → "N/A"
city / state / postCodejson.jsonLD.jobLocation.address.addressLocality / addressRegion / postalCode → location.orgAddress.city / state / postCode
geoPoint{lat, lon} from json.jsonLD.jobLocation.latitude/longitude → location.orgAddress.geoPoint (lat/lng)
countryCodesourceCC
language / localefirst two letters of locale / locale
timezone / timezoneOffsetjson.jsonLD.applicantLocationRequirements ("CEST Timezone" → "CEST") / location.orgAddress.timezoneOffset
workType / workPlace / careerLevel / contractTypejson.inferredTags.WORK_TYPES / WORK_PLACES / CAREER_LEVELS / CONTRACT_TYPES - unique, sorted, ["N/A"] if empty
skillsjson.jsonLD.skills → json.inferredTags.SKILLS
hasSalarytrue if salary.minValue, salary.maxValue or json.jsonLD.baseSalary is set
dateCreated / dateExpired / dateActivedateCreated / json.jsonLD.validThrough / dateExpired, else dateCreated + 1 month
source / portal / isDuplicate / isRecruiter / isDirectsame field names
description (CSV and Excel only)json.jsonLD.description

Need more columns, such as salary.minValue, company.info.companySize or the plain text? Add them in the mapping function (toApiJob / to_api_job / to_api) - every writer picks up new fields automatically except the fixed CSV and Parquet column lists, which you extend the same way.

2.2 RSS and Atom Elements

RSS items and Atom entries use the same elements as the API feeds. Values of N/A become empty elements, arrays are joined with commas.

RSS 2.0 <item>Atom <entry>Source (API job field)
titletitletitle
descriptioncontentjsonLD.description
pubDate (RFC 822)updated, published (ISO 8601)dateCreated
linklinkjsonLD.url
guididjsonLD.url
categorycategory term="..."occupation
locationlocationjsonLD.jobLocation.name
city, statecity, statecity, state
countrycountryjsonLD.jobLocation.address.addressCountry → countryCode
companycompanycompany
company_url, company_logocompany_url, company_logojsonLD.hiringOrganization.url, .logo
apply_linkapply_linkjsonLD.sameAs
salarysalaryjsonLD.baseSalary.name
workType, contractTypeworkType, contractTypeworkType joined with ", " → jsonLD.employmentType
industry, department, occupationindustry, department, occupationindustry, department, occupation
careerLevel, workPlacecareerLevel, workPlacecareerLevel, workPlace joined with ", "
skillsskillsjsonLD.skills joined with ", "

3. Output Formats

Pick a format to see what the converters produce and how to create only this format in each language.

CSV (.csv - API: format=csv)

One row per job posting and one column per top-level field of the API result. Nested objects and arrays (jsonLD, geoPoint, skills, workType, ...) are written as JSON strings, and jsonLD.description is copied into an extra description column at the end - exactly like format=csv of the API (json2csv). Strings are quoted, numbers and booleans are not. Use it for Excel, Google Sheets, pandas or database imports.

Create CSV only
node convert-jobs.js techmap_jobs_us_2026-09-06.jsonl.gz csv
python convert_jobs.py techmap_jobs_us_2026-09-06.jsonl.gz csv
java -cp "lib/*" ConvertJobs.java techmap_jobs_us_2026-09-06.jsonl.gz csv
./convert-jobs.sh techmap_jobs_us_2026-09-06.jsonl.gz csv
Example output (shortened)
"occupation","dateActive","city","timezone","contractType","language","industry","jsonLD","source","locale","geoPoint","title","skills","dateCreated","timezoneOffset","countryCode","company","state","isDuplicate","portal","department","workPlace","isRecruiter","hasSalary","careerLevel","workType","postCode","isDirect","dateExpired","description"
"Engineer","2026-10-06T00:00:00.000Z","Austin","CDT","[""N/A""]","en","IT","{""@context"":""https://schema.org"",""@type"":""JobPosting"",""title"":""Data Engineer"", ...}","linkedin_us","en_US","{""lat"":30.2672,""lon"":-97.7431}","Data Engineer","[""Python"",""SQL"",""AWS""]","2026-09-06T00:00:00+0000",,"us","Example Corp","Texas",false,"linkedin","IT","[""Hybrid""]",false,false,"[""N/A""]","[""FullTime""]","78701",false,"2026-10-06T00:00:00.000Z","We are looking for a Data Engineer ..."

4. Converter Source Code

Each converter is a single file you can copy, run and adapt. All four produce identical CSV, XML, RSS, Atom and Parquet output for the same input file (except for number formatting such as 0.0 vs. 0 inside JSON strings). Pass the formats you need as comma-separated second argument - the default is all of them.

Dependencies: npm: xml-js (XML, RSS, Atom) and parquetjs-lite (Parquet). Gzip, line reading and CSV use the Node.js standard library.

Install and run
# Node.js 18+ - xml-js and parquetjs-lite are the same libraries the Job Postings API uses
npm install xml-js@^1.6.11 parquetjs-lite@^0.8.7

node convert-jobs.js techmap_jobs_us_2026-09-06.jsonl.gz              # all formats
node convert-jobs.js techmap_jobs_us_2026-09-06.jsonl.gz csv,rss      # only CSV and RSS
convert-jobs.js
#!/usr/bin/env node
// convert-jobs.js - convert Techmap job postings (.jsonl.gz) to CSV, XML, RSS 2.0, Atom, Parquet and NDJSON
// Usage:  node convert-jobs.js techmap_jobs_us_2026-09-06.jsonl.gz [csv,xml,rss,atom,parquet,ndjson]
// Install: npm install xml-js parquetjs-lite   (the same libraries the Techmap Job Postings API uses)
const fs = require('fs');
const zlib = require('zlib');
const readline = require('readline');
const { once } = require('events');
const { json2xml } = require('xml-js');
const { ParquetSchema, ParquetWriter } = require('parquetjs-lite');

const input = process.argv[2];
if (!input) {
  console.error('Usage: node convert-jobs.js <techmap_jobs_cc_YYYY-MM-DD.jsonl.gz> [csv,xml,rss,atom,parquet,ndjson]');
  process.exit(1);
}
const formats = (process.argv[3] || 'csv,xml,rss,atom,parquet,ndjson').toLowerCase().split(',');
const base = input.replace(/\.jsonl?(\.gz)?$/, '');

// Fields of a job in the API result (api.techmap.io /api/v2/jobs/search)
const FIELDS = ['occupation', 'dateActive', 'city', 'timezone', 'contractType', 'language', 'industry', 'jsonLD',
  'source', 'locale', 'geoPoint', 'title', 'skills', 'dateCreated', 'timezoneOffset', 'countryCode', 'company',
  'state', 'isDuplicate', 'portal', 'department', 'workPlace', 'isRecruiter', 'hasSalary', 'careerLevel',
  'workType', 'postCode', 'isDirect', 'dateExpired'];
const BOOLEAN_FIELDS = ['isDuplicate', 'isRecruiter', 'isDirect', 'hasSalary'];
const NUMBER_FIELDS = ['timezoneOffset'];

// ---------- helpers ----------
const asList = (v) => (Array.isArray(v) ? v : v ? [v] : []);
const str = (v) => (typeof v === 'string' ? v : '');
const sanitize = (v) => (v === undefined || v === null || v === 'N/A' || v === '' ? '' : v);
const toDate = (s) => new Date(String(s).replace(/([+-]\d\d)(\d\d)$/, '$1:$2')); // "+0000" -> "+00:00"
const cleanXml = (s) => s.replace(/[\u0000-\u0008\u000B\u000C\u000E-\u001F]/g, ''); // chars not allowed in XML
const escAttr = (s) => String(s).replace(/&amp;/g, '&').replace(/&/g, '&amp;').replace(/</g, '&lt;'); // xml-js only escapes quotes in attributes
const toXml = (obj) => cleanXml(json2xml(JSON.stringify(obj), { compact: true, spaces: 2 }))
  .replace(/<@/g, '<AT').replace(/<\/@/g, '</AT'); // "@type" -> <ATtype> like the API
function tagList(raw, key) {
  const values = ((raw.json && raw.json.inferredTags && raw.json.inferredTags[key]) || [])
    .filter((v) => v && v !== 'N/A');
  return values.length ? [...new Set(values)].sort() : ['N/A'];
}

// Map one line of the S3/ADX file to the job structure returned by the API
function toApiJob(raw) {
  const json = raw.json || {};
  const jsonLD = json.jsonLD || json.schemaOrg || {};
  const tags = json.inferredTags || {};
  const place = jsonLD.jobLocation || {};
  const address = place.address || {};
  const org = (raw.location && raw.location.orgAddress) || {};
  const salary = raw.salary || {};
  let geoPoint = null;
  if (typeof place.latitude === 'number' && typeof place.longitude === 'number') geoPoint = { lat: place.latitude, lon: place.longitude };
  else if (org.geoPoint && org.geoPoint.lat != null) geoPoint = { lat: org.geoPoint.lat, lon: org.geoPoint.lng };
  const dateExpired = jsonLD.validThrough || null;
  let dateActive = dateExpired;
  if (!dateActive && raw.dateCreated) {
    const d = toDate(raw.dateCreated);
    d.setUTCMonth(d.getUTCMonth() + 1);
    dateActive = d.toISOString();
  }
  return {
    occupation: jsonLD.relevantOccupation || (raw.position && raw.position.name) || 'N/A',
    dateActive,
    city: address.addressLocality || org.city || '',
    timezone: (str(jsonLD.applicantLocationRequirements).match(/^(\S+) Timezone$/) || [])[1] || str(org.timezone),
    contractType: tagList(raw, 'CONTRACT_TYPES'),
    language: (raw.locale || '').slice(0, 2),
    industry: jsonLD.industry || (tags.INDUSTRIES || [])[0] || 'N/A',
    jsonLD,
    source: raw.source || '',
    locale: raw.locale || '',
    geoPoint,
    title: raw.name || jsonLD.title || '',
    skills: asList(jsonLD.skills).length ? asList(jsonLD.skills) : asList(tags.SKILLS),
    dateCreated: raw.dateCreated,
    timezoneOffset: typeof org.timezoneOffset === 'number' ? org.timezoneOffset : null,
    countryCode: raw.sourceCC || '',
    company: (raw.company && (raw.company.name || raw.company.nameOrg)) || (jsonLD.hiringOrganization || {}).name || '',
    state: address.addressRegion || org.state || '',
    isDuplicate: !!raw.isDuplicate,
    portal: raw.portal || '',
    department: jsonLD.employmentUnit || (tags.DEPARTMENTS || [])[0] || 'N/A',
    workPlace: tagList(raw, 'WORK_PLACES'),
    isRecruiter: !!raw.isRecruiter,
    hasSalary: !!(salary.minValue || salary.maxValue || jsonLD.baseSalary),
    careerLevel: tagList(raw, 'CAREER_LEVELS'),
    workType: tagList(raw, 'WORK_TYPES'),
    postCode: address.postalCode || org.postCode || '',
    isDirect: !!raw.isDirect,
    dateExpired,
  };
}

// ---------- format writers (same mapping as format=csv|xml|rss|atom|parquet in the API) ----------
// CSV: one column per top-level field, nested objects/arrays as JSON, jsonLD.description as extra "description" column
const csvCell = (v) => {
  if (v === undefined || v === null) return '';
  if (typeof v === 'number' || typeof v === 'boolean') return String(v);
  const s = typeof v === 'object' ? JSON.stringify(v) : String(v);
  return '"' + s.replace(/"/g, '""') + '"';
};
const csv = {
  ext: 'csv',
  header: () => [...FIELDS, 'description'].map(csvCell).join(',') + '\n',
  row: (job) => [...FIELDS.map((f) => job[f]), job.jsonLD.description || ''].map(csvCell).join(',') + '\n',
  footer: () => '',
};

const xml = {
  ext: 'xml',
  header: () => '<?xml version="1.0" encoding="UTF-8"?>\n<response>\n<api>Techmap.io Job Posting API</api>\n<result>\n',
  row: (job) => toXml({ job }) + '\n',
  footer: () => '</result>\n</response>\n',
};

// RSS item and Atom entry use the same element names as the API
function feedFields(job) {
  const ld = job.jsonLD;
  return {
    location: (ld.jobLocation || {}).name,
    city: job.city,
    state: job.state,
    country: ((ld.jobLocation || {}).address || {}).addressCountry || job.countryCode,
    company: job.company,
    company_url: (ld.hiringOrganization || {}).url,
    company_logo: sanitize((ld.hiringOrganization || {}).logo),
    apply_link: ld.sameAs,
    salary: (ld.baseSalary || {}).name,
    workType: sanitize(job.workType.join(', ') || ld.employmentType),
    contractType: sanitize(job.workType.join(', ') || ld.employmentType),
    industry: sanitize(job.industry),
    careerLevel: sanitize(job.careerLevel.join(', ')),
    workPlace: sanitize(job.workPlace.join(', ')),
    skills: sanitize(asList(ld.skills).join(', ')),
    department: sanitize(job.department || ld.employmentUnit),
    occupation: sanitize(job.occupation),
  };
}
const rss = {
  ext: 'rss.xml',
  header: () => '<?xml version="1.0" encoding="UTF-8"?>\n<rss version="2.0">\n<channel>\n' + toXml({
    title: 'Techmap.io Job Postings',
    link: 'https://api.techmap.io',
    description: 'Techmap job postings from ' + input + ' in RSS 2.0 Feed format.',
    pubDate: new Date().toUTCString(),
    docs: 'https://api.techmap.io',
    ttl: '60',
  }) + '\n',
  row: (job) => toXml({ item: {
    title: job.title,
    description: job.jsonLD.description || '',
    pubDate: toDate(job.dateCreated).toUTCString(),
    link: job.jsonLD.url,
    guid: job.jsonLD.url,
    category: sanitize(job.occupation),
    ...feedFields(job),
  } }) + '\n',
  footer: () => '</channel>\n</rss>\n',
};

const atom = {
  ext: 'atom.xml',
  header: () => '<?xml version="1.0" encoding="UTF-8"?>\n<feed version="1.0" xmlns="http://www.w3.org/2005/Atom">\n' + toXml({
    id: 'https://api.techmap.io',
    title: 'Techmap.io Job Postings',
    updated: new Date().toISOString(),
    subtitle: 'Techmap job postings from ' + input + ' in Atom 1.0 Feed format.',
    category: { _attributes: { term: 'jobs' } },
    docs: 'https://api.techmap.io',
  }) + '\n',
  row: (job) => toXml({ entry: {
    id: job.jsonLD.url,
    title: job.title,
    updated: toDate(job.dateCreated).toISOString(),
    link: job.jsonLD.url,
    content: job.jsonLD.description || '',
    published: toDate(job.dateCreated).toISOString(),
    category: { _attributes: { term: escAttr(sanitize(job.occupation)) } },
    ...feedFields(job),
  } }) + '\n',
  footer: () => '</feed>\n',
};

const ndjson = { ext: 'ndjson', header: () => '', row: (job) => JSON.stringify(job) + '\n', footer: () => '' };

// Parquet: booleans as BOOLEAN, numbers as DOUBLE, objects/arrays as JSON strings (UTF8) - like the API
function parquetWriter() {
  const fields = {};
  for (const f of FIELDS) {
    const type = BOOLEAN_FIELDS.includes(f) ? 'BOOLEAN' : NUMBER_FIELDS.includes(f) ? 'DOUBLE' : 'UTF8';
    fields[f] = { type, optional: true };
  }
  let writer;
  return {
    open: async () => { writer = await ParquetWriter.openFile(new ParquetSchema(fields), base + '.parquet'); },
    append: async (job) => {
      const row = {};
      for (const f of FIELDS) {
        const v = job[f];
        if (v === undefined || v === null) row[f] = null;
        else if (typeof v === 'object') row[f] = JSON.stringify(v);
        else row[f] = v;
      }
      await writer.appendRow(row);
    },
    close: () => writer.close(),
  };
}

// ---------- main: stream the gzip file line by line ----------
async function write(stream, text) {
  if (text && !stream.write(text)) await once(stream, 'drain'); // respect backpressure
}

async function main() {
  const textFormats = { csv, xml, rss, atom, ndjson };
  const outputs = formats.filter((f) => textFormats[f]).map((f) => ({
    fmt: textFormats[f],
    stream: fs.createWriteStream(base + '.' + textFormats[f].ext, 'utf8'),
  }));
  const parquet = formats.includes('parquet') ? parquetWriter() : null;
  if (parquet) await parquet.open();
  for (const o of outputs) await write(o.stream, o.fmt.header());

  const lines = readline.createInterface({
    input: fs.createReadStream(input).pipe(zlib.createGunzip()),
    crlfDelay: Infinity,
  });
  let count = 0;
  for await (const line of lines) {
    if (!line.trim()) continue;
    const job = toApiJob(JSON.parse(line));
    for (const o of outputs) await write(o.stream, o.fmt.row(job));
    if (parquet) await parquet.append(job);
    count++;
  }

  for (const o of outputs) {
    await write(o.stream, o.fmt.footer());
    o.stream.end();
    await once(o.stream, 'finish');
  }
  if (parquet) await parquet.close();
  console.log('Converted ' + count + ' job postings from ' + input + ' to: ' + formats.join(', '));
}

main().catch((err) => { console.error(err); process.exit(1); });

Converting many files

Download a month of files with the AWS CLI and loop over them - each file is converted independently:

aws s3 cp --recursive s3://YOUR_BUCKET_ALIAS/ . --exclude "*" --include "techmap_jobs_us_2026-09-*.jsonl.gz"

for f in techmap_jobs_us_2026-09-*.jsonl.gz; do
  python convert_jobs.py "$f" csv,parquet
done

5. Shortcuts: DuckDB and the API

Raw files to Parquet or CSV with one DuckDB command

If you want all original fields instead of the API structure, DuckDB reads the gzip files directly and keeps nested objects as STRUCT columns (Parquet) or JSON (CSV):

duckdb -c "COPY (SELECT * FROM read_json_auto('techmap_jobs_us_2026-09-*.jsonl.gz', union_by_name = true))
           TO 'techmap_jobs_us_2026-09.parquet' (FORMAT parquet)"

duckdb -c "COPY (SELECT sourceCC, dateCreated, name, url, company.name AS company, location.orgAddress.city AS city, text
           FROM read_json_auto('techmap_jobs_us_2026-09-06.jsonl.gz')) TO 'jobs.csv' (HEADER)"

Get CSV, XML, RSS, Atom or Parquet directly from the API

The Job Postings API returns the same formats without any conversion - add format=csv, xml, rss, atom or parquet to a search request (default is json). Sample responses in every format are available on the API sample data page.

curl -G "https://daily-international-job-postings.p.rapidapi.com/api/v2/jobs/search" \
  --data-urlencode "countryCode=us" \
  --data-urlencode "dateCreated=2026-09-06" \
  --data-urlencode "format=csv" \
  -H "X-RapidAPI-Key: YOUR_API_KEY" \
  -o techmap_jobs_us_2026-09-06_page1.csv

Use the API for filtered, near real-time queries and the data feeds for complete daily volumes per country - the converters on this page make both look the same.

6. References

Need another format or a custom field mapping? Contact us - we are happy to help.