Big Data Engineer - Python and Spark

Citi — Pune, Maharashtra

₹3–7 LPA (HireSetu estimate) · 0–2 yrs exp · Freshers eligible · Full-time

HireSetu listing reference 966dd8 — confirm requirements below, then apply on the employer site.

Quick answer: The posting says 0 to 2 years and calls the position trainee, then asks for at least three years designing compute-heavy systems. Read it as a mid-level big data job in Citi's Decision Management group, where a Master's and hands-on Spark, Hive and Hadoop are taken as given.

From Citi's job posting

The Data/Information Mgt Analyst is a trainee professional role. Requires a good knowledge of the range of processes, procedures and systems to be used in carrying out assigned tasks and a basic understanding of the underlying concepts and principles upon which the job is based. Good understanding of how the team interacts with others in accomplishing the objectives of the area.

Makes evaluative judgements based on the analysis of factual information. They are expected to resolve problems by identifying and selecting solutions through the application of acquired technical experience and will be guided by precedents. Must be able to exchange information in a concise way as well as be sensitive to audience diversity. Limited but direct impact on the business through the quality of the tasks/services provided.

Impact of the job holder is restricted to own job.

Qualifications

  • Master’s / Engineering Degree with 0- 2 years of experience in Big Data systems, Hive, Hadoop, Spark (Python/ scala) and cloud-based data management technologies
  • Hands-on experience in Unix Scripting, Python and Scala programing along with strong experience in SQL.
  • Comfortable working with completed unstructured, undocumented code and turning it around into best-in-class code redesigning costly compute and data processes and aligning to best development standards
  • Experienced in working with large and multiple datasets, data warehouses and ability to pull data using relevant programs and coding.
  • Well versed with necessary data preprocessing and application engineering skills
  • At least 3 years of experience designing software systems with intense computational needs across real time and batch process .
  • Experience and understanding of Supervised, unsupervised machine learning techniques
  • Exposure to data ingestion, ETL tools such as Talend, modeling tools, Performance Management tooling such as Pepper data, Cloudera stack will be a plus
  • Knowledge of data management, data governance, data security and regulatory practices
  • Ability to identify, clearly articulate and solve complex business problems and present them to the management in a structured and simpler form
  • Should have experience of working in onsite, offsite delivery model
  • Experience working with large and multiple datasets, data warehouses and ability to pull data using relevant programs and coding.
  • Previous related experience preferred
  • High attention to detail

Education

  • Bachelors/University degree or equivalent experience

This job description provides a high-level review of the types of work performed. Other job-related duties may be assigned as required.

------------------------------------------------------

Most Relevant Skills

Please see the requirements listed above.

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

View Citi’s EEO Policy Statement and the Know Your Rights poster.

The employer’s company overview, benefits and equal-opportunity statement are left out here; they are on the original posting.

Text from Citi's official job posting, formatted by HireSetu. The employer's own page is the final word on details.

HireSetu's note

The experience line contradicts itself

Two numbers sit in this listing. The tag says 0 to 2 years, and the opening paragraph calls it a trainee role. Then the requirements ask for at least three years of designing software systems with intense computational needs, with previous related experience preferred. Treat the second as the real bar. A Master's degree plus a couple of years of actual pipeline work is the profile that fits, and someone with only coursework will struggle to show the undocumented-code experience. With one or two years behind you, apply anyway, but know which side of that gap you're standing on.

Rewriting code someone else abandoned

The unusual line in this posting is about completed, unstructured, undocumented code you turn into something maintainable, redesigning costly compute and data processes as you go. That's the work. You inherit Hive and Spark jobs nobody has touched in years, work out what they actually produce, and rebuild them cheaper and cleaner. Bring a story about a pipeline or script you took over and fixed, and be able to say what it did, what it cost to run and what you changed.

Cost is measured on this team. Understand partitioning, shuffle and why a job that finishes successfully can still be a waste of a cluster.

Fundamentals before vendor names

Talend, Pepperdata, Cloudera and data governance appear as a plus, so don't spend the week memorising tool names. Spark internals matter more: RDDs against DataFrames, broadcast joins against shuffles, how a job splits into stages. SQL gets used heavily over big tables, so window functions, large joins and reading a query plan deserve an evening. Python and Scala are both listed, so lead with whichever you have genuinely written in production. Unix scripting comes up daily for scheduling and troubleshooting. Supervised and unsupervised techniques are on the list too, and you should be able to name one you applied and what came of it.

Drafted with AI from Citi's job posting and checked by Prashanth Jakkula before publishing. See our Editorial Policy.

How to apply

  1. Open Citi's careers site with the Apply button.
  2. Follow the application steps on the site and upload your resume.
  3. Keep the confirmation email or application number for follow-up.

Apply on Citi's site

No fees, ever. HireSetu never charges candidates, and no genuine employer asks for money to apply, interview or join. If anyone asks you to pay, stop and report the job.

Similar openings on HireSetu

View Big Data Engineer - Python and Spark on HireSetu · More Citi jobs · Data Science jobs · Jobs in Pune · Career Insights · FAQ · Companies

HireSetu
Career Intelligence Platform