MS Core and Cloud Specialist

Contractor

Job Description

The MS Core and Cloud Specialist is responsible for technical leadership, operational support, troubleshooting, lifecycle management, and continuous improvement of core network, cloud, and Deep Packet Inspection (DPI) platforms deployed across mobile and fixed broadband networks.

The role acts as a primary technical owner for the DPI and associated core/cloud technology domain, ensuring platform stability, availability, performance, security, and alignment with business, customer, and regulatory requirements.

The specialist will work closely with Packet Core, IP Transport, OSS/BSS, Network Operations, Automation, Analytics, Planning, and technology vendors to ensure efficient service delivery and rapid resolution of complex technical issues.

Key Responsibilities

1. Core, Cloud & DPI Platform Operations

  • Manage the operation, maintenance, troubleshooting, and lifecycle management of DPI and associated core/cloud platforms deployed across mobile and FTTH networks.
  • Ensure platform stability, availability, performance, and capacity.
  • Perform routine health checks, preventive maintenance, platform monitoring, and operational assessments.
  • Support software upgrades, migrations, configuration changes, expansions, and technology refresh activities.
  • Ensure operational activities comply with approved procedures, change-management processes, and service requirements.
  • Monitor platform resources, traffic, sessions, capacity, and service performance.

2. Technical Ownership

  • Act as the primary technical owner for the DPI and relevant core/cloud technology domain.
  • Provide technical leadership for complex operational and performance-related issues.
  • Work closely with Packet Core, IP Transport, OSS/BSS, Operations, Automation, Analytics, and vendor teams.
  • Ensure technical solutions meet agreed specifications and operational requirements.
  • Identify technical risks and proactively implement corrective and preventive measures.

3. Troubleshooting & Root Cause Analysis

  • Perform deep-dive Root Cause Analysis (RCA) using packet captures, flow logs, session traces, system logs, and platform diagnostics.
  • Investigate complex network, service, and platform issues.
  • Correlate platform-level logs with packet-level evidence to identify and isolate faults.
  • Use tools such as Wireshark and tcpdump for packet capture and analysis.
  • Develop systematic diagnostic approaches for complex network faults.
  • Identify the underlying cause of recurring problems and recommend permanent corrective actions.
  • Document troubleshooting results, root causes, resolutions, and lessons learned.

4. Customer Issue Management

  • Handle and resolve customer-impacting technical issues related to core, cloud, DPI, traffic management, and service performance.
  • Analyze customer-impacting incidents using available monitoring, probing, packet-analysis, and performance-management tools.
  • Coordinate with relevant technical teams to restore services within agreed timelines.
  • Provide technical updates and recommendations to stakeholders.
  • Identify recurring customer issues and implement measures to prevent future occurrences.

5. Emergency Recovery & Service Restoration

  • Support service restoration and recovery during major incidents and platform failures.
  • Execute approved emergency procedures, including:
    • Emergency rule bypass
    • Traffic re-routing
    • Graceful platform failover
    • Service restoration
    • Capacity redistribution
  • Assess potential service impact before implementing emergency actions.
  • Ensure recovery activities are performed safely and efficiently.
  • Participate in post-incident reviews and implement preventive measures.

6. Vendor & Technical Assistance Coordination

  • Liaise with technology vendors and Technical Assistance Centers (TAC) for complex platform defects and technical issues.
  • Provide detailed technical evidence, logs, traces, packet captures, and troubleshooting results to vendor support teams.
  • Manage technical escalations and track issues through to resolution.
  • Validate vendor recommendations, patches, fixes, and configuration changes.
  • Coordinate with vendors during upgrades, migrations, expansions, and platform optimization activities.

7. Automation & Analytics

  • Develop and maintain solutions based on network-operations automation use cases.
  • Identify opportunities to automate configuration management, health checks, log analysis, monitoring, alerting, and reporting.
  • Provide domain expertise to automation and analytics teams to support development of relevant use cases.
  • Improve automated service-delivery methodologies and operational processes.
  • Identify repetitive manual activities and develop appropriate automation solutions.
  • Support the implementation of proactive fault detection and predictive analytics.

8. Proactive Monitoring & Trend Analysis

  • Perform continuous monitoring of platform health, traffic, capacity, and service performance.
  • Conduct trend analysis to proactively detect potential failures and performance degradation.
  • Establish performance baselines and monitor defined thresholds.
  • Identify abnormal traffic patterns, capacity constraints, resource issues, and emerging service risks.
  • Initiate restoration and repair activities based on analysis, automated alerts, or domain-support processes.
  • Recommend proactive corrective and preventive actions.

9. Change & Implementation Management

  • Support implementation of approved technical changes, upgrades, migrations, and configuration activities.
  • Ensure solutions are implemented according to approved specifications and change requests.
  • Validate technical implementation and service performance following changes.
  • Ensure changes do not negatively impact service availability or customer experience.
  • Perform post-change validation and identify any unexpected service or platform behavior.
  • Coordinate with relevant technical teams during implementation and rollback activities.

Technical Responsibilities

The role will cover areas including:

  • Deep Packet Inspection (DPI)
  • Mobile Packet Core
  • Fixed Broadband / FTTH Networks
  • Cloud and Virtualized Network Functions
  • IP Networking
  • Traffic Management
  • Application Classification
  • Packet Capture and Analysis
  • Network Performance
  • Service Assurance
  • Platform Health Monitoring
  • Capacity Management
  • Incident Management
  • Root Cause Analysis
  • Automation
  • Analytics
  • Network Security and Traffic Policies

Education Requirements

  • Bachelor’s degree in Engineering, preferably in:
    • Electronics Engineering
    • Computer Engineering
    • Telecommunications Engineering
    • Electrical Engineering
    • Computer Science
    • Information Technology
    • Or an equivalent technical discipline.

Required Experience & Technical Skills

  • Experience in telecom network operations, core networks, cloud platforms, service assurance, DPI, or a related technical environment.
  • Hands-on experience with DPI systems or traffic-management platforms in a live telecommunications environment.
  • Good understanding of network architecture across 2G, 3G, 4G, and 5G technologies.
  • Strong understanding of Packet Core network architecture and protocols.
  • Strong understanding of DPI principles, including:
    • Flow inspection
    • Application classification
    • Protocol identification
    • Traffic policy enforcement
  • Understanding of application detection methodologies, including:
    • Shallow Packet Inspection
    • Deep Packet Inspection
    • Signature-based detection
  • Strong understanding of the TCP/IP protocol stack, including transport and application-layer protocols.
  • Proficiency with Wireshark and tcpdump for packet capture and analysis.
  • Ability to correlate packet-level evidence with platform logs and system diagnostics.
  • Good understanding of DPI-related KPIs and performance indicators.
  • Strong troubleshooting and Root Cause Analysis capabilities.
  • Good knowledge of IP networking and network troubleshooting.
  • Experience with network monitoring, service assurance, and network-probing solutions.

Networking & Infrastructure Skills

  • Strong understanding of IP networking concepts and protocols.
  • Experience with enterprise and telecom networking environments.
  • IP Networking certification or equivalent practical experience is preferred.
  • Experience with network infrastructure platforms such as Cisco, Arista, or equivalent technologies is preferred.
  • Hands-on experience with Linux server administration.
  • Experience with storage server solutions is an advantage.
  • Understanding of virtualization and cloud infrastructure.
  • Knowledge of cloud-native network functions and containerized environments is an advantage.

Automation & Scripting

  • Experience automating network operations tasks, including:
    • Configuration management
    • Platform health checks
    • Log analysis
    • Alerting
    • Monitoring
    • Reporting
  • Knowledge of Python, Bash, SQL, or similar scripting/programming languages is preferred.
  • Ability to identify automation opportunities and translate operational requirements into practical technical solutions.

Preferred Experience

  • Experience supporting large-scale mobile or fixed broadband networks.
  • Experience with DPI and traffic-management platforms.
  • Experience with Packet Core and IP networks.
  • Experience with cloud and virtualized telecom environments.
  • Experience troubleshooting complex multi-vendor telecommunications environments.
  • Experience with customer experience management tools or network-probing solutions.
  • Experience with automation and analytics platforms.
  • Experience supporting major incidents and emergency service-restoration activities.
  • Experience working with technology vendors and TAC support teams.
  • Understanding of Machine Learning (ML), Artificial Intelligence (AI), and Cloud technologies.

Core Competencies

  • Strong analytical and problem-solving skills.
  • Excellent troubleshooting and Root Cause Analysis capabilities.
  • Ability to decompose complex network faults into systematic diagnostic steps.
  • Strong understanding of telecom network architecture and protocols.
  • Ability to work effectively under pressure during major incidents.
  • Proactive approach to identifying and mitigating technical risks.
  • Strong customer and service-oriented mindset.
  • Strong communication and stakeholder-management skills.
  • Ability to work effectively with cross-functional technical teams.
  • Strong vendor-management and escalation skills.
  • Ability to work independently and as part of a 24×7 operational support team.
  • Strong documentation and technical reporting skills.

24×7 Support

  • Participate in a 24×7 on-call support rotation within the team.
  • Provide technical support during critical incidents and service-impacting events.
  • Respond to escalations within defined operational and service-level requirements.

Job Overview

All content copyrighted Tangent International © All rights reserved. Recruitment Website Design - RecWebs