ENGINEERING GUIDE · SECURITY

Ransomware Detection & Monitoring Guide

Build a practical ransomware detection capability from identity, endpoint, network, firewall, server and backup telemetry.

Engineering referenceRansomware · Detection

Related Ingenix tool

Assess ransomware resilience in the browser.

Ransomware Readiness Assessment →
Engineering principle: Good ransomware detection is not a single antivirus alert. Build a chain of evidence across identity, endpoint, network, server, firewall and backup systems. The strongest detections combine several weak signals into a high-confidence picture and give the response team enough context to act.

What a ransomware detection capability should achieve

The purpose of monitoring is not to collect the maximum possible number of logs. It is to identify the behaviours that indicate an attacker is progressing through the environment and to surface them early enough to contain the incident.

A useful detection capability should answer four questions:

  • What happened? Identify the account, device, process, connection or configuration change.
  • Where did it happen? Identify the host, user, server, network zone and affected service.
  • Is it expected? Compare the event with known administration, software deployment and backup activity.
  • What should we do? The alert should point to a practical containment action rather than simply reporting that something unusual occurred.

1. Identity monitoring

Identity is one of the most valuable telemetry sources because ransomware operators frequently need valid credentials to move through an environment. Monitoring should cover both successful and failed authentication, privilege changes and changes to the identity infrastructure itself.

What to monitor

  • Repeated authentication failures followed by a successful login.
  • Authentication from unusual locations, devices or networks.
  • Privileged logons from ordinary user workstations.
  • New administrative accounts or unexpected membership of privileged groups.
  • Changes to authentication policies, MFA configuration or conditional-access controls.
  • Unusual use of service accounts or accounts authenticating outside their normal function.
  • Large numbers of authentication attempts against multiple systems.

What suspicious activity can look like

A single failed login is rarely meaningful. A pattern such as one account attempting authentication against many systems, followed by successful access to a server it does not normally use, is much more significant.

Controls and response

  • Centralise identity logs so events can be correlated across systems.
  • Retain enough authentication history to establish a timeline.
  • Alert on privileged-account anomalies with lower thresholds than ordinary accounts.
  • Ensure responders can rapidly disable or restrict a compromised account.
  • Protect identity infrastructure and monitor changes to privileged groups and authentication controls.

Engineering test: Create a controlled test account and generate representative failed logons, unusual access and privilege changes. Confirm that the events reach the monitoring platform and generate useful alerts.

2. Endpoint monitoring

Endpoints are often where ransomware execution becomes visible. EDR or equivalent endpoint telemetry should provide process, file, user and security-control information rather than only reporting malware signatures.

What to monitor

  • Unexpected process creation and unusual parent-child process relationships.
  • Execution from temporary or user-writable locations.
  • Large-scale file modifications or rapid changes to many files.
  • Unexpected use of scripting and command interpreters.
  • Attempts to disable or modify security services.
  • New persistence mechanisms, scheduled tasks or services.
  • Unexpected administrative tools appearing on user workstations.

Why mass file activity matters

Encryption commonly produces a very different file-access pattern from normal office activity: one process may rapidly open and rewrite large numbers of files across many directories or network shares. This does not prove ransomware—backup, indexing, migration and legitimate data-processing applications can also generate high-volume file activity—but it is a valuable behavioural signal.

Controls and response

  • Deploy centrally managed EDR or equivalent protection to supported endpoints and servers.
  • Enable tamper protection where available.
  • Ensure endpoint telemetry is retained centrally rather than only on the device.
  • Give responders a rapid isolation mechanism.
  • Maintain an approved list of high-volume applications so legitimate activity can be distinguished from anomalous behaviour.

Engineering test: Validate that a test endpoint can be isolated quickly and that the monitoring platform retains the process and user context needed to investigate what happened immediately before isolation.

3. File server and SMB monitoring

File servers can provide an early warning when ransomware begins affecting shared data. Monitoring should focus on unusual access patterns rather than simply watching whether the server is online.

What to monitor

  • Sudden increases in file opens, writes, renames or deletes.
  • A workstation accessing an unusually large number of shares or files.
  • New clients accessing sensitive shares.
  • Large numbers of failed access attempts.
  • Unexpected administrative-share activity.
  • Rapid changes to file extensions or file metadata where telemetry supports this.

Controls

  • Restrict SMB between network zones to required paths.
  • Remove unnecessary share permissions and inherited access.
  • Separate administrative access from normal user access.
  • Enable appropriate file-access auditing on high-value data.
  • Monitor high-value shares more closely than ordinary user data.

Engineering test: Establish a normal baseline for file-server activity. Without a baseline, a detection such as “high file activity” will generate too many false positives to be useful.

4. Network and firewall monitoring

Network telemetry is particularly valuable when an attacker moves between systems. The objective is not to inspect every packet; it is to identify connections that violate the expected communication model.

What to monitor

  • Unexpected workstation-to-workstation connections.
  • New or unusual SMB, RDP, WinRM and other administration flows.
  • One source contacting many hosts in a short period.
  • Unexpected server-to-workstation communication.
  • Traffic crossing security zones that is not part of the documented application design.
  • Large outbound transfers from systems that do not normally send data externally.
  • New connections to external infrastructure following suspicious endpoint or identity activity.

Controls

  • Log inter-zone firewall decisions, particularly denied and unusual permitted traffic.
  • Retain useful source, destination, port and identity context where available.
  • Use network segmentation to make suspicious traffic easier to identify.
  • Build alerts around deviations from documented application flows rather than generic “any unusual traffic” rules.

Engineering test: From a representative user VLAN, verify that attempts to reach management and server services outside the user's required access generate the expected firewall events.

5. Privileged administration monitoring

Administrative activity deserves separate monitoring because attackers often use legitimate tools and credentials rather than obviously malicious malware.

What to monitor

  • New privileged accounts and changes to privileged group membership.
  • Interactive administrator logons to servers and workstations.
  • Remote administration from unexpected source systems.
  • Changes to Group Policy, security policies and endpoint-management configuration.
  • Changes to firewall rules, backup configuration and security tooling.
  • Bulk administrative actions across multiple hosts.

Administrative activity should be correlated with the person, device and change process responsible for it. A change made by an approved administrator from an approved management workstation during a planned maintenance window is very different from the same action originating from a normal user workstation at an unusual time.

6. Backup monitoring

Backup systems are a critical detection source because attackers may attempt to destroy recovery before deploying ransomware.

What to monitor

  • Deletion or unexpected modification of backup jobs.
  • Changes to retention periods or repository configuration.
  • Unexpected deletion of snapshots or recovery points.
  • New administrators or changes to backup permissions.
  • Sudden failures across multiple backup jobs.
  • Unusual access to backup repositories.
  • Changes that reduce the number or protection level of available recovery copies.

Controls

  • Forward important backup events to central monitoring.
  • Separate backup administration from ordinary domain administration.
  • Alert on destructive or security-relevant backup changes.
  • Periodically verify that protected recovery copies still exist and can be restored.

Engineering test: Perform a controlled change to a test backup policy and verify that the monitoring system records who made the change, what changed, when it happened and whether an alert was generated.

7. Security-tool and logging tampering

Attackers who understand the environment may attempt to reduce visibility before carrying out destructive actions. Loss of telemetry can therefore be a security event in its own right.

What to monitor

  • EDR or antivirus agents stopping or becoming unhealthy.
  • Security services being disabled or modified.
  • Logging agents stopping.
  • Firewall logging being disabled or materially changed.
  • Unexpected changes to SIEM, monitoring or alerting configuration.
  • Large gaps in telemetry from normally active systems.

A monitoring platform should not treat “no events” as automatically meaning “nothing happened”. Monitor the health of the monitoring system itself.

8. Correlation: turning individual events into an attack story

The most useful ransomware detections often come from combinations of events. A single unusual RDP connection may be legitimate. A suspicious authentication followed by unusual RDP, privilege use and mass file activity is much more concerning.

Useful correlation patterns

  • Credential abuse: authentication anomalies + unusual source device + privileged access.
  • Lateral movement: unusual authentication + new SMB/RDP connections + access to multiple hosts.
  • Preparation: privileged activity + security-tool changes + backup configuration changes.
  • Encryption: suspicious process + rapid file modifications + access to multiple shares.
  • Recovery attack: backup administrator activity + deletion/retention changes + unusual source host.

Correlation should increase confidence without hiding the underlying evidence. Analysts should still be able to see the individual events that caused an alert.

9. Prioritising alerts

Not every event deserves the same response. Prioritise alerts according to the potential impact and the number of independent signals supporting them.

  • Critical: confirmed ransomware behaviour, widespread encryption indicators, destructive backup activity or compromise of highly privileged identity infrastructure.
  • High: suspicious privileged activity combined with lateral movement, security-tool tampering or unusual access to sensitive systems.
  • Medium: isolated unusual authentication, scanning or administrative behaviour requiring investigation.
  • Informational: expected changes, routine administration and events retained primarily for investigation.

Thresholds should reflect the organisation's normal operating patterns. A school, small business and large enterprise will not have identical baselines.

10. Logging architecture and retention

Useful detection depends on logs being available when the incident occurs. Define the minimum telemetry required before choosing retention periods or SIEM rules.

High-value sources

  • Identity provider and directory authentication logs.
  • Endpoint/EDR process and security events.
  • Domain controller and privileged-account events.
  • Firewall and VPN logs.
  • DNS and relevant network telemetry.
  • File-server and SMB audit events for critical data.
  • Backup and storage administration logs.
  • Virtualisation and management-platform audit logs.

Protect central logs from alteration and ensure timestamps are consistent. During an incident, a reliable timeline is often as valuable as any individual alert.

11. Detection response workflow

A detection is useful only if someone can act on it. Define what happens after an alert before deploying large numbers of detection rules.

  1. Validate: determine whether the activity is expected, suspicious or clearly malicious.
  2. Scope: identify affected users, endpoints, servers, network zones and accounts.
  3. Contain: isolate affected endpoints, disable compromised accounts and restrict relevant network paths as appropriate.
  4. Preserve: retain relevant logs and evidence before rebuilding systems where practical.
  5. Escalate: involve the appropriate technical, management, legal, insurance or incident-response resources.
  6. Recover: restore systems using trusted recovery processes.
  7. Improve: identify which detection or control should have acted earlier and test the improvement.

12. Detection validation

Do not assume that a configured rule works. Detection engineering requires testing.

  • Test representative authentication anomalies.
  • Test a controlled endpoint-isolation workflow.
  • Test unusual administrative access.
  • Test security-tool and logging tamper alerts.
  • Test firewall events for prohibited east-west traffic.
  • Test backup deletion or policy-change alerts in a controlled environment.
  • Run tabletop exercises using realistic alert sequences.
  • Measure detection latency: how long between the behaviour occurring and the team receiving a useful alert?
  • Measure response latency: how long between the alert and containment?

Implementation checklist

  • Define the ransomware behaviours the organisation needs to detect.
  • Inventory identity, endpoint, firewall, server and backup telemetry.
  • Centralise high-value security logs.
  • Protect the logging and monitoring infrastructure itself.
  • Establish normal authentication and network baselines.
  • Alert on privileged-account anomalies.
  • Alert on unusual lateral movement and management traffic.
  • Monitor security-tool and logging health.
  • Monitor destructive backup and recovery changes.
  • Correlate multiple signals into higher-confidence detections.
  • Give responders practical containment actions for each high-priority alert.
  • Test every important detection and record the result.
  • Measure detection and containment time.
  • Review false positives and tune detections regularly.