Related Ingenix tool
Inspect captured traffic when the VPN is established but the application still fails.
In this guide
VPN troubleshooting terms
Understanding the terminology makes firewall logs and packet captures much easier to interpret.
| Term | Meaning | Why it matters |
|---|---|---|
| IKE | Internet Key Exchange. Negotiates authenticated security associations between VPN peers. | If IKE fails, the encrypted tunnel cannot be established. |
| IKE SA | The security association protecting IKE control traffic. | A healthy IKE SA does not necessarily mean user traffic can pass. |
| Child SA | The IPsec security association carrying protected user traffic. | Missing or incorrect Child SAs commonly indicate Phase 2/selector/proposal problems. |
| Traffic selector | The source/destination traffic definition that determines what traffic belongs to a policy-based IPsec SA. | Selectors must be compatible between peers. |
| ESP | Encapsulating Security Payload, the IPsec protocol normally used to protect payload traffic. | Seeing IKE but no ESP can indicate that the tunnel is not carrying user traffic. |
| NAT-T | NAT Traversal, which encapsulates IPsec ESP in UDP when NAT is detected. | Usually changes the observed data-plane traffic to UDP/4500. |
| Route-based VPN | Traffic is directed through a logical tunnel interface and normal routing decides what enters it. | Routing and tunnel-interface state become key evidence. |
| Policy-based VPN | Traffic selectors/policy determine which traffic is encrypted. | Policy order and selectors are critical. |
| PFS | Perfect Forward Secrecy. Uses additional key exchange for Child SAs. | A PFS mismatch can prevent Phase 2 negotiation. |
| PMTUD | Path MTU Discovery, used to determine the largest packet that can traverse a path without fragmentation. | Broken PMTUD can cause large transfers to fail while pings work. |
A repeatable troubleshooting workflow
- Define the exact test. Record source IP, destination IP, protocol and port, time and expected result.
- Check tunnel state. Determine whether IKE and Child SAs are established.
- Check the peer. Confirm public reachability and that the correct peer address is being used.
- Check negotiation logs. Look for the first meaningful failure rather than the final timeout.
- Check routing. Confirm the source traffic has a route towards the VPN.
- Check firewall policy. Confirm the intended security rule is matching.
- Check NAT. Confirm the packet is not being translated unexpectedly.
- Capture traffic. Compare what leaves the client/network with what reaches the VPN gateway.
- Test incrementally. Start with ICMP or a known TCP port, then test the actual application.
- Change one thing at a time. Record each change so a successful test can be reproduced.
1. Peer reachability
Before investigating cryptography, establish that the VPN peers can communicate. Confirm the configured public address, upstream routing, ISP connectivity and any intervening firewall or NAT device.
For IPsec, check whether IKE packets reach the peer. IKE commonly uses UDP/500. When NAT traversal is in use, UDP/4500 is normally used for the encapsulated exchange and data plane. Do not assume a successful ping to the peer proves that VPN traffic is permitted.
Useful evidence: gateway interface status, routing table, upstream firewall logs, WAN packet capture and the peer's IKE logs.
2. IKE negotiation
IKE negotiation establishes the control-plane security association. Compare both peers for:
- IKE version.
- Authentication method and credentials/certificates.
- Encryption algorithm.
- Integrity or PRF.
- Diffie-Hellman group.
- Peer identity/ID settings.
- Lifetime and rekey settings.
Do not treat a generic message such as negotiation failed as the diagnosis. Find the more specific reason immediately before it: proposal mismatch, authentication failure, invalid identity, certificate problem, unreachable peer or timeout.
Firewall log method: filter the VPN gateway logs by the peer public IP and the time of the test. Start a fresh test and note the exact timestamp. This makes it much easier to distinguish the current attempt from historical failures.
3. Child SAs and traffic selectors
Once the IKE SA exists, check whether the Child SA is established. The Child SA is where the parameters for protected user traffic are agreed.
Common causes of failure include incompatible encryption/integrity proposals, PFS mismatch, incompatible lifetimes and traffic selectors that do not match the intended networks.
For a policy-based VPN, compare the local and remote subnet definitions on both peers. For example, if Site A proposes 10.10.0.0/16 → 10.20.0.0/16, Site B must define the reverse relationship. A selector mismatch can produce a tunnel that appears partly healthy but carries no useful traffic.
4. Routing and firewall policy
A working tunnel does not guarantee that a packet will enter it. On a route-based VPN, inspect the routing table and confirm the destination uses the tunnel interface. On a policy-based VPN, confirm the traffic matches the VPN policy/selector.
Then inspect the security policy. Most enterprise firewalls provide a policy hit counter or session/log entry showing the rule that matched. Record:
- Source and destination addresses.
- Application/protocol and port.
- Ingress and egress interfaces/zones.
- Security policy/rule ID.
- Action: allow or deny.
- NAT rule and translated addresses, if applicable.
A useful technique is to generate exactly one test connection and immediately inspect the corresponding session/log entry. This is much more reliable than searching thousands of historic records.
5. NAT and NAT-T
VPN traffic frequently fails because a general outbound NAT rule catches it before the VPN policy. Compare the original and translated source/destination addresses in the firewall session or NAT log.
NAT-T is different from ordinary source NAT. It is an IPsec mechanism that allows encrypted traffic to traverse a NAT device, normally using UDP/4500. A packet capture can therefore show UDP/4500 rather than native ESP when NAT-T is active.
When troubleshooting, identify every device between the VPN endpoints that can perform NAT. The apparent peer address in the VPN configuration may not be the address actually visible to the other side.
6. Collecting useful evidence
Good VPN troubleshooting is evidence-driven. Collect evidence from both ends whenever possible.
Firewall and VPN logs
Capture the logs for a narrowly defined test window. Useful sources include:
- IKE negotiation logs.
- IPsec/Child SA logs.
- Authentication and certificate logs.
- Security-policy traffic logs.
- NAT/session logs.
- Routing or tunnel-interface state.
- System/event logs showing interface or WAN changes.
Export the relevant entries rather than relying on a screenshot. Preserve timestamps and timezone information.
VPN status output
Record the tunnel state immediately before and after the test. Useful fields include peer address, IKE SA state, Child SA state, negotiated algorithms, selectors, packet counters and tunnel uptime.
Client-side evidence
For remote-access VPNs, record the assigned VPN address, DNS servers, routes, client version and connection time. Avoid collecting credentials, private keys or other secrets.
7. Packet captures
A packet capture is often the fastest way to determine where a VPN problem occurs. Ideally capture at both ends of the suspected path.
What to capture
- Before encryption: the original client/server traffic, where the capture point permits it.
- Between VPN gateways: IKE, NAT-T or ESP traffic.
- After decryption: the original traffic emerging from the VPN gateway.
Comparing these views answers questions that logs alone cannot. For example, if the client sends TCP/443 but the firewall never sees it, the problem is upstream of the firewall. If the firewall sees it but does not encrypt it, investigate routing, policy, selectors or NAT. If encrypted traffic leaves but nothing reaches the remote gateway, investigate the WAN path or remote edge.
Useful filters
For IPsec investigations, filter around the peer addresses and the expected control/data protocols. UDP/500 is commonly used for IKE and UDP/4500 for NAT-T. Native ESP is IP protocol 50. For the protected application, filter on the original source/destination and TCP/UDP port when you have a capture point before or after encryption.
Do not capture indefinitely. Use a short window around a controlled test and record the exact start time. A packet capture tool such as the Packet Capture Analyser can then help inspect the captured traffic.
What packet patterns tell you
| Observation | Likely direction of investigation |
|---|---|
| Client sends traffic; gateway sees nothing | Upstream routing, ACL, VLAN, local firewall or wrong gateway. |
| IKE leaves but no response returns | Peer reachability, upstream filtering, NAT or wrong peer address. |
| IKE exchange repeats and fails | Negotiation/authentication/identity mismatch. |
| IKE is established but no Child SA | IPsec proposal, PFS, selector or policy mismatch. |
| Child SA exists but packet counters stay at zero | Routing, policy, selectors or traffic not actually reaching the tunnel. |
| Encrypted packets leave one gateway but never arrive | WAN path, NAT or upstream filtering. |
| Both sides see encrypted traffic but application fails | Remote policy, routing, NAT, MTU or application issue. |
8. MTU and MSS
VPN encapsulation adds overhead. A path that handles small packets may fail for larger packets if fragmentation or Path MTU Discovery is broken.
Symptoms include successful ping and DNS tests followed by stalled web pages, file transfers or application sessions. Test progressively smaller packet sizes and inspect ICMP fragmentation-needed/packet-too-big messages and TCP MSS negotiation in a packet capture.
If MSS clamping is used, document where it is applied and why. Do not reduce MTU or MSS blindly: first establish the actual path and encapsulation overhead.
9. DNS and application testing
Once IP connectivity works, test DNS separately from the application. A useful sequence is:
- Ping or otherwise test the destination IP where appropriate.
- Resolve the hostname using the expected DNS server.
- Test the specific TCP/UDP port.
- Test the application itself.
This prevents a DNS failure from being mistaken for a VPN failure. For remote access, check whether internal DNS servers and search suffixes are actually installed on the client.
10. Common failure patterns
| Symptom | Start here |
|---|---|
| Tunnel never establishes | Peer reachability, IKE logs, proposals, authentication and identities. |
| IKE established, Child SA fails | Phase 2 proposal, PFS, selectors and lifetime. |
| Tunnel established, no traffic | Routing, security policy, NAT, selectors and packet counters. |
| One direction works | Asymmetric routing, return route, remote policy or NAT. |
| Only some subnets work | Selectors, routes and security policy scope. |
| Only large transfers fail | MTU, MSS and PMTUD. |
| VPN drops after a period | Lifetime/rekey, DPD/peer detection, WAN stability and idle timers. |
| Remote user connects but cannot resolve internal names | VPN DNS configuration and split-tunnel DNS routing. |
11. Evidence checklist
- Exact source and destination IPs.
- Protocol and port being tested.
- Exact test timestamp and timezone.
- VPN peer public addresses.
- IKE SA and Child SA state.
- Negotiated encryption, integrity, DH/PFS and selectors.
- Relevant firewall/security-policy logs.
- NAT/session information.
- Routing and tunnel-interface state.
- Packet capture from one or more useful capture points.
- Client routes and DNS settings for remote-access VPNs.
- Exact changes made during troubleshooting.