Skip to main content

Error Handling and Retry

Interfaces fail. What separates a maintainable interface from a fragile one is deciding, in advance, which failures stop the feed and which ones do not.

Classify the failure first

CategoryExampleRetry helps?Response
TransientTimeout, connection reset, receiver restartingYesRetry with backoff
DataMissing required field, malformed payloadNoReject the message; the sender must fix it
ConfigurationWrong port, bad credentials, missing certificateNoFix the configuration
Business rejectionReceiver returns a negative ACK or a 409DependsFollow the receiver's contract

Retrying a data error just fails repeatedly at a slower rate. Refusing to retry a transient error turns a two-second network blip into a manual restart.

Raising an error in a script

Use error() when the message cannot be processed correctly:

local mrn = Msg.PID[3][1][1]:value()
if mrn == nil or mrn == "" then
error("PID-3 patient identifier is missing")
end

Name the field, not its value. Error text is stored and searchable, so a value in an error message becomes a patient identifier in a log.

WriteNot
error("PID-3 patient identifier is missing")error("bad MRN: " .. mrn)
error("receiver returned " .. resp.code)error("failed sending " .. Data)

Filtering is not an error

Returning without pushing drops a message deliberately. That is normal operation:

if MsgType ~= "ADT" then
print("Skipping message type " .. tostring(MsgType))
return -- not an error
end

Print a line when you do. Otherwise a working filter and a broken node look identical to whoever asks why nothing arrived.

Never swallow a failure

The most damaging pattern in an interface is treating a failed delivery as success:

-- Wrong: the message is lost and nobody knows
local resp = linkiir.link.web.post{ url = Url, body = Data }

-- Right: check both the transport and the response
local resp, err = linkiir.link.web.post{ url = Url, body = Data }
if not resp then
error("delivery failed: " .. err.message)
end
if resp.code >= 400 then
error("receiver returned " .. resp.code)
end

resp being present means the request completed, not that the receiver accepted it. Check the status code too.

Stopping versus continuing

When a script raises an error, the node reports the failure and stops processing. Unprocessed messages wait in the queue.

That is the safe default for clinical interfaces, and it is worth understanding why: stopping is not data loss. The queue is durable, order is preserved, and the failure is impossible to miss. A node that silently continues past errors can discard a day of data before anyone notices.

Where individual bad records are genuinely expected and must not halt a feed, handle them in the script rather than letting them raise:

local linkiir = require("linkiir")

function main(Data)
local ok, Msg = pcall(function()
return linkiir.data.extract{ schema = "adt.json", data = Data, type = "hl7" }
end)

if not ok then
-- Record it and move on, rather than stopping the feed.
print("Unparseable message skipped")
return
end

linkiir.flow.push{ data = Msg:text() }
end

Doing it this way makes the decision explicit in the script, where a reviewer can see it, and lets you distinguish the cases you tolerate from the ones you do not. Print a line whenever you swallow something, or a working filter and a broken node look identical to whoever asks why nothing arrived.

For destination-side rejections, the equivalent decision is a node field — Ack Error Handling on Destination LLP. See Destination Nodes.

Retry at the node, not in the script

Transport nodes already handle reconnection. Configure it rather than writing retry loops:

NodeFields
Destination LLPAttempt to Reconnect, Reconnect Attempts, Reconnection Interval, Resend on Ack Timeout, Resend Attempts, Ack Timeout
Source or Destination File/FTP, with FTP enabledAttempt to Reconnect, Reconnect Limit Times, Reconnection Interval

For an outbound call from a script, let the error propagate rather than looping. A retry loop inside main holds the worker and hides the failure from the node's state.

What happens when a node stops

GuaranteeDetail
Messages are not lostUnprocessed messages wait in the queue
Order is preservedDelivery resumes in order when the node restarts
No silent gapThe failure appears in the node state and in log search

This is why stopping is a safe response to an error. You are pausing a durable queue, not dropping traffic.

Duplicates are possible

Delivery is at-least-once. A message can be delivered twice if processing is interrupted between doing the work and confirming it.

Design the receiver's side for it:

DestinationProtection
Destination LLPOriginal ID and message type ACK verification; receiver deduplicates on message control ID
Destination File/FTPUnique ID file naming; FTP Overwrite Handling set to not be uploaded
Outbound HTTPA stable idempotency key derived from the message, if the API supports one
Database writeAn upsert or guarded insert keyed on a message identifier

Decide how duplicates are detected before go-live. A replay must not create a second clinical or financial transaction.

Investigating a failure

  1. Open the node and read its status detail — it carries the last error.
  2. Search log search for ERROR level records in the project.
  3. Open the failing record and note its correlation ID.
  4. Search that correlation ID to see how far the message got.
  5. Copy the archived payload into a sample and reproduce it with Run Test.

Step 5 is the one that saves time. Reproducing with the exact payload beats reasoning about what the sender might have sent.

Replay rather than rewind

To reprocess a message, replay that specific archived payload from log search.

Rewinding a queue consumer to an earlier position reprocesses everything from that point, which usually means re-delivering messages that already succeeded. Replay targets one message.

ActionScopeUse for
Replay an archived messageOne messageNormal recovery
Bulk replayA selected setA known batch that failed together
Rewind a consumer positionEverything from that pointRarely; only with a full understanding of the side effects

Confirm the receiver tolerates a duplicate before replaying anything.

Safe context to record

Include enough to diagnose without exposing patient data.

Safe: project, workflow, and node names; message ID and correlation ID; error category and code; retry attempt; queue position; timestamps.

Never: raw HL7, FHIR, CDA, or X12 payloads in general service logs; patient name, MRN, date of birth, health card number, address, or phone number in error text, metric labels, or email alerts.

Message payloads are archived deliberately, with access control, in the Log DB. That is where they belong — not in an error string. See Security.

Next