Chapter 22

Performance, memory, Safe Mode, upgrades, and troubleshooting

Use a small set of metrics and the first error to locate a problem. Distinguish queues, buffers, latency, rollback, and escalation, then troubleshoot within Safe Mode, backup, and version boundaries so every action remains reversible. This chapter performs offline checks only: it makes no connections and deploys nothing.

Why this chapter matters

“Node-RED is slow” could mean the editor is displaying too much Debug output, message objects are too large, a sequence buffer keeps growing, external I/O is blocked, or V8 is frequently collecting old space. It does not necessarily mean that memory is insufficient. Likewise, “it will not open after the upgrade” can indicate a failure in any layer: Add-on initialization, package installation, the runtime, the reverse proxy, the HA connection, or the flow schema.

This chapter is pinned to the behavior of Add-on 22.0.1 with Node-RED 5.0.2 and HA WebSocket nodes 0.80.3. It does not extrapolate from rolling documentation or other versions. Start with the symptom, timeline, and a bounded set of metrics, then narrow the scope one layer at a time. All network probes, HA calls, writes, package changes, and Deploy operations remain disabled or unexecuted.

Safety principle: Observe first, isolate next, change only one item at a time, and define a rollback point for every change. Before working in production, complete the Debug, Catch, and Status tests in Chapter 18 and the backup and restore drill in Chapter 21.

Locate faults by layer and timeline

LayerExample symptomCheck firstDo not do first
Supervisor / Add-onStartup loop or aborted initializationStartup time, first error, and recent option changesDo not clear settings or hide the first error with repeated restarts
Node.js / runtimeFrequent GC, event-loop latency, or process exitMemory trend, message rate, and old-space settingDo not set the heap close to the host’s total RAM
Flow / nodeDuplicate messages, buffer growth, or an unknown nodeNode status, Catch output, message shape, and Deploy timeDo not delete an unknown node or perform a Full Deploy
External I/OTimeouts, reconnects, or delayed dataDestination status, request duration, retry rate, and queue depthDo not run unauthorized probes or disable TLS verification
Reverse proxy / authenticationIngress opens, but an endpoint failsSeparate checks of Ingress/direct access, port, path, TLS, and authenticationDo not enable leave_front_door_open

The Add-on manifest builds only for aarch64 and amd64, and sets init=false. This manifest flag is not runtime Safe Mode. The image pins [email protected] and [email protected] directly, and includes the startup-support dependencies [email protected] and [email protected]. Confirm the affected layer and each version before diagnosing; do not report the Add-on version as the Node-RED version.

A step-by-step, reversible troubleshooting sequence

  1. Define one symptom and time window.

    Record when it began, which flow or route is affected, whether it is continuous or intermittent, and the last known good time. Replace URLs, entities, devices, areas, servers, MQTT topics, and client IDs with placeholders; do not collect complete payloads.

  2. Preserve the baseline without making changes.

    Record the Add-on, Node-RED, and HA WebSocket versions; startup time; CPU and memory trends; message and error rates; queue or buffer depth; and external-I/O latency. Capture only a limited number of lines around the error, then redact them so credentials, headers, cookies, and location data do not enter the ticket.

  3. Isolate external side effects.

    Disable Action, API, Fire Event, Update Config, HTTP Request, MQTT Out, File, email, Cast, InfluxDB, Modbus, serial, TCP/UDP/WebSocket output, and scheduled triggers. If a flow acts immediately at startup, use the Add-on’s safe_mode option to suppress its initial startup. Disable or disconnect every side-effect path before the first Deploy, because Deploy still starts deployed flows.

  4. Reproduce the issue with the smallest read-only path.

    Use a manual Inject node and small synthetic messages connected only to pure data-processing nodes such as Change, Switch, and Function, plus Debug nodes limited to specific fields. Add one node at a time to identify where latency, heap use, or queue depth begins to grow. Do not connect to the production HA instance or network.

  5. Change one controlled factor at a time.

    For example, limit Debug output first, then adjust a buffer or timeout, and only then evaluate old space. Record before-and-after metrics and the rollback value for each change. Confirm the restart scope before selecting Modified Nodes or Modified Flows; do not use Full Deploy as a routine troubleshooting tool.

  6. Roll back and observe.

    If the metrics do not improve, restore the previous value without stacking another change on top. If they do improve, first observe the same load for long enough in an isolated environment. Restoring network access, HA access, or writes in production requires separate approval and is not performed in this chapter.

End-of-chapter offline acceptance check

DeliverableExpected contentStop if
Incident worksheetOne redacted baseline, one falsifiable hypothesis, one changed factor, before-and-after results, and an explicit rollback value; external I/O and Deploy both remain unexecutedThe last known good time or first error is missing; multiple items changed together; there is no rollback value; evidence contains secrets or environment IDs; or interpretation requires a connection, deployment, or installation

Messages, buffers, context, and external I/O

Turn symptoms into comparable metrics first

SymptomRecord at leastCheck first
Unresponsive editorDebug messages per second, displayed length per message, and whether the sidebar stays openDisable unnecessary Debug nodes and output one sanitized field instead
Increasing latencyInput and output rates, processing time, and queue depthDelay rate limits, pending Trigger timers, and sequences waiting in Join
Stepwise memory growthProcess and container memory trends, message size, and pending-item countCloning of large objects, accumulating context, incomplete sequences, and retry queues
Periodic spikesSpike interval, schedules, and Poll State or Get History timingFlows triggered together, query range and result volume, and external timeouts
Missing or duplicate messages_msgid, msg.parts, and retry countSplit/Join boundaries, timeout paths, and the restart scope of a deployment

Debug output and message size

A Debug node should display only one sanitized field, not the complete message object. The Add-on settings template sets debugMaxLength to 1000 characters, but truncating the display neither removes a large upstream object nor proves that the complete object was not cloned or retained. Images, audio, very large JSON values, and live msg.req/msg.res objects do not belong in routine Debug output or context.

Node-RED may clone messages when they branch. Large nested objects, Buffers, and multiple branches magnify that cost. Pass only the fields required downstream, and do not create unbounded copies of historical arrays in a Function node. Use Function, Switch, Change, Range, Template, Exec, and RBE with a defined input shape and error path. Exec can run external commands and remains unexecuted in this chapter.

RBE (shown as Filter in the palette) can pass messages when a value changes or when a numeric deadband/narrowband rule permits them. It retains the previous comparison state within the configured topic boundary; it is not a stateless format conversion. Explicitly select the message property to compare, such as msg.payload, and the topic property. Define the deadband contract’s expected data type as a number. Because the runtime uses parseFloat, reject strings and objects first rather than allowing permissive parsing to accept them by accident. Acceptance tests must cover an unchanged value, values beyond and exactly at the threshold, different topics, msg.reset, and reconstruction of comparison state after a node restart or Deploy. A reset or restart can change whether the next message is emitted, so downstream side effects must not depend on untested prior state.

Delay, Trigger, and Join require explicit bounds

  • Delay: The rate limit and queue policy must handle peak load; input must not remain permanently faster than output. Whether to discard or retain intermediate values is a product decision that requires prior approval.
  • Trigger: Each topic or stream may retain a timer. If input keys are unbounded, the number of pending timers can also grow without bound. Set a verifiable timeout and a key allowlist.
  • Split / Join / Batch / Sort: Preserve correct msg.parts values, and limit the number of waiting sequences, items per group, and waiting time. The Add-on settings offer nodeMaxMessageBufferLength as an optional runtime limit. A value of 0 means unlimited and cannot be cited as evidence that a limit is in place.
  • HA Wait Until: Give every waiting message an explicit timeout and timeout branch; do not create unbounded waits. Avoid large numbers of simultaneous Time triggers.

Context and external I/O

Node, flow, and global context can use memory or a named localfilesystem store. Do not put unbounded arrays, complete events, or binary payloads in context. A disk-backed store is not durably written on every assignment. Define limits for keys, size, retention, and write frequency first.

HTTP, MQTT, WebSocket, TCP, UDP, serial, files, databases, and HA Get History or Poll State are all external I/O. Set a timeout, result-size limit, retry limit, and backoff policy for requests. Limit the time range and result count for Get History, and do not use Poll State where event subscriptions are appropriate. Replace every host, URL, broker, topic, server, and ID with a complete placeholder, and make no connections during diagnosis. See Chapter 20 for MQTT boundaries and Chapter 19 for HTTP routing.

max_old_space_size limits only V8 old space

The Add-on option max_old_space_size is an integer in MB. The startup script exports the following only when the option is present:

NODE_OPTIONS=--max_old_space_size=PLACEHOLDER_MB

This line is shown only for an exact comparison with runtime behavior; it is not a shell command to run in this chapter. It limits the old-memory section of the Node.js V8 heap, not total Node.js process memory, the container limit, or host RAM. The process also uses memory for the young generation, code, native add-ons, Buffers and other external allocations, and overhead. A value that is too high can starve the Home Assistant host; one that is too low can increase garbage collection or cause a heap failure.

Require evidence before and after an adjustment

  1. First record representative trends for process and container memory, message rate, latency, restarts, and queues.
  2. Fix unbounded Debug output, messages, context, Delay/Trigger/Join behavior, and external I/O first. More old space does not fix a leak or an unbounded buffer.
  3. If a justified need remains, choose a conservative value in an isolated environment, change only that setting, and repeat the same load.
  4. Observe GC, latency, peaks, and host headroom. If there is no improvement, roll back rather than continuing to increase the value.
There is no universal value: A safe value depends on the host, other Add-ons, flows, and load. Recommending a “best” number of MB without measurements would be fabrication.

Safe Mode, Deploy scope, and privacy-safe logs

The Add-on option safe_mode: true makes the startup script add --safe. The Node-RED runtime starts, but suppresses only the initial startup of flows so that you can repair them. Any Deploy can start the flows included in that deployment, even while safe_mode remains true. This is not Home Assistant Safe Mode, it does not stop the Add-on, and it is not an access-control mechanism. The editor and administration interface must still comply with the existing Ingress and direct-access authentication boundaries.

# Exact Add-on 22.0.1 option; this chapter does not apply it or restart the Add-on
safe_mode: true

This setting does not prove that a flow is safe. Before the first Deploy, disable nodes with network, HA, write, scheduling, or other side effects, or disconnect the wires leading to them. Do not Deploy first and remediate afterward: a post-Deploy check cannot undo actions that have already occurred.

The safe sequence is to begin with a restorable backup. During a maintenance window, an authorized operator can then enable Safe Mode and restart the Add-on. Without deploying, repair or disable unknown nodes, unbounded buffers, and every side-effect path. Treat the first Deploy as an operation that will start deployed flows, and approve it separately. Disabling Safe Mode can also run a flow immediately at the next startup, so merely opening the editor does not mean the work is complete.

Inject, Debug, Catch, Status, Complete, Link In/Out/Call, Comment, and Junction serve different purposes. Catch receives errors; Status receives node status; Complete indicates that a monitored node has finished processing an input, but does not guarantee that every downstream operation has finished. Link nodes alter message routing, while Comment and Junction must not be mistaken for runtime measurement points. Enable only the observation nodes needed for diagnosis so that observation itself does not create excessive load.

Deploy scope is not a performance switch

The Node-RED 5.0.2 editor provides Full, Modified Flows, and Modified Nodes deployment scopes. The selected scope determines which flows or nodes restart, which can affect timers, connections, context lifecycles, and in-flight messages. Understand the dependencies of the change before choosing the smallest correct scope. If the effect of a shared config node is unclear, do not select Modified Nodes merely because it seems faster.

Log only what is necessary

The log_level schema accepts trace, debug, info, notice, warning, error, and fatal, and the option may be omitted. The official documentation recommends info for normal use. The wrapper maps warning to Node-RED’s warn. Greater verbosity can expose payloads, URLs, headers, entity/device/area IDs, MQTT topics, file paths, or stack traces. Use it only for a defined period, then return to the necessary level after troubleshooting.

When sharing logs, retain only versions, timestamps, node types, error categories, and a correlation placeholder. Remove tokens, passwords, Authorization and Cookie headers, encrypted credentials, private keys, webhooks, location data, real hostnames and IP addresses, complete payloads, and environment details visible in screenshots. [email protected] can improve stack traces, but it does not sanitize sensitive values.

Upgrades, initialization, and reversible boundaries

Before upgrading

  • Complete an Add-on backup and isolated restore drill. Preserve the credential_secret or Project secret separately; confirm that node_modules is excluded and preserve a dependency inventory.
  • Export a sanitized flow inventory. Record the Add-on, Node-RED, and HA WebSocket versions; additional npm and system packages; and the optional and not bundled @flowfuse/[email protected]. Its metadata declares Node >=14 and Node-RED >=3.0.0, but that compatibility range does not replace testing in the target environment.
  • Read the target release notes and breaking changes, then confirm compatibility with HA, Node.js, the architecture, and all nodes. This chapter’s baseline establishes only 22.0.1/5.0.2/0.80.3 behavior; it makes no guarantee about other versions.
  • Disable HA, network, write, and scheduling nodes. Define stop conditions, the rollback version, and the backup to restore. An upgrade requires separate approval and is not performed in this chapter.

The exact 22.0.1 startup sequence

StageBehavior in this versionSafe response to failure
Data initializationIf a new /config has no settings, data may be migrated from /homeassistant/node-red; otherwise, the settings, flows, and nodes directories are created.Preserve inventories of both directories before doing anything else. Do not move or overwrite files manually; find the first migration error.
Theme migrationThe old dark theme is updated to dark-modern.Recognize this as a name migration; do not misdiagnose an appearance change as flow corruption.
Conflicting packagesThe wrapper attempts to remove node-red-contrib-home-assistant, node-red-contrib-home-assistant-llat, and node-red-contrib-home-assistant-ws.This is runtime-script behavior; do not issue a separate removal command. If it fails, preserve the log and package inventory.
Custom packagesAlpine system_packages are installed first, followed by npm npm_packages; any failure aborts startup.Find the first package, repository, or architecture error. Roll back the most recently added controlled option without clearing the data directory.
Custom commandsEach line of init_commands is passed to eval at every startup; any failure aborts startup.Keep the most recently added command disabled and return to known-good settings. Never put secrets in logs or commands.
Runtime flagssafe_mode becomes --safe; max_old_space_size becomes NODE_OPTIONS.Check that each option exists and has the correct type. Do not mistake the manifest’s init=false for Safe Mode.

After an upgrade and during rollback

In Safe Mode and without deploying, first verify the editor, node types, credential decryption, settings, and package loading. Only after every side-effect node is disabled or disconnected may a Deploy of the smallest pure-data flow be approved separately, because that Deploy will start deployed content. Do not start production flows while any unknown node, schema-migration warning, route error, or HA connection error remains unresolved.

Rollback means more than downgrading the image. Stop further changes, preserve sanitized logs and state from the failed version, and use the preapproved Home Assistant Add-on restore procedure to return to a matching backup, version, and secret. Do not feed data migrated by the new version directly to an older runtime, and do not manually edit the node type or version fields in flow JSON.

Safe, symptom-led responses

  • Add-on startup stops at a custom package or command: Find the first system_packages, npm_packages, or init_commands error in the log. Compare recent option changes and the architecture. Roll back only the latest change; do not run an ad hoc shell command, delete /config, or paste secrets into a command.
  • An HA node reports “disconnected” or “Unauthorized WebSocket”: First confirm that the Add-on still uses the preconfigured connection mode. The official known issue requires the “I use the Home Assistant Add-on” setting to be correct in the HA Server configuration. Check the 0.80.3 prerequisites: HA 2024.3+, Node-RED >=3.1.1, and Node >=18.2.0. These differ from the HA 2023.3.0 installation threshold in the Add-on manifest. Do not copy tokens, switch to an unencrypted URL, or send a test Action.
  • An HTTP node or Dashboard route is not found: First distinguish Ingress from direct access. The Add-on documentation states that HTTP nodes require a separately mapped network port and use paths under /endpoint/. Check the direct port, TLS certificate and key, path, and http_node authentication. Do not disable TLS or enable leave_front_door_open.
  • TLS handshake or certificate error: Perform only controlled checks of the filename, validity period, subject name, chain, and key/certificate match; the files must be under /ssl. Do not expose a private key, disable verification, or confuse Ingress TLS with the direct-access ssl option.
  • Unknown nodes or schema warnings appear after import or upgrade: Stay in Safe Mode. Do not Deploy, delete nodes, or edit their type manually. Identify the pinned package from the backup inventory and restore a compatible version in an isolated environment first. The legacy HA Entity node is deprecated; handle it only through the migration procedure for the pinned version, and do not create a new one.
  • Join or Wait Until emits nothing, or memory keeps growing: Sample msg.parts, sequence size per group, timeouts, pending keys, and input/output rates. Stop new input and preserve the smallest useful evidence first. Do not inject more test messages or clear production context directly.
  • An upgrade report or escalation is required: Stop making repeated changes. Provide the appropriate Add-on or node-package maintainer with pinned versions, architecture, a reproduction timeline, the first error, whether the problem remains when side effects are disabled, a minimal sanitized flow, and the single change already attempted. Never attach tokens, credentials, private keys, real endpoints or IDs, full payloads, browser storage, or unredacted screenshots.

Pinned sources

The options, scripts, nodes, and compatibility claims in this chapter rely only on the following exact commits and corresponding official documentation:

Frequently asked questions

Can max_old_space_size be set to the host’s available RAM?
No. It limits only V8 old space, not the whole process, container, or other services on the host. Find unbounded messages, buffers, or context first, then evaluate a conservative value under an isolated load and with adequate host headroom.
Does Safe Mode stop the entire Add-on?
No. It starts Node-RED with --safe and suppresses only the initial startup of flows. Any Deploy can start deployed flows even while safe_mode remains true. Disable or disconnect every side-effect path before the first Deploy.
Does truncating messages in the Debug sidebar reduce their memory cost?
Not necessarily. A display-length limit affects only the display; it does not prove that an upstream object was never created, cloned, or buffered. Reduce the message shape at its source, and limit both Debug frequency and fields.
Can I delete and recreate an unknown node?
Do not do so. Deletion loses the original configuration, and Deploy may start other flows. Remain in Safe Mode, identify the package from the pinned-version inventory, and restore a compatible node in an isolated environment or follow the official migration procedure.
Repeated restarts hid the first error and left only later timeouts. Which evidence matters?
Stop restarting. Treat the first error from the earliest startup, the last known good time, and the option differences as primary evidence; label later timeouts only as secondary symptoms. If the original record is gone, do not guess the root cause or stack additional fixes. Maintain isolation and reconstruct the smallest reproducible timeline through the established escalation path.