Skip to main content

Temporal Failures reference

View Markdown
info

For the causes the Temporal Service records when a Workflow Task fails, see the Workflow Task errors reference. For symptom-based guidance, see Troubleshooting Workflow Execution failures.

A Failure is Temporal's representation of an error in a Workflow, Activity, or Nexus Operation. Each type of Failure has its own class or error type in each SDK and its own information in the protobuf messages. The SDKs use these messages to communicate with the Temporal Service, and they appear in the Event History.

Failures are defined in the failure messages of the Temporal gRPC API. Each Failure type on this page lists its class or error type in each SDK and its proto message.

Temporal Failure​

Most SDKs have a base class that the other Failure types extend.

The base Failure proto message has these fields:

FieldDescription
messageThe error message.
stack_traceThe stack trace of the error.
sourceThe SDK the Failure came from, such as "TypeScriptSDK". Some SDKs use this field to rebuild the call stack into an exception object.
causeThe Failure message for the cause of this Failure, if there is one.
encoded_attributesThe encoded message and stack_trace fields when you use a Failure Converter that encodes them.

Application Failure​

Workflow, Activity, and Nexus Operation code use Application Failures to report application-specific errors. Application Failure is the only Failure type that your code creates and throws.

Errors in Workflows​

An error in a Workflow causes either a Workflow Task Failure or a Workflow Execution Failure. A Workflow Task Failure retries the Workflow Task. A Workflow Execution Failure closes the Workflow Execution with a Failed status.

Only exceptions that are Temporal Failures fail the Workflow Execution. All other exceptions fail the Workflow Task, and the Workflow Task is retried. In Go, any error the Workflow returns fails the Workflow Execution, and a panic fails the Workflow Task.

Temporal raises most Failure types for you, such as a Cancelled Failure when the Workflow is canceled or an Activity Failure when an Activity fails. To fail the Workflow Execution from your Workflow Definition, throw an Application Failure. In Go, return any error.

Workflow Task Failures​

A Workflow Task Failure means the Worker couldn't process a Workflow Task. It happens when your Workflow code throws an exception that isn't a Temporal Failure, or panics in Go. The Temporal Service retries the Workflow Task until the Workflow Execution Timeout, which is unlimited by default.

For the causes the Temporal Service records for each Workflow Task Failure, see the Workflow Task errors reference.

Workflow Execution Failures​

Throw an Application Failure in a Workflow to fail the Workflow Execution. The Workflow Execution moves to the Failed state, and the Temporal Service makes no more attempts to progress it.

To create a custom exception that fails the Workflow Execution, extend the Application Failure class for your SDK.

Errors in Activities​

To fail an Activity Task, throw an Application Failure or any other error. The SDK converts any other error to an Application Failure and sets these fields:

FieldValue
typeThe error's type name.
messageThe error message.
non_retryablefalse
detailsUnset.
causeA Failure converted from the error's cause property.
next_retry_delayUnset.

The SDK also copies the call stack.

When an Activity Execution fails, the Application Failure from the last Activity Task becomes the cause field of the Activity Failure. The Workflow's call to the Activity throws the Activity Failure, and your Workflow Definition can handle it.

Errors in Nexus Operations​

A Nexus Operation ends in one of four states: completed, failed, canceled, or timed out.

The caller's Nexus service splits an Operation into one or more StartOperation requests and completion callbacks. It retries these requests as long as they fail with retryable errors.

The Operation times out only when the schedule-to-close timeout set by the caller Workflow expires. The caller's Nexus service enforces this timeout.

The Operation reaches one of the other three states when one of these happens:

  • The Operation handler returns a synchronous response or error.
  • An asynchronous Operation, such as one backed by a Workflow, reaches a terminal state.

A Nexus Operation handler returns a retryable or non-retryable error to tell the caller's Nexus service whether to retry the request. If a request times out before the handler sends a response, the caller retries it.

Errors are retryable by default. These errors aren't retried:

  • Non-retryable Application Failures.
  • Unsuccessful Operation errors, which resolve the Operation as failed or canceled.
  • Handler errors of type BAD_REQUEST, UNAUTHENTICATED, UNAUTHORIZED, NOT_FOUND, or NOT_IMPLEMENTED.

Nexus Operation Task Failures​

A Nexus Operation Task Failure means the handler couldn't process a Nexus Operation Task. It happens when your Nexus handler code throws an unknown error. The Nexus Operation Task is retried.

Nexus Operation Execution Failures​

To fail the whole Nexus Operation Execution, throw a non-retryable Application Failure from the Nexus Operation handler. The Nexus Operation Execution moves to the Failed state, and no more attempts are made to complete it.

Propagation of Workflow errors​

When a Workflow started by a Nexus NewWorkflowRunOperation handler throws an Application Failure, the error reaches the caller as a non-retryable error. The Nexus Operation Execution fails.

Failures in a Nexus handler​

To fail a single Nexus Operation Task or the whole Nexus Operation Execution, throw an Application Failure, a Nexus error, or any other error from the handler.

The SDK converts unknown errors to a retryable Application Failure and sets these fields:

FieldValue
non_retryablefalse
typeThe error's type name.
messageThe error message.

Retryable failures​

The caller retries retryable Nexus Operation Task failures, such as an unknown error, with a built-in Retry Policy. When a Nexus Task fails, the caller Workflow records the failed attempt on the pending Nexus Operation and sets these fields:

FieldValue
stateThe new state, such as BackingOff.
attemptThe attempt count, incremented by one.
next_attempt_schedule_timeWhen the Nexus Task is retried.
last_attempt_failureThe error message in message and the Application Failure in failure_info.

For example, the Temporal CLI shows an unknown error thrown in a Nexus handler like this:

temporal workflow describe -w my-workflow-id
...
Pending Nexus Operations: 1

Endpoint myendpoint
Service my-hello-service
Operation echo
OperationToken
State BackingOff
Attempt 6
ScheduleToCloseTimeout 0s
NextAttemptScheduleTime 20 seconds from now
LastAttemptCompleteTime 11 seconds ago
LastAttemptFailure {"message":"unexpected response status: "500 Internal Server Error": internal error","applicationFailureInfo":{}}

Non-retryable​

When an Activity or Workflow throws an Application Failure, Temporal compares the Failure's type field to the Retry Policy's list of non-retryable errors. If the type is in the list, the Activity or Workflow isn't retried. To stop retries regardless of the Retry Policy, set the Application Failure's non_retryable field to true.

When a Nexus Operation handler throws an Application Failure, the caller retries it with a built-in Retry Policy that cannot be customized. To stop retries, set the Application Failure's non_retryable field to true. A non-retryable error from a Nexus handler fails the Nexus Operation Execution, and the caller's Workflow Execution receives it as a Nexus Operation Failure.

Next Retry Delay​

Set the Next Retry Delay on an Application Failure to control how long Temporal waits before it retries the Activity or Workflow. This delay overrides the interval the Retry Policy would have calculated for that failure.

Nexus errors​

Default mapping​

A Nexus Operation handler that throws an Application Failure returns one of these Nexus errors, depending on non_retryable:

non_retryableNexus errorHTTP status code
false (default)HandlerErrorTypeInternal500 Internal Server Error
trueUnsuccessfulOperationError424 Failed Dependency

Use Nexus errors directly​

Throw a Nexus error from your Nexus Operation handler instead of an Application Failure. A Nexus error carries the retry behavior listed in the following tables and maps to a more specific HTTP status code for external Nexus callers that support it.

For example, the Nexus Go SDK provides these errors:

  • nexus.HandlerError(nexus.HandlerErrorType, msg)
  • nexus.UnsuccessfulOperationError{state, failure}

Retryable Nexus errors​

Nexus error typenon_retryable
HandlerErrorTypeResourceExhaustedfalse
HandlerErrorTypeInternalfalse
HandlerErrorTypeUnavailablefalse

Non-retryable Nexus errors​

Nexus error typenon_retryable
HandlerErrorTypeBadRequesttrue
HandlerErrorTypeUnauthenticatedtrue
HandlerErrorTypeUnauthorizedtrue
HandlerErrorTypeNotFoundtrue
HandlerErrorTypeNotImplementedtrue
UnsuccessfulOperationErrortrue

Cancelled Failure​

When Cancellation of a Workflow, Activity, or Nexus Operation is requested, each SDK represents it in its own way. For example, in TypeScript, some Workflow API functions throw a Cancelled Failure directly, and others wrap it in a different Failure. Use the TypeScript isCancellation helper to check for both.

When a Workflow, Activity, or Nexus Operation is canceled, the Cancelled Failure is the cause field of the Activity Failure, Nexus Operation Failure, or "Workflow failed" error.

Activity Failure​

The Workflow Execution receives an Activity Failure when an Activity fails. It carries details about the Activity Execution, such as the Activity Type and Activity ID. The cause field holds the reason for the failure. For example, if the Activity Execution times out, the cause is a Timeout Failure.

Nexus Operation Failure​

The Workflow Execution receives a Nexus Operation Failure when a Nexus Operation fails. It carries details about the Nexus Operation Execution, such as the Operation name and Operation token. The message and cause fields hold the reason for the failure. The cause is usually an Application Failure or a Cancelled Failure.

A Nexus Operation Failure has these fields:

FieldValue
endpointThe name of the Nexus Endpoint.
serviceThe name of the Nexus Service.
operationThe name of the Operation.
operation_tokenThe Operation token, set for an asynchronous Operation. Use it to act on the Operation, such as to cancel it.
scheduled_event_idThe ID of the Event in the caller's Event History that scheduled the Operation.
messageA generic error message for an unsuccessful Operation.
causeThe underlying Application Failure, with non_retryable set to true, type set to the error's type name, and message set to the error message.
nexus_error_codeThe underlying Nexus error code.

Child Workflow Failure​

The Workflow Execution receives a Child Workflow Failure when a Child Workflow Execution fails. It carries details about the Child Workflow Execution, such as the Workflow Type and Workflow ID. The cause field holds the reason for the failure.

Timeout Failure​

A Timeout Failure represents the timeout of an Activity or Workflow. When an Activity times out, the Timeout Failure includes the last Heartbeat details the Activity sent.

Terminated Failure​

When a Workflow is terminated, a Terminated Failure is the cause of the error you receive in these places:

  • In a parent Workflow that's waiting for the result of a Child Workflow.
  • In a Client that's waiting for the result of a Workflow.

The SDK and proto types are:

Server Failure​

A Server Failure represents an error that comes from the Temporal Service.