Skip to main content
Latest Documentation
This is the latest documentation for the Cloud Posse Reference Architecture. To determine which version you're currently using, please see Version Identification.

datadog-logs-archive

This component provisions Datadog Log Archives. It creates a single log archive pipeline for each AWS account. If the catchall flag is set, it creates a catchall archive within the same S3 bucket.

Each log archive filters for the tag env:$env where $env is the environment/account name (e.g. sbx, prd, tools), as well as any tags identified in the additional_query_tags key. The catchall archive, as the name implies, filters for *.

A second bucket is created for CloudTrail, and a CloudTrail is configured to monitor the log archive bucket and log activity to the CloudTrail bucket. To forward these CloudTrail logs to Datadog, the CloudTrail bucket's ID must be added to the s3_buckets key for our datadog-lambda-forwarder component.

Both buckets support object lock, with overridable defaults of COMPLIANCE mode and a duration of 7 days.

Prerequisites

  • Datadog integration set up in the target environment
    • Relies on the Datadog API and App keys added by our Datadog integration component

Issues, Gotchas, Good-to-Knows

  • Destroy/reprovision process
    • Because of the protections for S3 buckets, destroying/replacing the bucket may require two passes or a manual bucket delete followed by Terraform cleanup. If the bucket has a full day or more of logs, deleting it manually first helps avoid Terraform timeouts.
    • Two-step process to destroy via Terraform:
      1. Set s3_force_destroy to true and apply
      2. Set enabled to false and apply, or run terraform destroy

CloudTrail KMS Encryption​

By default, this component creates a KMS key to encrypt CloudTrail logs for compliance and security. The KMS encryption can be configured using these variables:

  • cloudtrail_enable_kms_encryption (default: true) - Enable/disable KMS encryption for CloudTrail logs
  • cloudtrail_kms_key_arn (default: null) - Provide an existing KMS key ARN to use instead of creating a new one
  • cloudtrail_create_kms_key (default: true) - Create a new KMS key when cloudtrail_kms_key_arn is not provided
  • cloudtrail_kms_key_deletion_window_in_days (default: 10) - KMS key deletion window (7-30 days)
  • cloudtrail_kms_key_enable_rotation (default: true) - Enable automatic KMS key rotation

The created KMS key includes the required policy statements for CloudTrail to encrypt logs and for authorized principals to decrypt them.

Archive Tuning​

Four settings control how Datadog writes the archive and how much a search or rehydration is allowed to read back:

  • compression_method (default: ZSTD) - Compression Datadog uses when writing objects, either ZSTD or GZIP. ZSTD is the default and the recommendation in the Datadog console, especially where Archive Search is used: the objects are smaller, so they cost less to store, less to scan, and less in egress from the archive bucket. Note the Datadog provider itself defaults to GZIP.
  • partitioning_attributes (default: null) - Up to two low cardinality attributes used as partition keys, most frequently queried first. Logs sharing a partition value are co-located, so a search can skip partitions that cannot match before downloading them.
  • lookup_attributes (default: null) - Up to two high cardinality attributes, such as a trace, container or user ID, used to pinpoint individual logs within a data block.
  • rehydration_max_scan_size_in_gb (default: null) - Maximum volume, in GB, that a single job may scan against this archive. Despite the field name, which predates Archive Search, this one setting caps Archive Search queries and rehydration jobs alike. Left unset, a single wide search can scan the entire archive.

Partitioning is worth more than it first appears. Datadog applies the query filter after the matching files are downloaded, so scan size is driven by the length of the searched time range rather than by how selective the query is: a narrow filter over a wide window still scans the whole window. Partition attributes are the only mechanism that prunes files before download.

compression_method, partitioning_attributes and lookup_attributes are forward-only: they change how Datadog writes objects, so only logs archived after the setting is applied are affected and objects already written cannot be retrofitted. Changing compression_method on an existing archive only affects new files; objects already stored keep their original format and stay readable, since the two formats use different extensions.

rehydration_max_scan_size_in_gb is not forward-only. It is applied at query time, so it bounds every subsequent Archive Search and rehydration job, including jobs that read objects written before the limit was set.

These settings require the Datadog provider >= 4.13.0, the first release carrying all four attributes.

Excluding Logs From the Archive​

query_exclusions subtracts from the archive query, producing (<query>) -(<fragment> OR <fragment> ...).

It exists because neither existing input can express "archive everything except this". query_override replaces the whole query, including the AWS account id that is only resolved at apply time, so using it to add a negation means restating the query and hardcoding the account id. additional_query_tags appends with OR, where a negation matches nearly everything.

query_exclusions:
- '@http.useragent:"ELB-HealthChecker/2.0"'
- '@http.url_details.path:("/health" OR "/ready" OR "/livez")'

Worth knowing: index exclusion filters do not stop archiving. Per the Datadog documentation, excluded logs are discarded from indexes but still flow through Live Tail and are still archived. So health checks and similar high volume, low value traffic can be filtered out of the index and still fill the bucket. Combining an index exclusion filter with the matching query_exclusions fragment leaves that traffic visible in Live Tail for real time debugging while keeping it out of both the index and the archive.

The exclusions apply to the catchall archive too when catchall_enabled is true. Without that they would be pointless there, since the catchall query is * and would pick the excluded logs straight back up into the same bucket under /catchall.

Be deliberate about what goes in here. Dropping probe traffic is a cost decision; dropping audit logs such as CloudTrail or Kubernetes audit events is a compliance decision, and once a log is neither indexed nor archived nothing retains it.

Sponsorship​

This project is supported by the Datadog Open Source Program.

As part of this collaboration, Datadog provides a dedicated sandbox account that we use for automated integration and acceptance testing. This contribution allows us to continuously validate changes against a real Datadog environment, improving reliability and reducing the risk of regressions.

We are grateful to Datadog for supporting our open source ecosystem and helping ensure that infrastructure code for Terraform remains stable and well-tested


Usage​

Stack Level: Global

It's suggested to apply this component to all accounts from which Datadog receives logs.

Example Atmos snippet:

components:
terraform:
datadog-logs-archive:
settings:
spacelift:
workspace_enabled: true
vars:
enabled: true
# additional_query_tags:
# - "forwardername:*-dev-datadog-lambda-forwarder-logs"
# - "account:123456789012"

Variables​

Required Variables​

region (string) required

AWS Region

Optional Variables​

access_log_bucket_enabled (bool) optional

Whether to create a dedicated S3 bucket for CloudTrail bucket access logs


Default value: false

access_log_bucket_name (string) optional

Name of existing S3 bucket to use for CloudTrail bucket access logs. Only used when access_log_bucket_enabled is false


Default value: ""

additional_query_tags (list(any)) optional

Additional tags to be used in the query for this archive


Default value: [ ]

archive_lifecycle_config optional

Lifecycle configuration for the archive S3 bucket


Type:

object({
abort_incomplete_multipart_upload_days = optional(number, null)
enable_glacier_transition = optional(bool, true)
glacier_transition_days = optional(number, 365)
glacier_transition_storage_class = optional(string, "GLACIER_IR")
noncurrent_version_glacier_transition_days = optional(number, 30)
enable_deeparchive_transition = optional(bool, false)
deeparchive_transition_days = optional(number, 0)
noncurrent_version_deeparchive_transition_days = optional(number, 0)
enable_standard_ia_transition = optional(bool, false)
standard_transition_days = optional(number, 0)
expiration_days = optional(number, 0)
noncurrent_version_expiration_days = optional(number, 0)
})

Default value: { }

archive_name (string) optional

Name of the Datadog logs archive. Datadog logs archive names must be unique within a Datadog organization, so this defaults to the globally unique module ID (module.this.id) when null.


Default value: null

catchall_archive_name (string) optional

Name of the catchall Datadog logs archive. Datadog logs archive names must be unique within a Datadog organization, so this defaults to &lt;module.this.id&gt;-catchall when null.


Default value: null

catchall_enabled (bool) optional

Set to true to enable a catchall for logs unmatched by any queries. This should only be used in one environment/account


Default value: false

cloudtrail_create_kms_key (bool) optional

Create a new KMS key for CloudTrail encryption. Only used if cloudtrail_kms_key_arn is not provided and cloudtrail_enable_kms_encryption is true


Default value: true

cloudtrail_enable_kms_encryption (bool) optional

Enable KMS encryption for CloudTrail logs


Default value: true

cloudtrail_kms_key_arn (string) optional

ARN of an existing KMS key to use for CloudTrail log encryption. If not provided and cloudtrail_enable_kms_encryption is true, a new key will be created


Default value: null

cloudtrail_kms_key_deletion_window_in_days (number) optional

Duration in days after which the KMS key is deleted after destruction of the resource, must be between 7 and 30 days


Default value: 10

cloudtrail_kms_key_enable_rotation (bool) optional

Enable automatic rotation of the KMS key


Default value: true

cloudtrail_lifecycle_config optional

Lifecycle configuration for the cloudtrail S3 bucket


Type:

object({
abort_incomplete_multipart_upload_days = optional(number, null)
enable_glacier_transition = optional(bool, true)
glacier_transition_days = optional(number, 365)
noncurrent_version_glacier_transition_days = optional(number, 365)
enable_deeparchive_transition = optional(bool, false)
deeparchive_transition_days = optional(number, 0)
noncurrent_version_deeparchive_transition_days = optional(number, 0)
enable_standard_ia_transition = optional(bool, false)
standard_transition_days = optional(number, 0)
expiration_days = optional(number, 0)
noncurrent_version_expiration_days = optional(number, 0)
})

Default value: { }

compression_method (string) optional

Compression method Datadog uses when writing objects to the archive. One of ZSTD or GZIP.


Defaults to ZSTD, which is the default and the recommendation in the Datadog console,
especially where Archive Search is used. ZSTD objects are smaller than GZIP, so they cost
less to store, less to scan (Archive Search and rehydration are both billed on the volume
scanned) and less in egress from the archive bucket. The Datadog provider defaults to GZIP.



Default value: "ZSTD"

lifecycle_rules_enabled (bool) optional

Enable/disable lifecycle management rules for log archive s3 objects


Default value: true

lookup_attributes (list(string)) optional

Up to two high cardinality attributes (trace ID, container ID, user ID) used to pinpoint
individual logs within a data block, reducing both the volume scanned and egress from the
archive bucket.


Only logs archived after this is set benefit. Null disables lookup acceleration.



Default value: null

object_lock_days_archive (number) optional

Object lock duration for archive buckets in days


Default value: 7

object_lock_days_cloudtrail (number) optional

Object lock duration for cloudtrail buckets in days


Default value: 7

object_lock_mode_archive (string) optional

Object lock mode for archive bucket. Possible values are COMPLIANCE or GOVERNANCE


Default value: "COMPLIANCE"

object_lock_mode_cloudtrail (string) optional

Object lock mode for cloudtrail bucket. Possible values are COMPLIANCE or GOVERNANCE


Default value: "COMPLIANCE"

partitioning_attributes (list(string)) optional

Up to two low cardinality attributes used as partition keys for the archive, most frequently
queried first. Logs sharing a partition value are co-located, so a search can skip partitions
that cannot match before downloading them.


This is the only setting that decouples scan size from the length of the searched time range.
The query filter is applied after the matching files are downloaded, so an unpartitioned
archive scans the whole window regardless of how selective the query is.


Only logs archived after this is set are partitioned. Null leaves the archive unpartitioned.



Default value: null

query_exclusions (list(string)) optional

Query fragments to subtract from the archive query, combined as
(&lt;query&gt;) -(&lt;fragment&gt; OR &lt;fragment&gt; ...).


Unlike query_override, this keeps the derived query intact, including the AWS account id that
is only known at apply time, so an archive can drop a class of logs without the caller having to
restate the whole query. Unlike additional_query_tags, which appends with OR, these fragments
are negated as a group.


Logs matching a fragment are still ingested and still visible in Live Tail; they are only kept
out of the archive. Pair this with an index exclusion filter to keep them out of the index too,
since exclusion filters on their own do not stop archiving.


Applies to the catchall archive as well when catchall_enabled is true. Without that, an
excluded log would match the catchall query * and be written to the same bucket under
/catchall.


Null or empty leaves both queries untouched.



Default value: null

query_override (string) optional

Override query for datadog archive. If null would be used query 'env:{stage} OR account:{aws account id} OR {additional_query_tags}'


Default value: null

rehydration_max_scan_size_in_gb (number) optional

Maximum volume, in GB, that a single job may scan against this archive.


Despite the field name, which predates Archive Search, this is one per-archive setting that
caps Archive Search queries and rehydration jobs alike. Null means no limit, so a single wide
search can scan the entire archive and bill the corresponding egress.



Default value: null

s3_force_destroy (bool) optional

Set to true to delete non-empty buckets when enabled is set to false


Default value: false

Context Variables​

The following variables are defined in the context.tf file of this module and part of the terraform-null-label pattern.

additional_tag_map (map(string)) optional

Additional key-value pairs to add to each map in tags_as_list_of_maps. Not added to tags or id.
This is for some rare cases where resources want additional configuration of tags
and therefore take a list of maps with tag key, value, and additional configuration.


Required: No

Default value: { }

attributes (list(string)) optional

ID element. Additional attributes (e.g. workers or cluster) to add to id,
in the order they appear in the list. New attributes are appended to the
end of the list. The elements of the list are joined by the delimiter
and treated as a single ID element.


Required: No

Default value: [ ]

context (any) optional

Single object for setting entire context at once.
See description of individual variables for details.
Leave string and numeric variables as null to use default value.
Individual variable settings (non-null) override settings in context object,
except for attributes, tags, and additional_tag_map, which are merged.


Required: No

Default value:

{
"additional_tag_map": {},
"attributes": [],
"delimiter": null,
"descriptor_formats": {},
"enabled": true,
"environment": null,
"id_length_limit": null,
"label_key_case": null,
"label_order": [],
"label_value_case": null,
"labels_as_tags": [
"unset"
],
"name": null,
"namespace": null,
"regex_replace_chars": null,
"stage": null,
"tags": {},
"tenant": null
}
delimiter (string) optional

Delimiter to be used between ID elements.
Defaults to - (hyphen). Set to &#34;&#34; to use no delimiter at all.


Required: No

Default value: null

descriptor_formats (any) optional

Describe additional descriptors to be output in the descriptors output map.
Map of maps. Keys are names of descriptors. Values are maps of the form
\{<br/> format = string<br/> labels = list(string)<br/> \}
(Type is any so the map values can later be enhanced to provide additional options.)
format is a Terraform format string to be passed to the format() function.
labels is a list of labels, in order, to pass to format() function.
Label values will be normalized before being passed to format() so they will be
identical to how they appear in id.
Default is {} (descriptors output will be empty).


Required: No

Default value: { }

enabled (bool) optional

Set to false to prevent the module from creating any resources
Required: No

Default value: null

environment (string) optional

ID element. Usually used for region e.g. 'uw2', 'us-west-2', OR role 'prod', 'staging', 'dev', 'UAT'
Required: No

Default value: null

id_length_limit (number) optional

Limit id to this many characters (minimum 6).
Set to 0 for unlimited length.
Set to null for keep the existing setting, which defaults to 0.
Does not affect id_full.


Required: No

Default value: null

label_key_case (string) optional

Controls the letter case of the tags keys (label names) for tags generated by this module.
Does not affect keys of tags passed in via the tags input.
Possible values: lower, title, upper.
Default value: title.


Required: No

Default value: null

label_order (list(string)) optional

The order in which the labels (ID elements) appear in the id.
Defaults to ["namespace", "environment", "stage", "name", "attributes"].
You can omit any of the 6 labels ("tenant" is the 6th), but at least one must be present.


Required: No

Default value: null

label_value_case (string) optional

Controls the letter case of ID elements (labels) as included in id,
set as tag values, and output by this module individually.
Does not affect values of tags passed in via the tags input.
Possible values: lower, title, upper and none (no transformation).
Set this to title and set delimiter to &#34;&#34; to yield Pascal Case IDs.
Default value: lower.


Required: No

Default value: null

labels_as_tags (set(string)) optional

Set of labels (ID elements) to include as tags in the tags output.
Default is to include all labels.
Tags with empty values will not be included in the tags output.
Set to [] to suppress all generated tags.
Notes:
The value of the name tag, if included, will be the id, not the name.
Unlike other null-label inputs, the initial setting of labels_as_tags cannot be
changed in later chained modules. Attempts to change it will be silently ignored.


Required: No

Default value:

[
"default"
]
name (string) optional

ID element. Usually the component or solution name, e.g. 'app' or 'jenkins'.
This is the only ID element not also included as a tag.
The "name" tag is set to the full id string. There is no tag with the value of the name input.


Required: No

Default value: null

namespace (string) optional

ID element. Usually an abbreviation of your organization name, e.g. 'eg' or 'cp', to help ensure generated IDs are globally unique
Required: No

Default value: null

regex_replace_chars (string) optional

Terraform regular expression (regex) string.
Characters matching the regex will be removed from the ID elements.
If not set, &#34;/[^a-zA-Z0-9-]/&#34; is used to remove all characters other than hyphens, letters and digits.


Required: No

Default value: null

stage (string) optional

ID element. Usually used to indicate role, e.g. 'prod', 'staging', 'source', 'build', 'test', 'deploy', 'release'
Required: No

Default value: null

tags (map(string)) optional

Additional tags (e.g. {&#39;BusinessUnit&#39;: &#39;XYZ&#39;}).
Neither the tag keys nor the tag values will be modified by this module.


Required: No

Default value: { }

tenant (string) optional

ID element (Rarely used, not included by default). A customer identifier, indicating who this instance of a resource is for
Required: No

Default value: null

Outputs​

access_log_bucket_arn

The ARN of the bucket used for CloudTrail bucket access logs

access_log_bucket_domain_name

The FQDN of the bucket used for CloudTrail bucket access logs

access_log_bucket_id

The ID (name) of the bucket used for CloudTrail bucket access logs

archive_id

The ID of the environment-specific log archive

bucket_arn

The ARN of the bucket used for log archive storage

bucket_domain_name

The FQDN of the bucket used for log archive storage

bucket_id

The ID (name) of the bucket used for log archive storage

bucket_region

The region of the bucket used for log archive storage

catchall_id

The ID of the catchall log archive

cloudtrail_bucket_arn

The ARN of the bucket used for access logging via cloudtrail

cloudtrail_bucket_domain_name

The FQDN of the bucket used for access logging via cloudtrail

cloudtrail_bucket_id

The ID (name) of the bucket used for access logging via cloudtrail

cloudtrail_kms_key_alias

The alias of the KMS key used for CloudTrail log encryption (only if created by this module)

cloudtrail_kms_key_arn

The ARN of the KMS key used for CloudTrail log encryption

cloudtrail_kms_key_id

The ID of the KMS key used for CloudTrail log encryption (only if created by this module)

Dependencies​

Requirements​

  • terraform, version: >= 1.1.5
  • aws, version: >= 4.9.0, < 6.0.0
  • datadog, version: >= 4.13.0
  • http, version: >= 2.1.0

Providers​

  • aws, version: >= 4.9.0, < 6.0.0
  • datadog, version: >= 4.13.0
  • http, version: >= 2.1.0

Modules​

NameVersionSourceDescription
archive_bucket4.11.0cloudposse/s3-bucket/awsn/a
bucket_policy2.0.2cloudposse/iam-policy/awsn/a
cloudtrail0.24.0cloudposse/cloudtrail/awsn/a
cloudtrail_access_log_bucket4.11.0cloudposse/s3-bucket/awsn/a
cloudtrail_access_log_bucket_label0.25.0cloudposse/label/nulln/a
cloudtrail_bucket_label0.25.0cloudposse/label/nulln/a
cloudtrail_s3_bucket4.11.0cloudposse/s3-bucket/awsn/a
datadog_configurationv1.535.13github.com/cloudposse-terraform-components/aws-datadog-credentials//src/modules/datadog_keysn/a
iam_rolesv1.536.1github.com/cloudposse-terraform-components/aws-account-map//src/modules/iam-rolesn/a
this0.25.0cloudposse/label/nulln/a

Resources​

The following resources are used by this module:

Data Sources​

The following data sources are used by this module: