Abstract#
Acoustic environments are often treated either as physical exposures or as sources of subjective meaning. Both views are necessary, but neither is sufficient for studying deliberately composed acoustic conditions. This integrative review asks how the physical, perceptual and contextual properties of acoustic environments influence attention, arousal, affect, regulation and behaviour, and what methodological conditions are required to study designed environments as repeatable exposures. Evidence concerning environmental noise is comparatively mature: unwanted exposure is associated with annoyance, sleep disturbance, cardiovascular risk and some cognitive outcomes. Soundscape research adds a complementary object—the acoustic environment as perceived or experienced by people in context—but evidence for salutary or restorative effects is smaller, more heterogeneous and often dependent on stimulus selection, comparison condition and setting. Auditory attention, reward and music-regulation research provide plausible mechanisms and measurable constructs, yet they do not justify a generalized diagnosis of “sound addiction.” Likewise, heterogeneous findings from music and sound interventions cannot be transferred wholesale to designed acoustic environments or interpreted as therapeutic efficacy.
The paper therefore distinguishes six objects that are often collapsed: digital signal, playback exposure, room acoustic field, perceived condition, behavioural or physiological response, and clinical outcome. It proposes a reproducibility ladder from narrative description to human response and treats IFAH only as a candidate instrument for specifying, preserving and replaying acoustic recipes. An exact-release software audit finds strong version gating and parameter reconstruction, but only partial evidence for deterministic rendering; bit-identical live output, calibrated playback equivalence, room-field identity and repeated human response remain unestablished. A research agenda is proposed around calibrated exposure, contextual moderation, preregistered outcomes and independent replication. The central conclusion is methodological: deliberately composed sound can become a credible research object only when the acoustic condition, its presentation and the inference drawn from it are documented separately.
Keywords: acoustic environment; soundscape; environmental noise; auditory attention; music reward; regulation; reproducibility; Web Audio; research instrument
1. The invisible condition#
Sound is physically measurable and experientially elusive. A microphone can record pressure variation; a digital system can store samples; a loudspeaker can convert an electrical signal into acoustic energy. None of these operations, by itself, specifies what reached a listener, how that listener organized the scene, what the sounds meant in that situation, or what changed in attention, physiology or behaviour. Yet research and product language repeatedly compress this chain. “The same audio” is said to recreate “the same environment”; a calming rating becomes evidence of regulation; a plausible neural mechanism becomes evidence of efficacy; a repeated preference becomes “addiction.” Each compression makes a story easier to tell and harder to test.
Acoustic environments are not scientifically marginal. Environmental-noise research links exposure to annoyance, disturbed sleep, cardiovascular risk and cognitive outcomes, with evidence varying by source and endpoint [1–4]. Soundscape research has widened the object of study from unwanted level alone to the acoustic environment as perceived, experienced or understood by people in context [5–9]. Auditory-scene and attention research shows that a listener does not receive the environment passively: sound features are grouped, selected and suppressed over time [10–12]. Music cognition and intervention research further shows that auditory experiences can carry reward, memory, affective and regulatory functions [13–21]. These literatures establish that acoustic conditions matter. They do not establish that every designed condition has a stable effect, that pleasantness is therapeutic, or that the human response is reproducible because the source file is reproducible.
The term acoustic condition is used here for a specified arrangement of sound sources, signal properties, timing, spatial presentation and exposure parameters within a defined playback and situational context. It is deliberately narrower than soundscape, because a soundscape includes perception and context; and broader than audio file, because an exposure depends on transduction, level, space and listener position. A deliberately composed condition might include generated tones, noise bands, recorded material, temporal rules, spatial routing, reverberation and interactive parameters. Calling it a condition does not imply that it is beneficial. It makes the intended exposure available for description and test.
This paper asks:
How do the physical, perceptual and contextual properties of acoustic environments influence human attention, arousal, affect, regulation and behaviour—and what methodological conditions are required to study deliberately composed acoustic environments as repeatable exposures?
The paper is an integrative review and methodological perspective, not a systematic review and not an efficacy report. It connects evidence families that use different objects, outcomes and standards. Its governing distinction is:
The arrows are causal possibilities, not identities. Signal equality does not guarantee equal playback. Equal playback at a device does not guarantee an equal field at the ears. A similar field does not guarantee similar perception, and similar perception does not guarantee a repeated behavioural, physiological or clinical response. This separation is the paper’s analytic backbone.
The public title, Sound, Addiction & Environment, is therefore interrogative. “Addiction” names a contemporary anxiety about persistent stimulation and learned auditory cues; it is not the paper’s premise. The scientific frame is Acoustic Conditions, Stimulation and Human Regulation. Altered states of consciousness are outside scope because no construct, eligible induction condition, validated measure or safety protocol has been specified. IFAH also remains outside the explanatory argument until Section 8, where it is assessed as a candidate research instrument rather than introduced as a therapeutic solution.
2. From noise to soundscape#
2.1 The established pole: environmental noise as exposure#
Environmental-noise research begins with a defensible asymmetry: unwanted sound can be quantified as exposure, while its effects depend on source, timing, duration, sensitivity and outcome definition. The World Health Organization’s environmental-noise guidance synthesizes associations with annoyance, sleep disturbance, ischaemic heart disease, hypertension, hearing effects and cognitive impairment, while noting that evidence quality is not uniform across sources and endpoints [1,2]. Basner and colleagues’ review similarly describes auditory and non-auditory effects and emphasizes sleep fragmentation, stress pathways and cardiovascular consequences [3]. The European Environment Agency’s 2025 assessment shows that chronic transport-noise exposure remains a large population-health issue in Europe [4].
The strongest conclusion is not that “noise is bad” in the abstract. It is that specified exposures, particularly transport noise, can impose measurable population burdens. WHO’s evidence reviews illustrate why calibration matters. Road-traffic noise and some sleep outcomes measured with polysomnography have moderate-quality evidence, whereas several self-reported, child or source-specific sleep outcomes are lower certainty [1]. Evidence for cognitive outcomes varies by source, task and age group [1,22]. Cardiovascular evidence is most developed for road traffic noise and ischaemic heart disease, but even here the inference is epidemiological and exposure modelling is imperfect [23]. The credible statement is therefore outcome-specific: mature evidence supports harms from some noise exposures; it does not license a universal acoustic dose–response rule.
This literature supplies indispensable variables for designed-condition research: sound pressure level, frequency content, intermittency, temporal pattern, exposure duration and time of day. It also supplies a warning. If a study reports only the name of a stimulus—“nature sound,” “ambient music,” “sound bath”—then another researcher cannot know the dose or reproduce the condition. Conversely, a calibrated level alone does not explain why one source is noticed, welcomed or resisted.
2.2 The conceptual turn: perception and context enter the object#
ISO 12913-1 defines soundscape in relation to an acoustic environment as perceived, experienced or understood by a person or people in context [5]. This is not a cosmetic replacement of noise with a more positive word. It changes the object. The physical environment remains necessary, but the listener, activity, expectations, place and meaning become part of the inquiry. Parts 2 and 3 of the ISO series provide guidance on data collection, reporting and analysis [6,7]. A review of ISO-based studies found rapid adoption of the standardized perceptual attributes, particularly Method A, but rare strict compliance with the full standards [8]. A 2025 systematic review of soundscape context identified 14 theoretical and 52 empirical studies and described context through spatial–physical and socio-cultural components, while finding uneven coverage and continuing measurement fragmentation [9].
Perceived evaluation should therefore be treated as depending jointly on the source signal, playback exposure, room or spatial field, listener factors, situational context and temporal history. This is not a fitted causal model; it is a methodological reminder that perception cannot be inferred from the source signal alone. Noise sensitivity, hearing status, familiarity, expectations, social interaction, activity and perceived control may all alter evaluation [8,9,24]. In a large in-situ Montreal dataset, social interaction, noise sensitivity and extraversion were associated with soundscape judgments, with effect sizes generally small but non-zero [25]. The result is useful precisely because it frustrates a universal-listener model.
Indoor soundscape work reinforces the point. A systematic review of residential studies found varied assessment methods and recurring importance of source type, building characteristics, activities and personal factors [26]. In a controlled residential mock-up study, 35 participants rated 20 scenarios combining indoor sources and urban environments. Principal-component analysis yielded dimensions labelled Comfort, Content and Familiarity; acoustic indicators improved prediction, but semantic source categories added explanatory value [27]. The experiment did not prove a universal indoor model. It showed that level and meaning jointly structure evaluation.
2.3 Positive soundscapes: narrower and more heterogeneous evidence#
The turn to soundscape creates space to ask whether some acoustic environments support restoration or wellbeing. The evidence is promising but considerably smaller than the noise-harm literature. Aletta and colleagues identified only seven eligible studies in a 2018 review of positive soundscape and health relationships; the studies used different physiological and self-report outcomes and were synthesized qualitatively [28]. Buxton and colleagues’ 2021 systematic review identified 36 publications on natural sounds, with 18 entering meta-analysis. Across varied comparisons, natural sounds were associated with health and positive-affect outcomes and with reduced stress or annoyance, but subgroup cells were sparse and exposure methods heterogeneous [29].
Individual experiments are informative without being decisive. Alvarsson and colleagues exposed 40 participants to a stressor followed by one of four sound conditions. Skin-conductance recovery was faster during a nature-sound condition than during several noise conditions, while high-frequency heart-rate variability did not differ [30]. This supports a condition-specific recovery effect in a small laboratory design, not a general claim that natural sound heals. A 2024 meta-analysis comparing natural sounds with quiet found significant changes in some physiological measures but no significant pooled effect on subjective stress, illustrating how the comparator and outcome family can reverse a simple narrative [31].
The phrase positive soundscape should therefore describe appraisal or a design aim, not a guaranteed health intervention. A pleasant condition may reduce annoyance, mask an unwanted source, support a desired activity or be preferred to a control. Clinical significance requires a clinical endpoint, a suitable comparator, adequate duration and population, and a design capable of excluding expectation and demand effects.
2.4 What the transition permits—and what it does not#
Moving from noise to soundscape permits researchers to study experience without abandoning acoustics. It permits multiple legitimate outcomes: loudness and source identification; pleasantness and eventfulness; task performance; self-reported restoration; heart rate, electrodermal activity or sleep; and, where appropriate, clinical measures. These outcomes should remain separate. A change in pleasantness is not a proxy for a change in blood pressure; a momentary physiological difference is not necessarily meaningful regulation; and neither is automatically a clinical benefit.
The evidence also does not support an inverse-noise fallacy:
If uncontrolled noise can harm, it does not follow that deliberately designed sound will heal.
Design increases control over candidate variables. It can improve documentation, manipulate features and create repeatable contrasts. Benefit remains an empirical result.
3. Attention, arousal and embodied listening#
3.1 Auditory organization before outcome#
An acoustic field contains concurrent and successive energy patterns, not pre-labelled objects. Auditory-scene analysis describes how the nervous system groups features into streams and sources. Shamma, Elhilali and Micheyl proposed temporal coherence as a major principle: features whose neural responses covary over time are more likely to bind into an auditory object, while attention enhances selected streams [10]. The account is a theoretical synthesis supported by psychoacoustic and neurophysiological findings, not a complete model of everyday listening. It nonetheless identifies manipulable features—onset synchrony, modulation, spectral relations and temporal coherence—that can change perceptual organization without changing total energy.
Attention is not a single funnel. A listener may intentionally select a source, be captured by an unexpected event, or have serial recall disrupted by changing background sound. The duplex-mechanism account of auditory distraction distinguishes interference produced by a changing-state sequence from attentional capture by a deviant event [11]. These mechanisms predict different patterns and should not be compressed into “sound reduces focus.” A temporally varying but predictable scene may interfere with order processing; an isolated salient cue may redirect attention; a stable low-information background may have little effect for one task and matter for another.
This has immediate methodological consequences. Studies should state whether attention is the target of deliberate selection, a task-performance outcome, a subjective rating, or an inferred mechanism. They should also distinguish a sustained condition from discrete auditory events. A notification experiment, for example, can show that learned sounds interrupt an ongoing visual task. Upshaw and colleagues paired notification or control sounds with a visual attention task and observed slower responses and altered event-related potentials under notification conditions [12]. That is evidence of cue-specific distraction. It does not show addiction to sound, and it cannot be generalized to continuous composed environments.
3.2 Arousal is not valence#
Arousal and pleasantness are orthogonal enough to demand separate measurement. An acoustic condition may be pleasant and activating, unpleasant and activating, pleasant and low-arousal, or dull without being calming. The ISO soundscape framework’s pleasantness and eventfulness dimensions express this separation at an environmental-appraisal level [5–8]. Physiological measures add further dimensions but do not solve interpretation automatically. Heart rate, heart-rate variability, electrodermal activity, respiration and pupil responses have different time courses and sensitivities. The same change can reflect orienting, effort, emotional arousal, movement or respiration depending on design.
“Regulation” should therefore refer to a process or outcome that is explicitly operationalized. At least four uses must be distinguished:
- Subjective regulation: a reported change in felt stress, mood or ability to settle.
- Behavioural regulation: a change in task persistence, response inhibition, movement or sleep behaviour.
- Physiological regulation: a preregistered pattern in an autonomic or endocrine measure, interpreted with necessary controls.
- Clinical regulation: a meaningful change on a validated health outcome in an appropriate population.
These levels may converge, but convergence is a finding. Alvarsson et al.’s faster electrodermal recovery without an HF-HRV difference is a useful example [30]. A single phrase—“the nervous system calmed”—would erase the actual pattern.
3.3 Context, control and the body#
Listening is embodied in a literal sense: exposure reaches ears through a head and body in a physical space, and response occurs during an activity. Hearing thresholds, tinnitus, hyperacusis, age, fatigue, neurodivergence, medication, prior noise exposure and cultural familiarity can alter both tolerance and meaning. Posture and movement change spatial cues; respiration can influence both experience and physiological endpoints; headphones and loudspeakers create different fields. These are not nuisance details to be averaged away by default.
Perceived control is particularly important. The same sound may be sought voluntarily, accepted as useful, or experienced as intrusion. Soundscape context research repeatedly identifies activity, expectation, social setting and place as moderators [9,24–27]. Intervention studies also suggest that self-selection and researcher selection are not interchangeable, although preference can be confounded with familiarity and expectancy [18–21]. A design that lets participants adjust a condition tests a person–system interaction; a fixed exposure tests a standardized stimulus. Both can be valid, but they answer different questions.
Multisensory context complicates ecological validity further. A visually congruent scene can change source interpretation and environmental appraisal. Removing the visual, social and spatial context may improve control while changing the phenomenon. Laboratory research therefore needs two forms of fidelity: internal specification, which makes an exposure repeatable, and ecological correspondence, which asks whether responses relate to real-world listening. These goals may conflict, and neither should be used as a synonym for the other.
3.4 Measurement discipline#
A credible study of a designed acoustic condition should choose a primary outcome family before exposure. Subjective measures can capture the experience directly but are vulnerable to expectancy, scale framing and demand. Behavioural tasks can reveal attention or performance but may not represent rest or recovery. Physiological measures offer temporal resolution but require artefact control and restrained interpretation. Clinical outcomes have direct relevance but demand stronger designs, populations and intervention definitions.
The strongest near-term work is likely triangulated rather than maximal: one preregistered subjective outcome, one behavioural or physiological outcome tied to a specific hypothesis, and manipulation checks for audibility, pleasantness, eventfulness and familiarity. Adding ten weakly justified measures does not create depth. It creates undisclosed researcher degrees of freedom.
4. Stimulation, reward and habituation#
4.1 Why “addiction” must remain a question#
Persistent headphone use, constant background media and notification-rich environments make addiction an understandable public metaphor. A scientific diagnosis has a higher threshold. The World Health Organization describes disorders due to addictive behaviours as clinically significant syndromes involving distress or interference with functioning. Its gaming-disorder criteria include impaired control, increasing priority over other activities, continuation or escalation despite negative consequences and significant impairment [32,33]. Repetition, pleasure, craving-like language, dopamine involvement or discomfort with silence do not by themselves meet that threshold.
Music clearly recruits reward processes. Salimpoor and colleagues combined PET and fMRI to relate intense musical pleasure and anticipation to dopamine release in striatal regions [13]. Ferreri and colleagues used a double-blind within-participant pharmacological design with 27 healthy adults and found opposite changes in musical pleasure and motivation after dopaminergic enhancement and antagonism [14]. These studies support a causal role for dopamine in aspects of musical reward. They do not diagnose addiction. Dopamine participates in learning, movement, motivation and ordinary reward; its presence is not a pathology marker.
Direct evidence for “music addiction” is sparse. Schmuziger and colleagues compared 50 non-professional pop/rock musicians with 50 matched controls using criteria adapted from substance-dependence instruments; more musicians met the adapted threshold [34]. The study makes dependence-like loud-music behaviour empirically discussable, but its exploratory, cross-sectional design, selected population and adapted measures cannot validate a generalized disorder. Hazardous acoustic dose is also a distinct issue: WHO’s safe-listening standard addresses risk from sound level and duration, not compulsive behaviour [35].
Better-developed adjacent constructs exist. The Healthy–Unhealthy Music Scale was developed across three adolescent samples and distinguishes music engagement associated with wellbeing from engagement associated with rumination, stress and depressive vulnerability [15]. Broader reviews of music engagement and mental health find both adaptive and maladaptive patterns, with much of the evidence correlational and heterogeneous [16]. Auditory cues can also acquire learned significance and capture attention, as notification studies demonstrate [12]. In each case, the object may be affect regulation, a social message, a device-checking loop, identity or avoidance—not sound itself.
The publication-safe conclusion is:
Contemporary acoustic life may involve persistent stimulation, learned cue-reactivity, reward-seeking and maladaptive listening practices; current evidence does not justify collapsing these phenomena into a generalized addiction to sound.
4.2 Habituation, prediction and dose#
Repeated exposure can produce habituation, sensitization, learning, preference or boredom depending on the stimulus and listener. These processes should not be conflated. Habituation is a reduced response to repeated stimulation under specified conditions; tolerance in addiction is a diagnostic-style construct involving diminished effect or escalation in a broader syndrome. Familiarity can increase liking through learning, but repeated unwanted exposure can maintain annoyance or sleep disturbance. Prediction can reduce surprise without removing energetic masking or physiological disruption.
For designed conditions, repetition is both a variable and a threat to inference. An initial response may reflect novelty. A repeated-session response may weaken through habituation, strengthen through learned association, or vary with day and state. A single exposure can characterize immediate appraisal; it cannot establish durable regulation. Longitudinal designs should record dose, adherence, concurrent listening and whether participants self-select other audio between sessions.
4.3 A construct map for future work#
Rather than starting with addiction, future studies can separate four testable branches:
| Branch | Example construct | Suitable evidence | Boundary |
|---|---|---|---|
| Reward | pleasure, wanting, anticipation | ratings, choice, effort, validated reward tasks | Reward is not pathology. |
| Cue-reactivity | orienting, checking urge, attentional capture | reaction time, eye/EEG measures, cue-specific behaviour | The cue may signal another rewarding object. |
| Regulation | mood management, rumination, avoidance | validated regulation/use scales, experience sampling | Adaptive and maladaptive use may coexist. |
| Hazardous exposure | level–duration dose, hearing symptoms | dosimetry, audiology, symptom measures | Acoustic harm is not behavioural addiction. |
Only if a coherent syndrome of impaired control, priority, persistence despite harm and functional impairment emerged would an addiction construct become appropriate. That investigation would require clinical and behavioural-addiction expertise, validated instruments, differential diagnosis and an ethics framework. Research 01 does not cross that threshold.
5. Sound and music in care#
5.1 Categories that must not be collapsed#
Evidence about sound in care spans practices with different agents, materials and causal claims. Music therapy is a credentialed professional practice in which music is used within a therapeutic relationship to pursue individualized goals [36,37]. Music medicine commonly refers to pre-recorded or researcher-selected music delivered in a medical setting without the same relational process. Music-based interventions is a broad umbrella that can include active music making, receptive listening, individualized programmes and standardized playlists. Environmental sound interventions may use natural, broadband, masking or spatial material without being music. A designed acoustic condition may be generated by rules and may not have conventional musical form.
These categories cannot donate evidence to one another without qualification. If a trial depends on therapist attunement, song choice, biography and interaction, its result does not estimate the effect of a fixed recording. If an exposure is a fixed perioperative playlist, it does not test a responsive generative environment. Reporting guidance for music-based interventions has emerged precisely because intervention content, provider, dose, tailoring and fidelity are often incompletely described [38,39].
5.2 What the intervention literature supports#
The intervention literature is large enough to show that music-based interventions can affect selected outcomes, and heterogeneous enough to preclude a universal efficacy claim. Hole and colleagues synthesized 72 randomized trials involving nearly 7,000 adults undergoing surgery. Music was associated with lower postoperative pain and anxiety, reduced analgesic use and greater satisfaction, but not shorter length of stay [17]. Blinding was difficult, intervention forms varied and the findings concern a defined perioperative context.
Meta-analyses of music interventions for stress report small-to-moderate or moderate effects across psychological and physiological outcomes, along with substantial heterogeneity and moderator effects [18,19]. The number of included studies and participants strengthens the conclusion that some music interventions can alter stress-related outcomes. It does not identify a single effective acoustic recipe. Studies differ in active versus receptive form, provider, preference, duration, population, comparator and measurement timing.
Sleep evidence is similarly conditional. A 2024 systematic review of adults with mental-health problems found 15 eligible studies and potential benefits for subjective sleep, with smaller meta-analytic subsets and risk-of-bias concerns [20]. A separate review of adults without diagnosed sleep disorder found no overall improvement across six studies and judged the evidence low quality, while suggesting duration-dependent effects in subgroups [21]. The contrast warns against writing “music improves sleep” without population, dose and outcome.
Updated Cochrane syntheses add necessary constraint. For people with dementia, music-based therapy may reduce depressive symptoms in some comparisons but probably has little or no effect on agitation or aggression relative to usual care; longer-term and comparator-specific findings remain uncertain [40]. In cancer care, 81 studies involving 5,576 participants suggested possible benefits across several outcomes, with music therapy tending to yield more consistent results, but evidence certainty and outcome coverage varied [41]. These are not failures of the field. They show why the intervention, comparison and endpoint must stay visible.
5.3 Why calming is not clinical efficacy#
Three inferential steps are routinely skipped:
- A participant prefers or describes a condition as calming.
- A short-term subjective or physiological measure changes.
- A clinically meaningful outcome improves.
Each step requires evidence. Pleasantness can be a valid endpoint when the question concerns appraisal or acceptability. Acute recovery can be a valid endpoint when measured against a stressor and control. Clinical efficacy requires an appropriate population, validated outcome, credible comparator, adequate follow-up and analysis of harm. A laboratory condition should not be marketed with language borrowed from clinical trials before those steps are completed.
This boundary is especially important for individualized systems. Personal control and preference may increase adherence or immediate liking; they also make the intervention less standardized and can unblind the hypothesis. The right response is not to eliminate personalization automatically, but to specify its rule. A study can compare fixed versus user-adaptive conditions, or treat selection behaviour as part of the intervention. It cannot describe a freely adjusted experience as though every participant received the same exposure.
5.4 Reporting as part of the intervention#
The Reporting Guidelines for Music-Based Interventions (RG-MBI) and its checklist identify components such as intervention rationale, content, delivery, provider, recipient, setting, schedule, tailoring and fidelity [38,39]. For acoustic-condition research, the same principle must be extended down the signal chain: source assets or algorithms, renderer version, random seeds, sample rate, channel layout, level, device, room, position and deviations. The intervention is not fully identified until another team can determine both what was intended and what was presented.
6. Natural, urban and deliberately designed conditions#
6.1 Refusing the nature–city moral binary#
Natural sounds often perform well against transport noise in laboratory comparisons [29–31]. This does not establish a universal hierarchy in which nature is regulating and the city is dysregulating. “Natural” includes gentle water, dense insect choruses, thunder, wind and predator cues; “urban” includes traffic, voices, ventilation, bells and valued social activity. Source label, level, temporal structure, biodiversity, familiarity and meaning all vary within the categories. A quiet urban courtyard may be preferred to a loud natural scene. A water sound may mask traffic effectively for one listener and evoke discomfort for another.
The comparator is decisive. Natural sound versus loud road noise tests replacement of an aversive exposure and may combine changes in level, spectrum, semantics and control. Natural sound versus silence tests a different question. Natural sound mixed with noise tests masking and scene composition. Meta-analysis can pool only at the cost of such distinctions, and sparse subgroup cells make broad category effects unstable [29].
6.2 Urban and indoor soundscapes as activity systems#
Urban sound is intertwined with public life. Voices can signal safety, sociability or crowding. Traffic may be unwanted but informative. Eventfulness can support a marketplace and undermine a reading room. Soundscape research therefore includes appropriateness—fit between the acoustic environment and place or activity—rather than treating pleasantness as the sole goal [5–9]. The Montreal in-situ work showing differences associated with social interaction is a reminder that users partly constitute the evaluated environment [25].
Indoor conditions add building envelopes, mechanical systems, privacy and neighbour sounds. Residential research suggests that control and source meaning matter alongside loudness [26,27]. A deliberately designed indoor condition will interact with this existing field rather than replace it cleanly. Researchers should measure background level and report salient external sources; otherwise, a “fixed” intervention is mixed with an undocumented environment.
6.3 Designed conditions as parameterized exposures#
A deliberately designed condition can combine recorded and synthesized materials, parameterized timing, spatial relations and adaptive rules. Its scientific advantage is not intrinsic efficacy but variable addressability. Features can be changed independently: event density, spectral centroid, modulation depth, predictability, reverberation, spatial movement, foreground–background ratio, semantic content and user control. If the system records those changes, it can support controlled contrasts unavailable in a static category label.
Design also creates new confounds. A complex composition may change many variables simultaneously. Generative rules can make every exposure different. Adaptive behaviour can couple the stimulus to the participant. High production quality can increase expectancy. Researchers should decide whether the intended object is:
- a fixed waveform;
- a fixed recipe that deterministically generates a waveform;
- a distribution of possible outputs under a rule set; or
- an interactive policy whose output depends on the participant.
These are different interventions. A generative system can be reproducible at the policy level without replaying an identical signal, but then the research question concerns a class of exposures and requires distributional checks. Calling all four “the same sound environment” hides the design.
6.4 A design-space hypothesis, not a benefit claim#
The literature supports a restrained hypothesis: acoustic properties and contextual fit can influence appraisal, attention and short-term regulation-related outcomes; deliberately designed systems can manipulate those properties with greater control than loosely named sound categories. It does not yet identify a universal restorative grammar. Any proposed “calming” feature—slow change, low event density, harmonic stability, reduced sharpness—must be tested against level, familiarity, preference, task and population.
7. The reproducibility problem#
7.1 “Same audio” is underspecified#
Reproducibility in acoustic research is a chain of conditional claims. A file can be copied exactly while playback differs by decoder, resampling, operating-system processing, device response, gain, channel mapping and room. A browser can execute the same graph while its AudioContext adopts a different hardware sample rate. Headphones reduce room effects but introduce model response, fit and level uncertainty. Loudspeakers permit spatial fields but make room geometry and position consequential.
Online auditory research demonstrates both the opportunity and the limitation. Woods and colleagues developed a headphone-screening task to improve control in unsupervised listening [42]. Wycisk and colleagues studied listening-device effects in 40 participants and found high test–retest reliability for level adjustment and somewhat lower reliability for stereo-related measures, while device conditions still mattered [43]. Seow and Hauser found that online ratings of affective sounds could show good reliability and broad comparability to laboratory data [44]. These studies support feasible remote perception research; they do not make remote playback calibrated.
7.2 Signal repetition and field repetition are different#
A room field depends on source directivity, loudspeaker placement, reflections, absorption, listener location and background sound. Spatial reproduction adds format, decoding and array geometry. Tarlao, Steele and Guastavino compared in-situ soundscape evaluations with first-order Ambisonic laboratory reproduction and different questionnaire modes. They found metric invariance for computer and paper administration and recovered important soundscape dimensions, while contextual factors such as time, day and location affected ratings [45]. The study supports laboratory reproduction for some evaluation purposes and simultaneously shows that representation depends on context and stimulus selection.
Ecological validity and technical repeatability are therefore distinct axes. A tightly controlled headphone stimulus may be repeatable but unlike the target environment. An immersive field reproduction may better approximate spatial experience while adding hardware and room dependencies. A field study has high situational realism but limited experimental control. Strong programmes use complementary designs rather than declaring one setting universally valid.
7.3 Perception and response are not renderer outputs#
Even a bit-identical waveform is an exposure opportunity, not a response guarantee. Hearing, attention, state, expectations, order effects and prior exposure vary within a person. Between-person variability adds sensitivity, age, culture, clinical status and preference. Repeated response should be measured with reliability statistics and analysed for both group and individual patterns. If a system creates a highly repeatable signal but the outcome has low test–retest reliability, that is a scientific result rather than an engineering failure.
This yields a six-level reproducibility ladder:
| Level | Object repeated | Minimum evidence |
|---|---|---|
| R0 — Description | Intended experience in words | Stable narrative record |
| R1 — Composition specification | Sources, parameters, timing and authored relations | Versioned recipe, schema, round-trip tests |
| R2 — Digital rendering | Output under a named renderer | Fixed dependencies, seeds, rate and PCM conformance |
| R3 — Playback exposure | Signal at the listener interface | Device, channel path, level and calibration record |
| R4 — Room field | Spatiotemporal field in a defined space | Room/IR, source and listener geometry, calibration |
| R5 — Perceived condition | Appraisal or perceptual organization | Standardized context and repeated perceptual measures |
| R6 — Human response | Behavioural, physiological or clinical outcome | Controlled design, valid measures and replication |
The ladder is cumulative only in documentation, not in guaranteed effect. Higher levels require new evidence. R2 does not entail R3; R4 does not entail R5; R5 does not entail R6.
7.4 Standards, open systems and conformance#
Standards reduce ambiguity but do not prove adoption. The ISO 12913 series provides a common soundscape framework, collection methods and analysis guidance [5–8]. RG-MBI improves intervention reporting [38,39]. The Web Audio specification defines processing interfaces, but underlying implementation and hardware still require empirical conformance testing [46]. The Auditory Modeling Toolbox offers a useful model for computational reproducibility: named versions, documented implementations, included data, demonstrations and experiments intended to reproduce published outputs [47].
For a designed acoustic-condition system, a credible R2 claim requires more than a project version string. It requires a content-addressed dependency manifest; fixed renderer profile and sample rate; deterministic pseudorandom sources; explicit time-zero semantics; embedded or hashed assets; and fixtures that compare rendered PCM across fresh runs and supported runtimes. Numerical tolerance may be appropriate if bit identity is not technically attainable, but the tolerance and comparison metric must be declared before testing.
7.5 The minimum acoustic-condition record#
At minimum, a study should publish:
- recipe/schema version and immutable identifier;
- source-code release, renderer profile and dependency manifest;
- asset hashes, pseudorandom seeds, sample rate and channel layout;
- rendered-file hashes when fixed renders are used;
- presentation device, operating chain, gain and calibration method;
- room, source and listener positions for loudspeaker presentation;
- session timing, background sound, activity and instructions;
- participant hearing/sensitivity variables justified by the question;
- primary and secondary outcomes, exclusions and analysis plan;
- deviations from the intended condition.
Without this chain, replication failure cannot be localized. With it, a later team can ask whether divergence occurred in rendering, playback, context, perception or response.
8. IFAH as a specified acoustic-condition instrument#
8.1 What the system can reasonably be#
IFAH enters the argument here, after the evidence and methodological problem have been established. It is not presented as a therapy, a state-reproduction device or proof that designed sound regulates people. Its candidate contribution is narrower:
IFAH is a system for composing, preserving and replaying specified acoustic conditions. It is designed to improve consistency and documentation of the stimulus while leaving the human response open to observation.
In the current implementation, a sealed .soma Edition is represented internally by a “moment” that stores an authored recipe and engine identity. The system can seal and reopen recipes, reject incompatible engine hashes and restore a broad parameter set. A separate offline path uses fixed-rate, seeded construction for a synthesis subset. These are meaningful R1 and partial-R2 capabilities. They could support controlled acoustic stimuli, versioned study fixtures and transparent reporting if their limits are made explicit.
8.2 Exact-release audit#
A read-only audit examined two exact repository releases on 14 August 2026:
- Player: commit
7dd8a253aa5ca813b67e36f9ecc2e981e5d22a64 - Studio: commit
bcbcdd6c3d1559788f1828e11fe56e7c6986af03
Both identify engine SM-01, engine version 2.4.0 and recipe version 2. The engine contract, moment implementation, tests, offline renderer, chamber and material modules are byte-identical across the inspected releases; Player-specific shell and later application modules differ. Detailed blob identifiers and permalinks appear in Appendix C.
The audit confirms four strengths:
- sealed objects carry engine name, version and hash;
- strict opening rejects missing, foreign, future and forged legacy identities;
- recipe round-trip and compatibility behaviour are regression-tested;
- offline synthesis fixes 44.1 kHz and uses seeded deterministic construction for its internal noise and synthesized impulse.
It also confirms material limits:
- the engine hash is FNV-1a over engine version plus suite material; it is not a cryptographic content hash of DSP source, wavetables, chamber impulse responses or audio assets, despite broader contract wording;
- live breath noise uses
Math.random(); - live
AudioContextuses the default device sample rate; - live and offline graphs differ, including waveshaper oversampling (
4xlive,2xoffline) and chamber behaviour; - offline
renderMixis explicitly synthesis-only while real-time tape can include ambient material; - external assets are loaded by path and decoded at runtime, without content addressing;
- Player restart reapplies the recipe to a persistent graph and does not reset all oscillator, modulation, noise or loop phases to signal time zero;
- playback level, device response, room field and listener position are not sealed guarantees.
The defensible present verdict is R2-partial: version-gated parameter reconstruction plus a deterministic offline synthesis subset. Bit-identical live output across runs, devices or platforms is not established.
8.3 Conformance status#
A conformance harness was prepared against the exact Studio renderer. It defines a fixed 16-second, two-channel, synthesis-only fixture at 44.1 kHz; launches two fresh browser processes; renders from the exact release modules; hashes each Float32 PCM channel with SHA-256; and records equality, peak and RMS. This is an appropriate first within-runtime R2 test because it begins from fresh renderer processes rather than reusing a graph.
The test did not produce a render result in this audit environment. The installed automation package had no browser executable, so execution stopped before the fixture ran. The result must therefore be recorded as pending, not pass or fail. No statement about PCM identity is inferred from source inspection alone. The next step is to test the declared deployment/runtime matrix, attempting exact equality first and preregistering any tolerance-based criterion before use.
8.4 Route to a defensible research renderer#
A defensible progression is:
- replace broad identity language with the present R2-partial claim;
- generate a SHA-256 manifest covering renderer and DSP sources, suite data, wavetables, impulse responses, audio bytes and clamp semantics;
- derive every stochastic source from a sealed seed;
- declare renderer profile, sample rate, channel layout and manifest root in the seal;
- define restart as a deterministic reconstruction from signal time zero;
- unify live and offline graphs or name them as separate render products;
- embed or content-address recorded assets and exclude volatile live input from deterministic claims;
- run and publish conformance fixtures across the supported runtime matrix;
- add a study exposure record for device, gain, calibrated level, room and position.
Completing these steps would strengthen R2 and permit documented R3 protocols. It would still not establish a repeated perceived condition or human response.
9. Research agenda#
9.1 Programme logic#
The research programme should progress from instrument validation to perceptual characterization and only then to regulation or clinical hypotheses.
The sequence prevents a common product-shaped design in which a desired health claim determines the stimulus and measures. It also creates useful stopping points. If the renderer is not conformant, instrument work continues. If appraisal is unstable, the stimulus may still be an artistic object but not a standardized perceptual condition. If acute effects do not replicate, no clinical escalation is justified.
9.2 Study 0: render and exposure validation#
The first study is technical. Predeclare fixtures representing pure synthesis, recorded assets, chamber processing and any generative randomness. Test same-run, fresh-run, same-runtime and cross-runtime PCM identity. Publish hashes, difference signals, peak absolute error and spectral error. Separately measure the playback chain with a calibrated microphone or coupler, recording level, transfer function, channel mapping and background sound. The primary outcomes are conformance metrics, not participant responses.
9.3 Study 1: perceptual map and reliability#
Recruit a heterogeneous non-clinical sample with documented hearing status, noise sensitivity and relevant musical familiarity. Compare a small number of preregistered acoustic conditions with matched controls. Match or explicitly manipulate level. Collect pleasantness, eventfulness, appropriateness and perceived control, plus a limited set of semantic descriptors. Repeat conditions across sessions to estimate within-person and group test–retest reliability. The primary question is whether the specified condition produces a distinguishable and repeatable perceptual profile—not whether it heals.
9.4 Study 2: acute attention or stress recovery#
Choose one mechanism. For attention, compare conditions on a task selected to distinguish changing-state interference from deviant capture. For stress recovery, use a standardized mild stressor, preregister one subjective and one physiological primary outcome, and include quiet plus an acoustically and semantically informative comparator. Counterbalance order, blind analysts where possible and measure expectancy. Report null and divergent measures rather than compressing them into “regulation.”
9.5 Study 3: context and agency#
Test fixed, participant-selected and adaptive variants while holding the available design space constant. Manipulate or record visual congruence, activity and perceived control. Model participants and sessions hierarchically. The target is moderation: who responds, under which context, and whether agency changes appraisal, adherence or outcome. This study should be powered for interactions only after plausible effect sizes emerge from earlier work.
9.6 Study 4: repeated exposure#
If acute results are credible, examine several weeks of exposure with adherence logs and ecological momentary measures. Distinguish habituation, learning and selection. Record concurrent listening and adverse effects such as irritation, headache, tinnitus aggravation or sleep disruption. A repeated-exposure study should not be called clinical unless it enrolls a defined clinical population and uses validated clinical outcomes under appropriate oversight.
9.7 Open-science commitments#
For each study, preregister the primary question, stimulus identities, comparison conditions, exclusions, outcome hierarchy and analysis. Deposit source releases, manifests, recipes, rendered fixtures, calibration records, anonymized data and analysis code where ethical and legally possible. Report researcher involvement in system development as a conflict of interest. Seek independent replication by a laboratory that did not author the acoustic conditions or software.
10. Limits and open question#
This review is integrative and targeted. It did not use a database-complete systematic search, dual screening or formal risk-of-bias instrument across every evidence family. Its role is to establish construct boundaries, compare evidence maturity and design a testable methodological frame. The literature is also structurally heterogeneous: epidemiology, soundscape appraisal, auditory cognition, music reward, intervention trials and software conformance answer different questions. Quantitative pooling across them would be inappropriate.
The paper does not claim that sound is neutral. Some environmental exposures are harmful, auditory cues can capture attention, music can be deeply rewarding, and selected interventions can improve selected outcomes. The limits concern transfer. Harm from noise does not prove benefit from designed sound. Reward does not establish addiction. Preference does not establish regulation. Acute regulation does not establish clinical efficacy. A deterministic signal does not establish a repeated room field, perception or response.
IFAH’s present value is accordingly methodological. It can make an authored acoustic recipe more inspectable and replayable than an undocumented live session. The exact-release audit also shows why this value must not be inflated: identity metadata are not yet a full dependency manifest; the deterministic offline path is incomplete; live time-zero semantics and cross-runtime PCM conformance remain open; and exposure calibration lies outside the seal.
The central open question is not whether sound “works.” It is:
Which specified acoustic properties, delivered through which calibrated playback conditions, are associated with which perceptual or regulatory outcomes, for whom, during which activity and over what time—and how much of that relationship survives independent replication?
The value of that question is methodological. It turns a designed acoustic environment from a persuasive experience into a research object while preserving the most important uncertainty: even when a condition is carefully composed, preserved and replayed, the human response remains something to observe.
References#
- World Health Organization Regional Office for Europe. (2018). Environmental noise guidelines for the European Region. https://www.who.int/europe/publications/i/item/9789289053563
- World Health Organization. (2024). Guidance on environmental noise. https://www.who.int/tools/compendium-on-health-and-environment/environmental-noise
- Basner, M., Babisch, W., Davis, A., Brink, M., Clark, C., Janssen, S., & Stansfeld, S. (2014). Auditory and non-auditory effects of noise on health. The Lancet, 383(9925), 1325–1332. https://doi.org/10.1016/S0140-6736(13)61613-X
- European Environment Agency. (2025). Environmental noise in Europe 2025. https://www.eea.europa.eu/en/analysis/publications/environmental-noise-in-europe-2025
- International Organization for Standardization. (2014). ISO 12913-1:2014 Acoustics—Soundscape—Part 1: Definition and conceptual framework. https://www.iso.org/standard/52161.html
- International Organization for Standardization. (2018). ISO/TS 12913-2:2018 Acoustics—Soundscape—Part 2: Data collection and reporting requirements. https://www.iso.org/standard/75267.html
- International Organization for Standardization. (2025). ISO/TS 12913-3:2025 Acoustics—Soundscape—Part 3: Data analysis. https://www.iso.org/standard/86955.html
- Aletta, F., & Torresin, S. (2023). Adoption of ISO/TS 12913-2:2018 protocols for data collection from individuals in soundscape studies. Current Pollution Reports. https://doi.org/10.1007/s40726-023-00283-6
- Zhang, R., Ma, H., Wang, C., Zhang, Y., & Kang, J. (2025). Soundscape and its context: A framework based on a systematic review. The Journal of the Acoustical Society of America, 157(6), 4417–4436. https://doi.org/10.1121/10.0036882
- Shamma, S. A., Elhilali, M., & Micheyl, C. (2011). Temporal coherence and attention in auditory scene analysis. Trends in Neurosciences, 34(3), 114–123. https://doi.org/10.1016/j.tins.2010.11.002
- Bell, R., Röer, J. P., Marsh, J. E., Storch, D., & Buchner, A. (2017). The effect of cognitive control on different types of auditory distraction. Experimental Psychology. https://doi.org/10.1027/1618-3169/a000372
- Upshaw, J. D., Stevens, C. E., Ganis, G., & Zabelina, D. L. (2022). The hidden cost of a smartphone: The effects of smartphone notifications on cognitive control from a behavioral and electrophysiological perspective. PLOS ONE, 17, e0277220. https://doi.org/10.1371/journal.pone.0277220
- Salimpoor, V. N., Benovoy, M., Larcher, K., Dagher, A., & Zatorre, R. J. (2011). Anatomically distinct dopamine release during anticipation and experience of peak emotion to music. Nature Neuroscience, 14, 257–262. https://doi.org/10.1038/nn.2726
- Ferreri, L., Mas-Herrero, E., Zatorre, R. J., Ripollés, P., Gomez-Andres, A., Alicart, H., Olivé, G., Marco-Pallarés, J., Antonijoan, R. M., Valle, M., Riba, J., & Rodriguez-Fornells, A. (2019). Dopamine modulates the reward experiences elicited by music. Proceedings of the National Academy of Sciences, 116(9), 3793–3798. https://doi.org/10.1073/pnas.1811878116
- Saarikallio, S., Gold, C., & McFerran, K. (2015). Development and validation of the Healthy–Unhealthy Music Scale. Child and Adolescent Mental Health, 20(4), 210–217. https://doi.org/10.1111/camh.12109
- Gustavson, D. E., Coleman, P. L., Iversen, J. R., Maes, H. H., Gordon, R. L., & Lense, M. D. (2021). Mental health and music engagement: Review, framework, and guidelines for future studies. Translational Psychiatry, 11, 370. https://pmc.ncbi.nlm.nih.gov/articles/PMC8257764/
- Hole, J., Hirsch, M., Ball, E., & Meads, C. (2015). Music as an aid for postoperative recovery in adults: A systematic review and meta-analysis. The Lancet, 386(10004), 1659–1671. https://doi.org/10.1016/S0140-6736(15)60169-6
- de Witte, M., Spruit, A., van Hooren, S., Moonen, X., & Stams, G.-J. (2020). Effects of music interventions on stress-related outcomes: A systematic review and two meta-analyses. Health Psychology Review, 14(2), 294–324. https://doi.org/10.1080/17437199.2019.1627897
- de Witte, M., Pinho, A. da S., Stams, G.-J., Moonen, X., Bos, A. E. R., & van Hooren, S. (2022). Music therapy for stress reduction: A systematic review and meta-analysis. Health Psychology Review, 16(1), 134–159. https://doi.org/10.1080/17437199.2020.1846580
- Zhao, N., Nystrup Lund, H., & Vibe Jespersen, K. (2024). A systematic review and meta-analysis of music interventions to improve sleep in adults with mental health problems. European Psychiatry, 67(1), e62. https://doi.org/10.1192/j.eurpsy.2024.1773
- Tang, Y. W., Teoh, S. L., Yeo, J. H. H., Ngim, C. F., Lai, N. M., Durrant, S. J., & Lee, S. W. H. (2022). Music-based intervention for improving sleep quality of adults without sleep disorder: A systematic review and meta-analysis. Behavioral Sleep Medicine, 20(2), 241–259. https://doi.org/10.1080/15402002.2021.1915787
- Clark, C., & Paunovic, K. (2018). WHO environmental noise guidelines for the European Region: A systematic review on environmental noise and cognition. International Journal of Environmental Research and Public Health, 15(2), 285. https://doi.org/10.3390/ijerph15020285
- van Kempen, E., Casas, M., Pershagen, G., & Foraster, M. (2018). WHO environmental noise guidelines for the European Region: A systematic review on environmental noise and cardiovascular and metabolic effects. International Journal of Environmental Research and Public Health, 15(2), 379. https://doi.org/10.3390/ijerph15020379
- Mitchell, A., Oberman, T., Aletta, F., Erfanian, M., Kachlicka, M., Lionello, M., & Kang, J. (2022). The Soundscape Indices protocol: A method for assessing acoustic environments in urban parks. https://pmc.ncbi.nlm.nih.gov/articles/PMC9690752/
- Tarlao, C., Steele, D., & Guastavino, C. (2019). Investigating factors influencing soundscape evaluations across multiple urban spaces in Montreal. INTER-NOISE 2019. https://www.sea-acustica.es/INTERNOISE_2019/Fchrs/Proceedings/2032.pdf
- Torresin, S., Albatici, R., Aletta, F., Babich, F., & Kang, J. (2019). Assessment methods and factors determining positive indoor soundscapes in residential buildings: A systematic review. Sustainability, 11(19), 5290. https://doi.org/10.3390/su11195290
- Torresin, S., Albatici, R., Aletta, F., Babich, F., Oberman, T., Siboni, S., & Kang, J. (2020). Indoor soundscape assessment: A principal components model of acoustic perception in residential buildings. Building and Environment, 182, 107152. https://doi.org/10.1016/j.buildenv.2020.107152
- Aletta, F., Oberman, T., & Kang, J. (2018). Associations between positive health-related effects and soundscapes perceptual constructs: A systematic review. International Journal of Environmental Research and Public Health, 15(11), 2392. https://doi.org/10.3390/ijerph15112392
- Buxton, R. T., Pearson, A. L., Allou, C., Fristrup, K., & Wittemyer, G. (2021). A synthesis of health benefits of natural sounds and their distribution in national parks. Proceedings of the National Academy of Sciences, 118(14), e2013097118. https://doi.org/10.1073/pnas.2013097118
- Alvarsson, J. J., Wiens, S., & Nilsson, M. E. (2010). Stress recovery during exposure to nature sound and environmental noise. International Journal of Environmental Research and Public Health, 7(3), 1036–1046. https://doi.org/10.3390/ijerph7031036
- Fan, L., & Baharum, M. R. (2024). The effect of exposure to natural sounds on stress reduction: A systematic review and meta-analysis. Stress, 27(1), 2402519. https://doi.org/10.1080/10253890.2024.2402519
- World Health Organization. (n.d.). Addictive behaviour. https://www.who.int/health-topics/addictive-behaviour
- World Health Organization. (n.d.). Gaming disorder: Frequently asked questions. https://www.who.int/standards/classifications/frequently-asked-questions/gaming-disorder
- Schmuziger, N., Patscheke, J., Stieglitz, R.-D., & Probst, R. (2012). Is there addiction to loud music? Findings in a group of non-professional pop/rock musicians. Audiology Research, 2(1), e11. https://doi.org/10.4081/audiores.2012.e11
- World Health Organization & International Telecommunication Union. (2019). Safe listening devices and systems: A WHO–ITU standard. https://www.who.int/publications/i/item/9789241515276
- American Music Therapy Association. (n.d.). About music therapy. https://www.musictherapy.org/about/musictherapy/
- World Federation of Music Therapy. (2026). A new definition of music therapy—July 2026. https://www.wfmt.info/post/a-new-definition-of-music-therapy-july-2026
- Robb, S. L., Springs, S., Edwards, E., Golden, T. L., Johnson, J. K., Burns, D. S., Belgrave, M., Bradt, J., Gold, C., Habibi, A., Iversen, J. R., Lense, M., MacLean, J. A., & Perkins, S. M. (2025). Reporting Guidelines for Music-based Interventions: An update and validation study. Frontiers in Psychology, 16, 1551920. https://doi.org/10.3389/fpsyg.2025.1551920
- Robb, S. L., Story, K. M., Harman, E., Burns, D. S., Bradt, J., Edwards, E., & others. (2025). Reporting Guidelines for Music-based Interventions checklist: Explanation and elaboration guide. Frontiers in Psychology, 16, 1552659. https://doi.org/10.3389/fpsyg.2025.1552659
- Cochrane. (2025). Does music-based therapy help people with dementia? https://www.cochrane.org/evidence/CD003477_does-music-based-therapy-help-people-dementia
- Cochrane. (2021). Can music interventions benefit people with cancer? https://www.cochrane.org/evidence/CD006911_can-music-interventions-benefit-people-cancer
- Woods, K. J. P., Siegel, M. H., Traer, J., & McDermott, J. H. (2017). Headphone screening to facilitate web-based auditory experiments. Attention, Perception, & Psychophysics, 79, 2064–2072. https://pmc.ncbi.nlm.nih.gov/articles/PMC5693749/
- Wycisk, Y., Kopiez, R., Bergner, J., Sander, K., Preihs, S., Peissig, J., & Platz, F. (2023). The Headphone and Loudspeaker Test—Part I: Suggestions for controlling characteristics of playback devices in internet experiments. Behavior Research Methods, 55, 1094–1107. https://doi.org/10.3758/s13428-022-01859-8
- Seow, T. X. F., & Hauser, T. U. (2022). Reliability and validity of online affective auditory experiments. Behavior Research Methods. https://doi.org/10.3758/s13428-021-01643-0
- Tarlao, C., Steele, D., & Guastavino, C. (2022). Assessing the ecological validity of soundscape reproduction in different laboratory settings. PLOS ONE, 17(6), e0270401. https://doi.org/10.1371/journal.pone.0270401
- World Wide Web Consortium. (2025). Web Audio API 1.1. https://www.w3.org/TR/webaudio-1.1/
- Majdak, P., Hollomey, C., & Baumgartner, R. (2022). AMT 1.x: A toolbox for reproducible research in auditory modeling. Acta Acustica, 6, 19. https://doi.org/10.1051/aacus/2022011
TECHNICAL APPENDICES · A–G OPEN
Appendix A — Review and synthesis methods#
A1. Design#
This work is an integrative review and methodological perspective. It was designed to connect evidence families whose questions and methods are not commensurable enough for a single pooled estimate. It is not labelled a systematic review because it did not conduct exhaustive database searches, duplicate title/abstract screening, formal study-by-study risk-of-bias scoring or PRISMA flow accounting.
A2. Governing source hierarchy#
Sources were prioritized in this order:
- current standards, international health guidance and classification sources;
- systematic reviews, meta-analyses and evidence syntheses;
- preregistered, randomized or otherwise strong controlled primary studies;
- validated measurement and reporting frameworks;
- mechanistic reviews and targeted experiments;
- exact-release source code, tests and contracts for system claims.
Public summaries were used only when they came from the issuing organization and directly represented a current review or standard. Search-result snippets were not treated as evidence when the source document could be inspected.
A3. Search structure#
Targeted searches were conducted and updated through 14 August 2026 across official WHO, EEA, ISO, W3C, AMTA, WFMT, Cochrane, PubMed/PMC, journal and DOI records. Search families included:
- environmental noise + sleep/cardiovascular/cognition;
- soundscape + ISO 12913 + context + indoor;
- natural sound + stress/recovery + meta-analysis;
- auditory attention + auditory distraction + notification cues;
- music reward + dopamine + addiction/problematic use;
- music-based intervention + stress/sleep/dementia/cancer;
- intervention reporting + reproducibility;
- audio playback + online experiments + ecological validity;
- Web Audio + deterministic render + conformance.
The direct addiction search applied a construct threshold before interpreting adjacent evidence: impaired control, priority, persistence despite harm and functional impairment. Studies of reward, loudness risk, notification capture, music use in substance-use care and maladaptive regulation were recorded as adjacent evidence rather than direct confirmation.
A4. Extraction fields#
For each central source, extraction focused on:
- study/review type and population;
- exposure or intervention definition;
- comparator;
- outcome family;
- central result;
- heterogeneity, certainty or design limits;
- the strongest claim the source permits;
- the tempting claim it does not permit.
The synthesis deliberately preserved disagreement between outcomes. A significant physiological measure did not overwrite a null subjective result, and a preferred condition did not overwrite absent clinical evidence.
A5. Evidence grades used in the claim ledger#
| Grade | Meaning in this paper |
|---|---|
| High | Supported by current authoritative guidance and/or multiple converging high-level syntheses for a well-specified claim. |
| Moderate | Supported by a systematic review or several controlled studies, with material heterogeneity or scope limits. |
| Preliminary | Supported by small, selected, exploratory or narrow primary evidence; suitable for a hypothesis, not a broad conclusion. |
| Boundary | Evidence primarily constrains wording or prevents category error rather than estimating an effect. |
| Technical pass/partial/fail/pending | Direct source-code or conformance status; not a grade of human-outcome evidence. |
A6. Publication-integrity rules#
- “Associated with” is not rewritten as “causes” unless the design supports causal inference.
- Mechanism evidence is not used as efficacy evidence.
- Self-report, behavioural, physiological and clinical outcomes remain labelled.
- Evidence from music therapy, music medicine, passive listening, natural sound and designed conditions is not merged without stating the transfer.
- No unpublished IFAH human result is described or implied.
- Exact-release technical claims cite immutable commits and file blobs.
- A blocked conformance run is reported as pending.
Appendix B — Updated claim ledger#
| ID | Claim | Evidence grade | Permitted formulation | Principal limit |
|---|---|---|---|---|
| C01 | Some environmental-noise exposures adversely affect health. | High | Transport noise is associated with annoyance, sleep disturbance, cardiovascular risk and selected cognitive outcomes. | Evidence varies by source, exposure model and endpoint. |
| C02 | Environmental noise has one universal effect curve. | Unsupported | No universal claim permitted. | Source, timing, level, outcome and population differ. |
| C03 | Soundscape is the positive inverse of noise. | Unsupported | Soundscape adds perception, experience and context to the object of study. | It is not inherently positive or health-promoting. |
| C04 | Context moderates soundscape appraisal. | Moderate | Place, activity, multisensory cues, social setting and person factors can influence evaluation. | Indicators are inconsistent and causal directions can be reciprocal. |
| C05 | Natural sounds can support short-term recovery-related outcomes. | Moderate | Selected natural-sound conditions outperform selected noise or quiet comparators on some outcomes. | Heterogeneous stimuli, sparse subgroups and outcome disagreement. |
| C06 | Natural sound is universally regulating. | Unsupported | No universal claim permitted. | Category breadth, semantics, level, familiarity and context. |
| C07 | Temporal organization influences auditory grouping and attention. | Moderate | Temporal coherence, changing state and deviants provide testable accounts of different attentional effects. | Models are task- and paradigm-dependent. |
| C08 | Notification sounds can capture attention. | Moderate, narrow | Learned notification cues can disrupt performance in defined tasks. | Does not generalize to continuous sound or diagnose addiction. |
| C09 | Music engages reward-related dopaminergic processes. | Moderate | Dopamine causally contributes to aspects of musical pleasure and motivation. | Selected stimuli/listeners; mechanism is not diagnosis or efficacy. |
| C10 | Reward-system engagement establishes addiction. | Unsupported / boundary | Reward and addiction must be separated. | Addiction requires impaired control, priority, persistence despite harm and impairment. |
| C11 | A general addiction to sound is currently established. | Not supported | Addiction remains a public question, not a scientific premise. | Direct evidence is sparse, exploratory and constructually indirect. |
| C12 | Music use can be adaptive or maladaptive. | Moderate | Music engagement can relate to wellbeing, rumination, stress or depressive vulnerability. | Much evidence is correlational; HUMS is not an addiction instrument. |
| C13 | Hazardous sound exposure and addictive behaviour are the same. | Unsupported / boundary | Treat level–duration dose and impaired-control syndromes as separate risk dimensions. | They may co-occur but require different measures. |
| C14 | Music-based interventions can affect selected health-related outcomes. | Moderate | Some perioperative, stress, sleep, dementia and cancer outcomes show benefits in defined contexts. | Heterogeneity, bias, comparator effects and inconsistent longer-term evidence. |
| C15 | Any calming sound is a therapy. | Unsupported | “Calming” may describe appraisal or an acute outcome only. | Therapy requires a professional/intervention frame and efficacy evidence. |
| C16 | Music-therapy evidence transfers directly to fixed designed audio. | Unsupported / boundary | Transfer requires an explicit intervention-component argument. | Therapist relationship, active participation and tailoring may be causal components. |
| C17 | Deliberate design increases control over candidate acoustic variables. | High, methodological | Parameterized design can enable specified contrasts and records. | Control does not imply benefit or ecological validity. |
| C18 | The same digital file ensures the same listener exposure. | Unsupported | File identity is one element of an exposure chain. | Decoder, sample rate, gain, device, fit, room and position. |
| C19 | Repeated digital output ensures repeated perception or response. | Unsupported | Perception and response remain empirical outcomes. | Within- and between-person state and context vary. |
| C20 | IFAH presently preserves a version-gated recipe and restores authored parameters. | Technical pass / partial-strong | Sealed moments support compatibility identity and parameter reconstruction. | Completeness and external dependencies still require audit. |
| C21 | IFAH presently guarantees bit-identical live output. | Technical fail / not established | No such claim permitted. | Unseeded live noise, default sample rate, persistent graph, runtime differences. |
| C22 | IFAH has a deterministic offline synthesis subset. | Technical partial | The exact release contains a fixed-rate, seeded offline synthesis path. | It is synth-only, differs from live graph and lacks completed PCM conformance. |
| C23 | The IFAH engine hash covers all reproduction-critical bytes. | Technical fail | The hash gates declared version and suite material. | It omits DSP/source/assets/IR bytes despite broader contract wording. |
| C24 | IFAH reproduces a human state. | Unsupported and prohibited | IFAH may specify an acoustic condition; the human response is observed. | R5/R6 require empirical measurement and replication. |
Appendix C — Exact-release technical audit#
C1. Scope and reproducibility statement#
Repository: echofield/somasound
Audit mode: read-only source and test inspection plus attempted local conformance execution
Audit date: 14 August 2026
Player release: 7dd8a253aa5ca813b67e36f9ecc2e981e5d22a64
Studio release: bcbcdd6c3d1559788f1828e11fe56e7c6986af03
The two releases share the audited engine contract, moment logic, tests, renderer, chamber and material modules at identical blob hashes. Player has later application/audio-shell modules and Player-only edition/shell files. The inspected tests pin the current declared engine hash as 101c9e3e. The scientific verdict does not depend on branch names; it is pinned to the commits and blobs below.
C2. File and blob register#
| File | Player blob | Studio blob | Audit relevance |
|---|---|---|---|
contract/SOMA_CONTRACT.md |
1211a1b2c23e1603784141e4195618342f314dea |
same | Intended engine-pinning and manifest contract |
src/engine/moment.js |
0b7c04c5e5e3387c887edf42297befa7d5f0d8e0 |
same | Identity, FNV-1a hash, seal/open and refusal semantics |
tests/moment.test.mjs |
68bfe41c1dbd0c798482f3422d3c5de92ddcbfb7 |
same | Serialization and compatibility regression tests |
src/main.js |
c53bcbb1a4f694fa06ebb19d673e6f16fd56501322 |
b46acef945c7af6b06d98af1b2673884cf5a25b92 |
Capture/application and offline-render scope |
src/engine/audio.js |
3a4de2df694ff8d3da8387a1d0ddef664b5a5e5c |
ba3d6e56eb515a4ea5201777adbc94a38dbeac91 |
Live graph, randomness, default sample rate, assets, 4× oversampling |
src/engine/render.js |
fa34d3c11899333a76839fd59fa72f59dd664f99 |
same | Offline fixed-rate deterministic subset; 2× oversampling |
src/engine/chamber.js |
e14a2d19cf4108dcf059cf52fff8c333220621b0 |
same | Synthesized and recorded impulse responses |
src/engine/material.js |
edca73fe8e68a8cfc52b1b22f5766e84177529d5 |
same | External asset paths and material definitions |
src/player/edition.js |
56abb568062f0b1e1b95b47ee1e7a6d95ca4e7d5 |
absent | Strict opening and edition metadata |
src/player/shell.js |
098fd3af92fcea38b2b6f4d79633bce66746b5ff4 |
absent | Open/restart and recipe reapplication |
Additional exact engine dependencies recovered for the conformance fixture were src/engine/suites.js (0291cea42a164c4b45b828a34643a1d49a3b8975) and src/engine/arrange.js (d1bf452faa71c8bdcc057118503cb7e11d57a416).
C3. Audit verdict by claim unit#
| Claim unit | Status | Direct finding |
|---|---|---|
| Recipe schema/version | Pass | Recipe version 2 is sealed and opened. |
| Engine identity | Pass | SM-01, ENGINE_VERSION = '2.4.0', current hash pinned. |
| Strict mismatch refusal | Pass | Missing, foreign, future and forged legacy identities are refused. |
| Parameter round trip | Partial–strong | Broad authored state captured/applied; completeness against all render-affecting state not formally proven. |
| Reproduction-critical manifest | Fail against contract wording | Actual hash covers version + suite material, not all named reproduction-critical bytes. |
| Offline deterministic construction | Partial | Fixed 44.1 kHz and seeded internal construction; no completed fresh-process PCM result. |
| Offline/live graph equivalence | Fail | Oversampling and chamber/material paths differ. |
| Complete deterministic field | Fail | Offline render is synthesis-only; live/tape can include ambient material. |
| Live source determinism | Fail | Live breath noise calls Math.random(). |
| Live sample-rate identity | Fail | Default AudioContext() adopts runtime/device conditions. |
| Time-zero restart | Fail / unverified | Reapplication occurs on a persistent graph; source and modulation phases are not rebuilt. |
| Content-addressed assets | Fail | Assets are path-addressed and runtime-decoded. |
| Playback/room calibration | Absent | Not sealed into the inspected research object. |
| Perception/response repeatability | Not established | No code property can guarantee R5/R6. |
C4. Conformance fixture and execution record#
Fixture: 16 seconds, stereo, 44,100 Hz, synthesis-only, no external suite assets, synthesized chamber path, fixed recipe values.
Isolation: two fresh browser processes.
Output checks: SHA-256 of each Float32 PCM channel, exact channel equality, peak and RMS; difference metrics reserved for a predeclared tolerance path.
Runtime target: exact Studio release modules listed above.
Execution result: pending.
The automation package was present, but its expected Chromium executable was absent. The harness failed at browser launch, before renderer import or PCM generation. This is an environment failure, not an audio result. No pass/fail judgment about cross-run identity is made.
C5. Acceptance criteria for the next audit#
- Two fresh processes on the same pinned runtime produce bit-identical channel arrays for deterministic fixtures.
- Fixtures containing every supported deterministic source and effect are added incrementally.
- Runtime version and binary hash are recorded.
- Chromium, Firefox and WebKit are tested only if they are supported products; otherwise, test the declared deployment runtime matrix.
- If cross-runtime exact identity fails, inspect difference signals before selecting a scientifically justified numerical tolerance.
- Live and offline profiles are either made graph-equivalent or declared as distinct, separately conformed renderers.
- R3 calibration tests begin only after the R2 profile is named and stable.
Appendix D — Inference boundaries#
| Tempting inference | Why it fails | Publication-safe replacement |
|---|---|---|
| Noise harms, therefore designed sound heals. | Harm and benefit are not mathematical inverses; exposure, comparison and endpoints differ. | Noise evidence motivates careful study of acoustic conditions; benefit requires direct evidence. |
| Pleasantness is wellbeing. | Pleasantness is an appraisal, not a health outcome. | Report pleasantness as pleasantness and test wellbeing separately. |
| A lower heart rate means regulation. | Heart rate is affected by movement, respiration, attention and baseline. | Interpret a preregistered multi-measure pattern within the design. |
| A physiological difference is clinically meaningful. | Statistical and clinical importance are different. | Use validated clinical outcomes and meaningful-effect thresholds. |
| Dopamine means addiction. | Dopamine supports ordinary reward, learning and motivation. | Reward mechanisms justify reward hypotheses, not diagnoses. |
| Frequent listening is compulsive listening. | Frequency does not establish impaired control or functional harm. | Measure control, priority, persistence despite harm and impairment. |
| Loud listening is addiction. | Acoustic dose and behavioural syndrome are distinct. | Measure hearing risk and problematic behaviour separately. |
| Notifications prove addiction to sound. | The sound may cue social information or device checking. | Describe cue-specific attentional capture. |
| Music therapy proves any audio therapy. | Relational, active and individualized components may matter. | Identify which intervention components plausibly transfer. |
| Natural sounds are universally good. | Natural sources vary; context and comparator shape response. | State the exact sound, level, population, context and comparator. |
| Urban sounds are universally bad. | Appropriateness and social meaning can be positive. | Treat source, activity and context as variables. |
| Designed means controlled. | Complex designs may vary many variables or generate unique outputs. | Declare the controlled object: waveform, recipe, distribution or policy. |
| Same file means same signal. | Decoding and resampling can differ. | Hash decoded PCM under a named renderer. |
| Same signal means same exposure. | Gain, device and transfer function intervene. | Publish device and calibrated exposure records. |
| Same exposure means same field. | Room, position and source directivity intervene. | Document or measure the room field. |
| Same field means same perception. | Listener and context moderate organization and appraisal. | Treat perception as measured outcome. |
| Same perception means same response. | State, history and measurement error remain. | Estimate reliability and moderation empirically. |
| A version string is a content manifest. | Unhashed dependencies can change under the same version. | Seal a cryptographic dependency-manifest root. |
| Parameter replay means waveform replay. | Randomness, phase, timing, assets and runtime remain. | Claim parameter reconstruction until PCM conformance passes. |
| Offline determinism certifies live output. | The inspected graphs differ and live timing/device paths add variables. | Name and test offline and live profiles separately. |
| Restart means time zero. | UI/runtime reset may not reset oscillator and modulation phase. | Define restart as deterministic graph reconstruction. |
| IFAH reproduces a state. | A system renders sound; state is a participant outcome. | IFAH specifies conditions and leaves response open to observation. |
Appendix E — Public website copy#
E1. Website abstract (152 words)#
Acoustic environments can disturb sleep, capture attention, carry meaning and support pleasure—but those observations do not prove that designed sound heals, regulates everyone or constitutes an addiction. IFAH Research 01 connects evidence from environmental noise, soundscape studies, auditory attention, music reward, care interventions and reproducible audio research. Its central proposal is methodological: separate the digital signal from the playback exposure, room field, perceived condition and human response. Each layer requires its own record and evidence.
The paper treats “addiction” as a question rather than a diagnosis and distinguishes reward, cue-reactivity, maladaptive music use and hazardous listening. It also audits the current IFAH/SOMA software against exact releases. The system can preserve a version-gated acoustic recipe and reconstruct authored parameters, while deterministic live output, calibrated room exposure and repeated human response remain unestablished. IFAH is therefore proposed as a candidate research instrument—not a therapy—for composing, documenting and testing specified acoustic conditions under transparent limits.
E2. Website introduction (76 words)#
What would it take to study a composed acoustic environment as seriously as any other experimental exposure? This first IFAH research paper maps what is known about noise, soundscape, attention, reward, regulation and music-based care, then focuses on the missing methodological link: reproducibility. It asks how a digital recipe becomes an exposure, a perception and a measurable response—and refuses to treat those stages as interchangeable. The result is a research agenda built to be challenged, tested and replicated.
E3. Collaborator invitation#
We invite acousticians, auditory and environmental psychologists, soundscape researchers, neuroscientists, music therapists, clinical methodologists, statisticians and research-software engineers to challenge this framework. The immediate work is not to endorse IFAH. It is to test whether the proposed construct boundaries, exposure records and reproducibility ladder survive expert scrutiny. Contributions are especially welcome on playback calibration, spatial-field measurement, hearing and sensitivity variables, preregistered perceptual and autonomic outcomes, ethical repeated-exposure design, and independent renderer conformance. Collaborators should be able to disagree with the premise, inspect the exact stimulus pipeline and publish null or adverse results.
Appendix F — Unresolved questions#
Concept and theory#
- Is acoustic condition the most useful term, or should the field distinguish composed condition, rendered condition and presented condition more explicitly?
- Which acoustic variables have the strongest theory-backed relationship with pleasantness, eventfulness, attentional demand or acute recovery?
- When is a generative recipe the research object, and when must every realized waveform be retained?
- How should perceived control be modelled: moderator, intervention component, outcome or all three in distinct analyses?
- Can appropriateness provide a more context-sensitive target than pleasantness for designed environments?
Measurement#
- Which minimum hearing, tinnitus, hyperacusis and noise-sensitivity screen is scientifically justified without overburdening participants?
- Which physiological endpoint is most defensible for the first acute recovery study, and what respiration/movement controls are required?
- What test–retest interval best separates instrument reliability from habituation and state change?
- Which manipulation checks are necessary to show that conditions differ as intended without revealing the hypothesis?
- How should adverse acoustic responses be defined and monitored?
Engineering and acoustics#
- Must the research renderer target cross-browser identity, or should it freeze a single audited runtime image?
- Is bit identity necessary for every profile, or is a perceptually and numerically justified tolerance more appropriate for some DSP paths?
- How should impulse responses, decoded audio assets and runtime binaries enter the manifest root?
- What are the minimum R3 calibration requirements for headphones versus loudspeakers?
- How large may a listener zone be before an R4 room field becomes meaningfully different?
- How should interactive participant adjustments be logged so the policy and realized exposure remain reconstructible?
Study design and governance#
- What control conditions can separate level, semantics, expectation, novelty and agency without producing an impractically large design?
- What effect size would justify moving from perceptual characterization to repeated-exposure research?
- When does a regulation study require clinical oversight even if it enrolls a non-clinical sample?
- Which independent laboratory could run the first renderer and perceptual replication without involvement in IFAH development?
- How should artistic authorship and system-development conflicts be disclosed and governed?
- At what point should the public title change from Sound, Addiction & Environment to avoid diagnostic misreading?
Appendix G — Critical review record#
Before completion, the draft was read against the strongest plausible objections. The resulting revisions are recorded here so that remaining vulnerability is visible.
| Critical objection | Disposition in this draft | Remaining vulnerability |
|---|---|---|
| “This is systematic-review language without systematic-review method.” | The article is explicitly labelled an integrative review and methodological perspective; Appendix A states absent systematic procedures. | A later formal review would need protocol registration, database-complete searches and dual screening. |
| “The title diagnoses a condition the paper cannot support.” | Addiction is restricted to one bounded interrogation; the abstract and conclusion reject a generalized diagnosis. | Public readers may still infer a claim from the title alone. |
| “The paper reverse-engineers science to sell IFAH.” | IFAH first appears in Section 8; adverse technical findings are retained; no efficacy claim or human result appears. | Author/system conflicts must be formally disclosed on submission. |
| “Noise evidence is being used as a back door to wellness.” | The inverse-noise fallacy is named repeatedly; the asymmetry in evidence maturity is explicit. | Website or promotional reuse could remove the qualification. |
| “Mechanisms decorate a weak intervention claim.” | Attention and reward mechanisms are tied to narrow constructs and explicitly separated from efficacy. | The theoretical section still requires specialist review in auditory cognition. |
| “Music therapy evidence is irrelevant to a generative sound engine.” | Intervention categories and transfer limits are stated before any synthesis. | Component-level transfer has not yet been formally modelled. |
| “Natural-sound findings are a category error.” | Comparator, source heterogeneity and null subjective-stress evidence are included. | The natural-sound literature itself has publication and selection biases not formally scored here. |
| “The reproducibility ladder is invented terminology.” | Every level has a testable object and evidence requirement; it is proposed as a framework, not a validated scale. | External users must test whether the levels are complete and usable. |
| “The code audit mistakes reading source for testing output.” | Source inspection and PCM conformance are separated; the blocked run is pending. | No empirical PCM result is available yet. |
| “An FNV hash plus version string is being oversold.” | The contract–implementation gap is a central audit finding, and a SHA-256 manifest is required. | Manual version discipline may hide historical dependency changes. |
| “The paper’s proposed studies are under-specified.” | A staged agenda identifies objects, primary families, controls and open-science commitments without inventing sample sizes. | Power, populations, measures and ethics require expert protocol co-design. |
| “Null, heterogeneous and adverse results will disappear once the project seeks impact.” | The draft requires preregistration, adverse-event recording, conflict disclosure and independent replication. | Governance and publication commitments need institutional enforcement. |
G1. Readability pass#
The final prose pass removed unneeded framework repetition, converted the core signal-to-response distinction into one compact figure, separated public and scientific titles, and kept software detail in Section 8 and Appendix C. Technical terms are defined at first use. Tables are used where exact mappings matter; narrative remains primary where evidence needs qualification.
G2. Final integrity statement#
This manuscript contains no invented human result, no claim that IFAH is a therapy, no claim that its current live renderer is bit-identical, and no claim that an acoustic condition determines a human state. Its strongest contribution is a falsifiable separation of objects and a concrete path toward better stimulus documentation. Its strongest unresolved weakness is equally concrete: the proposed renderer conformance test has not yet run.