Ambient acoustic simulation speech generation method and device, equipment and medium

By separating, transforming, and identifying the mixed speech data of the original environment, determining the acoustic labels of the target environment, and generating a simulated speech generation method, the environmental acoustic technology problems that cannot be solved in the prior art are solved. The simulated speech generation device solves the problem that the prior art cannot specifically replace the original environmental noise according to the target geographical location claimed by the user. The generated simulated speech data is more realistic in terms of acoustic distribution and geographical consistency, and it is difficult to identify camouflage traces, thus improving the acoustic camouflage effect of the speech content under the target geographical location.

CN120612917BActive Publication Date: 2025-11-18PING AN TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510844797.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-11-18
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

Existing technologies cannot replace the original environmental noise in a targeted manner based on the user's claimed target geographical location. They lack an acoustic tag recognition and environmental acoustic information matching mechanism that is collaboratively associated with voice content and geographical information, resulting in the replacement results not matching the target location and the simulation effect being unrealistic.

Method used

The system acquires and separates mixed speech data to generate original speech content and environmental acoustic information, converts it into text information, identifies environmental acoustic tags, determines target acoustic tags based on the target geographical location, extracts target acoustic information from a preset sound data set, adjusts amplitude characteristics to match the original acoustic information, and synthesizes analog speech data.

Benefits of technology

The generated simulated speech data is more realistic in terms of acoustic distribution and geographical consistency, making it difficult to detect spoofing traces and improving the acoustic spoofing effect of speech content in the target geographical location.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612917B_ABST
    Figure CN120612917B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of speech processing, which can be applied to business scenarios such as financial technology and medical health, and discloses an environmental acoustic simulation speech generation method, device, equipment and medium, comprising: obtaining and separating mixed speech data to generate original speech content and original environmental acoustic information; converting the original speech content into first text information; determining a target environmental acoustic label in combination with the original environmental acoustic label, the first text information and target geographic location information; obtaining target environmental acoustic information from a preset sound data set based on the label and adjusting its amplitude characteristics to match the original environmental acoustic information; and combining the adjusted target environmental acoustic information with the original speech content to form simulation speech data. The present application introduces target geographic location information to participate in acoustic feature determination and amplitude adjustment, so that the generated simulation speech data is more consistent in geographic semantics and acoustic performance, effectively improving the authenticity and concealment of speech environment camouflage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing technology, and in particular to a method, apparatus, device, and storage medium for generating environmental acoustic simulated speech. Background Technology

[0002] With the continuous development of audio processing and voice communication technologies, simulating or replacing the environmental acoustic features of a call has been gradually applied to various scenarios such as user privacy protection, voice data enhancement, and virtual interaction masquerading. However, existing technologies still have many shortcomings in simulating and replacing environmental acoustic information, especially in telephone call scenarios. When the user's actual location does not match their claimed location, environmental noise often contains rich geographical clues, which can be easily used to identify the true location of the call participants, and existing replacement technologies struggle to effectively conceal this.

[0003] In the fintech sector, voice interaction systems are widely used in remote identity verification, customer service, and risk control management. In these applications, the acoustic environment in which a customer calls a voice hotline is often used to help assess risk levels. For example, the system can use background noise to determine if the user is indoors or in a public place, or if it matches their declared address. However, existing audio simulation technologies often use generic template-based ambient sound replacement methods, failing to consider the actual environmental characteristics of the user's claimed location. This results in a lack of semantic and acoustic consistency between the replaced ambient noise and the actual location, making risk identification models in financial systems more prone to alerting to abnormal locations, thus impacting service processes and customer experience.

[0004] In the healthcare sector, scenarios such as remote consultations, online psychological counseling, and voice interaction monitoring place high demands on the authenticity and privacy protection of voice data. Patients often choose to conceal their geographical location for privacy reasons, and some products attempt to achieve "virtual scene camouflage" by simply replacing background noise. However, because current technology lacks the ability to model the acoustic characteristics of localized natural environments during simulation, such as regional bird calls and background noise of human voices in specific contexts, the replaced ambient sounds often sound harsh and distorted. This makes it easy for medical professionals or systems to detect traces of human manipulation, which is detrimental to protecting user information and affects the naturalness and trustworthiness of the interaction.

[0005] In the field of traditional voice spoofing technology, some systems have attempted to combine speech recognition and voiceprint forgery to "mimic" the user's environment. However, these systems typically treat environmental noise as an isolated background signal, lacking a systematic fusion analysis of semantic content, acoustic energy, temporal characteristics, and the user's claimed location. This results in a logical disconnect between the simulated environment and the user's voice. Furthermore, the replaced ambient sound lacks temporal dynamics and amplitude adaptability, making it difficult to transition naturally with the original speech content and further exposing the spoofing intent.

[0006] Therefore, existing environmental acoustic feature replacement or simulation technologies have significant shortcomings in terms of semantic consistency, preservation of regional acoustic features, dynamic energy adaptation, and speech fusion processing, making it difficult to meet the current practical needs for high-reliability spoofed speech data. Summary of the Invention

[0007] The main objective of this invention is to provide an environmental acoustic simulation speech generation method, apparatus, device, and storage medium, aiming to solve the technical problems of existing technologies that cannot specifically replace the original environmental noise according to the user's claimed target geographical location, lack an acoustic tag recognition and environmental acoustic information matching mechanism that is collaboratively associated with speech content and geographical information, resulting in the replacement result not matching the target location and the simulation effect being unrealistic.

[0008] To achieve the above objectives, the present invention provides a method for generating simulated environmental acoustic speech, comprising:

[0009] Acquire mixed speech data and perform a separation operation on the mixed speech data to generate original speech content and original environmental acoustic information;

[0010] The original speech content is converted into first text information;

[0011] The original environmental acoustic label is identified from the original environmental acoustic information, and the target environmental acoustic label is determined by the intelligent processing module based on the original environmental acoustic label, the first text information, and the target geographical location information.

[0012] Based on the target environment acoustic tag, extract the corresponding target environment acoustic information from a preset sound data set;

[0013] Adjust the amplitude characteristics of the target environmental acoustic information to match the amplitude characteristics of the original environmental acoustic information;

[0014] The original speech content is combined with the adjusted target environmental acoustic information to form analog speech data.

[0015] Furthermore, to achieve the above objectives, the present invention provides an environmental acoustic simulation speech generation device, comprising:

[0016] The sound source separation module is used to acquire mixed speech data and perform separation operations on the mixed speech data to generate original speech content and original environmental acoustic information;

[0017] The speech recognition module is used to convert the original speech content into first text information;

[0018] An environmental label recognition module is used to identify original environmental acoustic labels from the original environmental acoustic information, and to determine the target environmental acoustic label through an intelligent processing module based on the original environmental acoustic labels, the first text information, and the target geographical location information.

[0019] The acoustic retrieval and matching module is used to extract the corresponding target environment acoustic information from a preset sound data set based on the target environment acoustic tag.

[0020] An acoustic modulation module is used to adjust the amplitude characteristics of the target environmental acoustic information so that the amplitude characteristics of the target environmental acoustic information match the amplitude characteristics of the original environmental acoustic information.

[0021] The speech synthesis module is used to synthesize the original speech content and the adjusted target environmental acoustic information into analog speech data.

[0022] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and an environmental acoustic simulation speech generation program stored in the memory and executable on the processor, wherein when the environmental acoustic simulation speech generation program is executed by the processor, it implements the steps of the environmental acoustic simulation speech generation method as described above.

[0023] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an environmental acoustic simulation speech generation program, wherein when the environmental acoustic simulation speech generation program is executed by a processor, it implements the steps of the environmental acoustic simulation speech generation method described above.

[0024] Beneficial Effects: This invention relates to the field of speech processing technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for generating simulated environmental acoustic speech, comprising: acquiring and separating mixed speech data to generate original speech content and original environmental acoustic information; converting the original speech content into first text information; determining a target environmental acoustic label based on the original environmental acoustic label, the first text information, and target geographical location information; determining the target environmental acoustic information from a preset sound data set through label matching; adjusting the amplitude characteristics of the target environmental acoustic information to match the amplitude characteristics of the original environmental acoustic information; and synthesizing the original speech content and the adjusted target environmental acoustic information into simulated speech data. This invention introduces target geographical location information into the process of selecting environmental acoustic labels and generating acoustic information, enabling the synthesized environmental noise to possess acoustic characteristics consistent with the target location. Furthermore, by adjusting the amplitude characteristics, it achieves acoustic fusion between the simulated environment and the original environment, making the final generated simulated speech data more realistic in terms of acoustic distribution and geographical consistency, and more difficult to detect camouflage traces, thereby improving the acoustic camouflage effect of speech content at the target geographical location. Attached Figure Description

[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0026] Figure 1 This is a schematic diagram of an application environment for the environmental acoustic simulation speech generation method in one embodiment of the present invention;

[0027] Figure 2 This is a flowchart illustrating an embodiment of the environmental acoustic simulation speech generation method of the present invention;

[0028] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the environmental acoustic simulation speech generation device of the present invention;

[0029] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0030] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0031] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0032] The environmental acoustic simulation speech generation method provided in this invention can be applied to, for example... Figure 1In this application environment, the user terminal communicates with the server via a network. The server can acquire and separate mixed speech data from the user terminal to generate original speech content and original environmental acoustic information; convert the original speech content into first text information; determine the target environmental acoustic label based on the original environmental acoustic label, the first text information, and the target geographical location information; determine the target environmental acoustic information from a preset sound data set through label matching; adjust the amplitude characteristics of the target environmental acoustic information to match the amplitude characteristics of the original environmental acoustic information; and synthesize the original speech content and the adjusted target environmental acoustic information into simulated speech data. This invention introduces target geographical location information into the process of selecting environmental acoustic labels and generating acoustic information, enabling the synthesized environmental noise to have acoustic characteristics consistent with the target location; further, by adjusting the amplitude characteristics, it achieves acoustic fusion between the simulated environment and the original environment, making the final generated simulated speech data more realistic in terms of acoustic distribution and geographical consistency, and more difficult to detect camouflage traces, thereby improving the acoustic camouflage effect of the speech content under the target geographical location. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0033] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the environmental acoustic simulation speech generation method provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0034] like Figure 2 As shown, the environmental acoustic simulation speech generation method proposed in this invention includes the following steps:

[0035] S10, acquire mixed speech data, and perform a separation operation on the mixed speech data to generate original speech content and original environmental acoustic information;

[0036] In this embodiment, acquiring mixed speech data requires an audio acquisition module with multi-source signal acquisition capabilities, preferably a multi-channel audio input device with an array structure. Mixed speech data refers to raw waveform data that includes unstructured ambient sound signals in addition to the dominant human voice information. This mixed signal may be acquired by a user's communication device in a real-world scenario, such as a telephone microphone, remote voice terminal, or embedded communication module. Its characteristic is that human voice and background sound overlap to varying degrees in the frequency and time domains, making them indistinguishable directly.

[0037] To improve the stability and specificity of subsequent signal processing, signal enhancement is first required to eliminate uncontrollable interference. Noise suppression can be achieved through spectral subtraction, adaptive filtering, or deep learning enhancement models. The goal is to reduce unidentifiable, unstructured background noise interference in mixed signals while maintaining the integrity of the recognizable structure of ambient sound. For example, convolutional time-frequency masking networks can perform mask estimation of target signal components in the STFT domain, thereby suppressing irrelevant noise.

[0038] After noise reduction, a neural network capable of jointly separating speech and non-speech signals is needed for dual-source decomposition. This model typically comprises two parallel but co-trained branches: one branch receives the enhanced mixed speech as input and outputs speech components related to the human voice's speaker structure; the other branch constructs environmental acoustic output based on the residual sound waveform. The human voice extraction branch may be based on a gated convolutional network and a dual-channel attention mechanism to extract clear speech trajectories and reconstruct semantically stable original speech content. The environmental acoustic branch focuses on extracting spatial localization features, low-frequency background noise patterns, and periodic event sounds, outputting original environmental acoustic information with temporal resolution and geographical distribution significance. The two outputs maintain synchronization and temporal alignment in their data structure for easy subsequent joint processing.

[0039] In actual processing, the input window size and overlap ratio of the separation model can be determined according to the input signal frame length and sampling rate. For example, if the input sampling rate is 16kHz, a short-time Fourier transform with a 512-point window length and a 256-point sliding window is used as the basis for signal feature extraction, and the decoupling of human voice and background sound is completed through frame-by-frame inference.

[0040] A sound source separation network based on depthwise separable convolution, combined with a gated recursive module, can be used to perform sound source separation processing on multi-channel input signals. By introducing a channel attention mechanism, the model's sensitivity to non-stationary noise in the environment is enhanced, making it suitable for multi-source reverberation environments such as financial service counters and remote medical consultations. When deployed on edge devices, the model structure can be compressed and run in quantized inference mode, adapting to resource-constrained communication terminals. During remote calls, frame buffering and low-latency streaming mechanisms can also be introduced to ensure real-time output separation and maintain frame-level time synchronization between speech and ambient sound.

[0041] Alternatively, the signal separation process can be embedded within the audio preprocessing chip, pre-converting the mixed input signal into Mel-spectrum features, and using a separation model to perform inter-frame dimension mapping prediction, ultimately outputting two decoupled audio data streams. Another option is to introduce a real-world regional ambient sound database during the training phase for enhanced background sound matching training, making the environmental branch more likely to capture acoustic property changes in real-world geographical scenes.

[0042] Example Explanation: In healthcare scenarios, remote voice consultations may expose the user's true location due to background noise. For example, background audio elements in emergency rooms, waiting areas, or wards often exhibit distinct medical environment characteristics. By structurally acquiring and separating multi-channel voice data captured during calls, clear human voice data can be extracted without affecting the user's expression, and the accompanying environmental audio can be separated into independent background information. This provides controllable environmental noise input material for subsequently constructing a geographic acoustic atmosphere consistent with specific non-medical locations (such as home care scenarios, rehabilitation centers, or private consultation areas), thereby meeting the need for precise geographic camouflage.

[0043] In fintech business scenarios, customers sometimes wish to simulate the acoustic experience of a conversation in a branch, wealth management center, or private service area during voice identity authentication, remote account opening, or sensitive business transactions. By finely separating the original speech content from the accompanying background noise in the mixed speech signal, uncontrollable background elements in the original office environment or open area, such as crowd noise, street noise, or television broadcasts, can be removed to the greatest extent possible. While preserving a clean speech track, replaceable background sound sources are output, making the subsequent environmental construction more closely match the acoustic characteristics of the target location, forming a natural and imperceptible geolocation mapping effect.

[0044] This embodiment effectively extracts human voice information with complete semantic trajectories and environmental acoustic background containing geographical distribution characteristics by performing multi-branch structure sound source separation processing on mixed speech signals. It also maintains the time alignment characteristics of the mixed data before noise reduction and separation, thereby providing a clear and energy-balanced acoustic input foundation for subsequent semantic recognition and geographic camouflage processing.

[0045] S20, convert the original speech content into first text information;

[0046] In this embodiment, converting the original speech content into first text information first involves the identification and processing of semantically carrying regions in the speech data. The original speech content refers to the user's speech signal channel extracted after sound source separation, which may contain brief silences, non-semantic interjections, or speech rate fluctuations. To avoid these non-semantic information interfering with the speech recognition quality, the entire speech data can be processed by speech activity detection. Speech activity detection relies on multi-dimensional signal features such as short-time energy and short-time zero-crossing rate, which can effectively segment active speech regions and remove silent segments, generating multiple speech segments with temporal continuity and expressive integrity.

[0047] Each speech segment needs to be sequentially input into a speech recognition model capable of dialect identification. This model is typically trained on a deep neural network architecture and optimized using multilingual and multi-regional corpora. The intermediate text data output by the model may contain non-standard expressions introduced by different regional terms or speech recognition errors. Therefore, dialect consistency correction processing needs to be performed based on a unified language standardization strategy. This processing relies on standard dialect dictionaries, word synonym maps, and contextual semantic matching rules to map regional accents to standard expressions and eliminate potential semantic ambiguity.

[0048] After speech recognition and correction are completed, in order to preserve the correspondence between semantics and time, it is also necessary to timestamp each corrected text segment according to the temporal structure of the original speech content. This process can be generated by mapping the start and end times of each speech segment to the corresponding recognition results. All corrected texts containing time information are sorted and concatenated along the temporal dimension to form a structured language information set with a time axis index, which is finally aggregated into the first text information.

[0049] An end-to-end speech recognition network can be used to directly map speech segments to standard Chinese Pinyin, and then the speech can be transcribed using a multilingual dictionary, thereby improving the recognition accuracy for non-Mandarin speakers. Furthermore, a region-specific vocabulary can be loaded based on the user's target geographical location, dynamically adjusting the language prior distribution during language model decoding to achieve regional vocabulary adaptation. During timestamp annotation, the corresponding time periods of the speech signal and text sequence can be accurately obtained using a frame-aligned CTC (Connectionist Temporal Classification) output format, thus generating frame-segmented structured language output.

[0050] When correcting recognition errors, external knowledge graphs can be loaded to perform contextual consistency checks on identified technical terms, place names, equipment names, and other fields, thereby replacing vague and non-standard expressions. Furthermore, adaptive front-end modules, such as speech rate normalization and frequency band enhancement filtering, can be introduced into the recognition model to improve recognition performance in noisy or dialect-heavy contexts.

[0051] Example Explanation: In a healthcare scenario, when a user remotely contacts a hospital via voice and claims to be near an emergency access point, the system needs to verify whether their environment matches this claim. By performing speech activity detection and structured segmentation on the original speech content, the system can effectively extract its continuously expressed semantic content. After transing each speech segment into text, dialect correction is further performed using a standard language expression model specific to the medical region. This ensures that the final generated text information is not only semantically clear but also possesses an expression consistent with the language characteristics of the medical environment. This information can then be used to match the semantic scenarios that the target medical region may involve, helping to determine whether the user's voice possesses the language characteristics expected in that environment, and providing a semantic basis for subsequently generating replacement information that fits the acoustic characteristics of the target medical region.

[0052] In fintech scenarios, users may claim to be conducting in-person transactions at a securities brokerage via remote voice calls. To support this location spoofing, it's crucial to determine if their language exhibits characteristics of typical brokerage conversations. By segmenting and recognizing the original speech content and applying a model of commonly used brokerage terminology, the system can identify keywords such as "counter," "on-site matching," and "subscription" in the initial text. The resulting text not only forms a structured output but also aligns with the timestamps of the original speech segments. This allows for precise control of subsequent acoustic replacement time periods, ensuring that the constructed simulated speech data matches the intended scenario of the target geographical location in both semantic content and temporal structure. This enhances the completeness and credibility of the location spoofing.

[0053] This embodiment divides the speech signal into multiple complete speech segments and corrects the dialect expressions based on the geographical location context, generating semantically standardized, structurally clear, and timestamped language text content. This processing not only improves the accuracy of subsequent semantic analysis or environmental labeling inference modules but also provides linguistic layer support for the construction of geographical location acoustic camouflage. Since the text information can be used to extract geographic keywords and link them with the semantic rules of the target location, the accuracy of the processing directly determines the ability to construct a consistent simulated environment.

[0054] S30, Identify the original environmental acoustic label from the original environmental acoustic information, and determine the target environmental acoustic label through the intelligent processing module based on the original environmental acoustic label, the first text information and the target geographical location information;

[0055] In this embodiment, after receiving the raw environmental acoustic information, it needs to undergo spectral analysis processing. A short-time Fourier transform frame is constructed using a time window sliding method to generate a spectrogram data structure. This spectral structure serves as the acoustic input signal, fed into a recognition network with multi-level classification capabilities. First, coarse-grained environmental category recognition is performed, extracting broad acoustic labels such as natural sounds, mechanical noise, and speech interference. Each label is associated with multiple dedicated sub-models through an indexing mechanism. These models, after targeted training, can perform fine-grained differentiation of acoustic samples under specific broad categories. For example, bird calls are subdivided into sparrow calls and pigeon wing flapping, while mechanical sounds can be further identified as subway operation sounds and elevator motor sounds. To improve recognition accuracy, feature extraction based on Mel-frequency cepstral coefficients is performed on the acoustic segments to construct multi-dimensional acoustic vectors, which are then input into the corresponding sub-models for deep classification, outputting a set of raw environmental acoustic labels.

[0056] After obtaining the original environmental acoustic labels, the geographic semantic keywords are extracted from the structured first text information to identify potential location markers. To improve the logical consistency of label correction, target geographic location information is introduced as spatial background knowledge. A structured input vector containing environmental labels, geographic keywords, and target geographic location information is constructed and input into a pre-trained large-scale language model. The model uses a geographic knowledge graph to perform linkage judgment between acoustic labels and semantic content, identifying whether the current acoustic label has a spatial conflict with the claimed geographic location.

[0057] When the language model identifies a label in the label set that is inconsistent with the target geographical location, it automatically executes a label removal or replacement strategy. The replacement strategy is calculated and generated based on dimensions such as the historical co-occurrence probability between the target geographical location and the acoustic label, natural environmental characteristics, and urban background noise distribution. For example, the label "tidal sound" appearing in urban office areas is identified as a contradiction and replaced with "printer fan sound" as an equivalent. The final output is a set of environmental acoustic labels with a corrected structure, which serves as the semantic basis for subsequent acoustic replacement operations.

[0058] In one implementation, an acoustic segment classification module based on a hybrid structure of convolutional neural networks and attention mechanisms can be used for label recognition. The extracted acoustic feature vectors are input into a shared backbone network, and discriminative layers for different label subcategories are loaded at the tail. The original environmental acoustic labels are then output through a softmax layer. Text keyword extraction can be achieved using an LSTM-CRF network based on part-of-speech tagging, labeling the first text information as containing geographical indications such as place names, common directional words, and organization names. Target geographic location information can be provided in the system's preset configuration or by the front-end user input field, and is mapped to a unified spatial coordinate system through geocoding. The above three types of data are merged into a structured prompt template, such as a label array, a set of geographic keywords, and spatial location information forming a joint Prompt input, which is accessed through an API to a large language model server, calling the language model to complete geographic consistency inference and label correction.

[0059] Graph neural networks can also be used to model the spatial co-occurrence relationship between geographic locations and acoustic labels, constructing a location-label co-occurrence map to enhance the inference results of language models, thereby improving the accuracy and logical closure of replacements. In addition to direct replacement mechanisms, label replacement strategies can also include rules that retain labels with high confidence and replace only contradictory labels, adapting to the flexible correction needs in multi-label uncertainty scenarios.

[0060] Example Description: In a financial services scenario, a customer claims via voice call that they are in a bank lobby in a prime urban business district applying for a loan. To verify the validity of their claimed location, the system performs spectrum segmentation and label recognition on the collected background noise from the call, initially labeling it with raw environmental acoustic tags such as "wave sound" and "outdoor vendor noise." Simultaneously, the system identifies keywords such as "queue number," "filling out forms," ​​and "applying for a card" from the text within the voice call. These keywords are then associated with the claimed location, "city bank branch," and a language model determines that the labels are inconsistent with the spatial environment of a bank in a prime urban area. Given a structured input containing labels, text keywords, and geographic information, the intelligent processing module modifies the label set, replacing them with labels such as "air conditioner fan noise," "counter printing sound," and "background human voices." By completing this operation, subsequent data used for environmental synthesis will be generated based on the updated labels, thereby improving the consistency between the acoustic output content and the semantics of the target space, thus supporting the overall geographic location spoofing objective.

[0061] In healthcare scenarios, when a user claims to be waiting for a consultation in a hospital during a call, the background noise detected by the system contains tags such as "bus announcements" and "street performances," while the text recognition results contain geographic semantic clues such as "registration," "waiting," and "waiting hall." Through tag recognition and language model inference, it is determined that these tags significantly conflict with the acoustic characteristics of a hospital waiting area. The system then performs an acoustic tag replacement operation in its intelligent processing module, replacing the urban street sound elements in the initial tag set with acoustic tags more closely resembling a medical environment, such as "calling announcements," "wheelchair movement sounds," and "coughing sounds." This provides input that matches the target location characteristics for subsequent generation of environmental sound data that can be used to simulate the hospital sound field. This processing improves the consistency of the tag base used in the environment synthesis stage, thereby enhancing the acoustic realism of location masquerading in medical dialogue scenarios.

[0062] This embodiment establishes a chain mechanism linking acoustic tag recognition with semantic and geographic reasoning, effectively filtering environmental acoustic elements that do not match the target geographic location and generating an acoustic tag set consistent with geographic semantics. The semantic consistency of the acoustic tags directly determines the authenticity and concealment of subsequent acoustic data replacement, avoiding masquerading failures caused by contradictions between acoustic and semantic content, thereby improving the credibility and technical closure of the entire geographic location masquerading process.

[0063] S40, based on the target environment acoustic tag, extract the corresponding target environment acoustic information from the preset sound data set;

[0064] In this embodiment, the target environment acoustic tag refers to a set of acoustic scene description units that correspond to the target geographical location, output by the preceding tag analysis module. In form, it can be a tag set (such as "crowd conversation," "broadcast prompt," or "elevator announcement"), with each tag representing the acoustic category of a specific sound source component. This tag set serves as an index in subsequent retrieval operations, guiding the data system to extract audio segments with consistent acoustic structure and semantic meaning from the constructed sound data set.

[0065] The sound dataset is a pre-built and structured audio resource library. Each environmental sound sample carries multi-dimensional descriptive fields such as labeled acoustic tags, geographic location metadata, sampling time period, sound source category, spectral distribution characteristics, and loudness curve. The acoustic samples undergo standardization during the acquisition phase, including various spatial structures (such as indoor, open, and enclosed spaces) and regional backgrounds to improve the coverage of subsequent matching.

[0066] The retrieval process begins with a first-round screening based on the target environment's acoustic tags, extracting a set of audio samples that perfectly match the tag semantics, serving as a candidate acoustic segment set. Building upon this, to improve acoustic coherence and the naturalness of the output sound, each candidate segment undergoes a comparative calculation based on the original environmental acoustic information, including but not limited to temporal energy envelope shape, dominant frequency band distribution range, loudness change slope, and spectral centroid shift. Each candidate segment is scored using a similarity function, and a similarity threshold is set to select high-fidelity matching segments as acoustic materials suitable for synthesis.

[0067] During the acoustic parameter matching process, if the candidate segments do not meet the full parameter similarity requirements, amplitude adjustment can be performed based on the temporal energy characteristics in the original environmental acoustic information to make the overall loudness curve of the audio segment consistent with the original background. In terms of temporal arrangement, the segments will be arranged and combined according to the segment length, pause interval, and energy fluctuation rhythm of the original environmental acoustic information to construct an audio sequence that conforms to natural listening perception.

[0068] The final generated target environment acoustic information has a label structure that matches the target geographical location, an energy distribution similar to the original environmental volume, and an acoustic rhythm that is synchronized with the temporal structure of the dialogue behavior, which enables it to be smoothly embedded into the original speech content in the subsequent synthesis stage, forming spatially consistent composite speech data.

[0069] When implementing this in an urban office environment, the target environment acoustic labels can be limited to categories such as "keyboard typing," "meeting discussion," and "printer operation." After selecting multiple high-quality segments containing the above sound source components from the sound data set, linear normalization is performed based on the average energy level extracted from the original environmental acoustic information. The Dynamic Time Warping (DTW) algorithm is then used to optimize the arrangement of the segments to ensure a natural transition with the timing of speech behavior.

[0070] In high-dynamic environments such as train stations and airports, background reverberation features can be introduced as an acoustic matching indicator. Post-processing with preset reverberation parameters can be added to the audio samples that match the acoustic labels of the target environment to enhance the spatial realism. In this case, multi-source sound mixing technology can also be used to superimpose multiple label segments (such as "broadcast voice", "pedestrian footsteps", "wheelchair gliding") in parallel to simulate a complex environment with multiple sound sources.

[0071] In implementation, a dynamic sampling mechanism can also be adopted, which means that based on the regional classification corresponding to the acoustic label of the target environment, the latest samples are dynamically pulled from the online acoustic library at runtime and combined with the original environmental volume structure for real-time adaptation to ensure the timeliness and diversity of the generated results.

[0072] Example Explanation: In the financial services sector, if a user claims to be at a commercial bank branch over the phone, to make the environmental simulation more closely resemble the actual acoustic characteristics of that location, the system first retrieves audio segments from the pre-built sound dataset based on the target environmental acoustic labels (such as "printer operation sound," "queue call prompt," and "lobby background chatter") output by the label correction module. These audio segments either perfectly match the labels or are highly similar in spectral characteristics. After acoustic feature comparison, the audio segments are selected as basic units representing the sound field composition of the target location. The structure of each audio segment is precisely identified, including its temporal energy distribution and spectral dynamics, ensuring good auditory continuity with the previously identified original environmental noise and avoiding abrupt volume changes or frequency jumps. Ultimately, these selected and modulated audio segments constitute the target environmental acoustic information used for synthesis, providing a reliable sound source basis for subsequently achieving voice masquerading effects consistent with the bank branch environment.

[0073] In the healthcare field, if a user claims to be making a call in a hospital waiting area, the system must ensure that the constructed ambient noise possesses typical acoustic elements of that type of space. When the identified target environment acoustic labels are "electronic queuing," "wheelchair gliding," and "quiet conversation among people," the system will retrieve real hospital audio samples corresponding to these labels from the sound dataset. Through an acoustic feature comparison algorithm, segments similar to the original background noise in terms of energy envelope, frequency band proportion, and temporal structure are prioritized for subsequent environment generation. This matching mechanism effectively ensures the continuity of each segment in temporal structure and perceived energy during environment construction, giving the synthesized speech the auditory quality of a specific hospital area and enhancing the credibility of the spatial masquerading.

[0074] This embodiment utilizes tag-driven data retrieval and acoustic feature matching to enable environmental sound replacement operations to move beyond static resources or template audio. It allows for the dynamic construction of environmental acoustic information that is semantically consistent with the target geographic location, has a coordinated spectral morphology, and similar energy rhythm. By leveraging structured audio tags and a similarity filtering mechanism, it effectively preserves geographically oriented acoustic elements while achieving the disguised replacement of environmental sounds without disrupting the naturalness of dialogue. This provides a realistic, smooth, and difficult-to-detect audio foundation for simulating acoustic perception in specific geographic spaces.

[0075] S50, adjust the amplitude characteristics of the target environmental acoustic information to match the amplitude characteristics of the original environmental acoustic information;

[0076] In this embodiment, the target environmental acoustic information is output by a tag-driven data retrieval module. Its sound source structure and spectral characteristics are consistent with geospatial semantics, but its performance in terms of amplitude dimensions such as volume dynamics, energy rhythm, and loudness peak differs from the original environmental acoustic information. To improve the overall coherence and camouflage quality of the final synthesized speech, structural alignment in the amplitude dimension needs to be further performed.

[0077] First, the target environment acoustic information needs to be decomposed at the frame level, dividing the entire audio segment into a series of short-time acoustic frames using fixed-duration windows. This facilitates frame-by-frame analysis and dynamic adjustment. The original environmental acoustic information is then processed simultaneously using the same frame division strategy to generate the original short-time acoustic frame sequence, ensuring that the two signal sequences are comparable in the time dimension during subsequent operations.

[0078] Subsequently, amplitude features are extracted for each frame, which typically include, but are not limited to, root mean square energy, peak amplitude, short-time energy gradient, and loudness change rate. Root mean square energy measures the overall loudness level of the frame, peak amplitude reflects the instantaneous volume extremes, and gradient and rate of change measure the inter-frame variation trend. These metrics form the basis for comparing the original and target signals in amplitude space.

[0079] In the calculation phase, the amplitude characteristics of the original short-time acoustic frame are used as a reference benchmark and compared with the corresponding characteristics of the target short-time acoustic frame. Based on the difference, a sequence of gain adjustment coefficients is constructed. The gain adjustment coefficients can be linear coefficients (used to amplify or compress the volume of the target frame), or an adaptive sliding window smoothing strategy can be adopted to keep the inter-frame gain changes consistent and avoid abrupt volume jumps during the synthesis process.

[0080] Gain adjustment coefficients are applied to the amplitude dimension of the target short-time acoustic frames. Through point-to-point amplitude scaling, the target frames maintain a high degree of consistency with the original frames in terms of energy envelope, loudness slope, and other characteristics. The adjusted set of target frames is then reassembled into a continuous audio sequence to construct background acoustic data that matches the amplitude characteristics of the original environment.

[0081] To ensure that the adjusted signal does not introduce distortion, amplitude consistency verification of the generated audio is also required. Specifically, an evaluation function based on energy envelope similarity can be used. If the deviation between the continuous audio and the original environment on the energy change curve is less than a set threshold, the acoustic information of the target environment is determined to meet the amplitude alignment requirements and can be used for subsequent synthesis.

[0082] In static environments (such as conference rooms or hospital corridors), amplitude feature extraction can be performed using fixed-length frames (e.g., 25ms per frame with a 10ms frame shift). Amplitude adjustment is achieved through weighted averaging and interpolation smoothing to enhance the gradual change in loudness. To further reduce manual intervention, a self-supervised model based on energy similarity can be used to automatically calibrate the target amplitude range, improving the efficiency of gain coefficient calculation.

[0083] In dynamic environments (such as traffic scenarios or shopping mall calls), amplitude characteristics may fluctuate significantly due to environmental noise interference. In such cases, a dynamic gain control strategy based on time-varying entropy constraints can be used to control the gain adjustment amplitude of local frames to not exceed the maximum dynamic range, thereby avoiding instantaneous volume disturbances. In this situation, speech activity information can also be introduced as a reference to avoid excessive amplification of ambient sounds in speech-dominant regions.

[0084] For samples with uneven speech distribution or long background silence intervals, a gain distribution reconstruction mechanism based on rhythm structure can be adopted, which refits the loudness rhythm of the target frame into a segmented loudness model similar to the original frame in terms of time scale, thereby improving the realism of the overall acoustic rhythm.

[0085] Example Explanation: In a financial services scenario, if a user intends to access a bank's back-end system remotely via voice, claiming to be in the lobby while actually in their residential area, the system needs to construct an audio background similar to the lobby's acoustic environment to prevent background noise from revealing their actual location. An amplitude adjustment module can adjust the target ambient acoustic information collected from the lobby to match the amplitude characteristics of the original ambient acoustic information in the user's call background, avoiding issues such as fluctuating ambient sound volume or rhythmic breaks, thus ensuring the disguised audio is both complete and deceptive.

[0086] In healthcare services, if a user attempts to make a remote appointment via voice system and claims to be in a waiting area, to avoid the system questioning the unnatural background volume caused by the static recording environment, this adjustment mechanism can be used to meticulously align the ambient sound of the target hospital's waiting area with the background loudness structure of the user's actual call environment, generating a naturally transitioning background noise stream. This enhances the immersive experience of the simulation and prevents the system from judging the discrepancy between the user's location information and the claim based on abnormal amplitude.

[0087] This embodiment constructs a frame-by-frame amplitude feature alignment mechanism, combined with precise calculation of gain adjustment coefficients and inter-frame smoothing strategies, ensuring that the target environmental acoustic information not only maintains consistency with the target geospatial information in terms of label semantics, but also closely matches the original environment in perceptual dimensions such as loudness, rhythm, and energy flow. This operation ensures that the environmental acoustic substrate possesses sufficient volume naturalness and camouflage continuity before speech and background sound synthesis, thereby effectively suppressing forgery traces caused by energy imbalance, making the subsequently constructed speech data more realistic in sound, and significantly enhancing the ability to camouflage geographical locations.

[0088] S60, the original speech content and the adjusted target environmental acoustic information are combined into analog speech data.

[0089] In this embodiment, the original speech content has been extracted during the sound source separation process and is typically represented in the form of a time-domain waveform or a spectral matrix. It includes the user's speech components but does not carry the original background sound pressure characteristics. The adjusted target environment acoustic information has undergone energy distribution, rhythm, and amplitude feature alignment processing in the preceding operations, possessing a fusionable loudness structure and spectral compatibility. The synthesis of the two must simultaneously meet the requirements of natural transition and camouflage within the temporal structure, energy control, and frequency domain space.

[0090] First, the two signals should be divided into short-time frames to ensure consistent processing granularity across the time dimension. Typically, each frame is 20 to 30 milliseconds long, and a frame shift of 10 to 15 milliseconds can achieve higher resolution. After frame division, the temporal features of the original speech content need to be extracted, including speech activity detection results, principal component energy change trajectories, and acoustic pause intervals. Based on these features, the target environment acoustic frames can be time-aligned, ensuring that background noise remains restrained in speech intervals and smoothly fills in silent intervals.

[0091] The dynamic energy balancing stage then begins. Energy analysis needs to be performed on both short-time speech frames and short-time environmental frames. By calculating the root mean square energy and loudness waveform of the short-time frames, the gain compensation coefficient to be applied to the target ambient sound frame within the speech interval is determined. This coefficient is not a uniform scaling value, but is adjusted frame-by-frame based on dynamic factors such as speech activity, ambient noise floor, and desired camouflage level. For example, in areas with active human voices, the energy level of the ambient background should be in a suboptimal state dominated by speech; conversely, in quiet areas, the continuity of background noise filling should be ensured to simulate a natural transition in real space.

[0092] After energy balancing, the speech frame and the gain-compensated ambient sound frame are frequency-domain mixed. Specifically, the two frames are subjected to Fast Fourier Transform (FFT) to obtain complex spectra, which are then superimposed. The superposition method can be based on an amplitude-phase synthesis structure, where the speech retains its original phase, and the ambient sound incorporates frequency components with suppressed amplitude, thereby enhancing the spectral consistency of the ambient sound camouflage while preserving speech intelligibility.

[0093] After frequency domain superposition, an inverse Fourier transform is performed on each frequency domain frame to restore it to a time domain waveform, resulting in a mixed time domain frame sequence. To avoid discontinuities between frame boundaries, a windowing and overlapping addition method is used for frame-level stitching. Typical window functions include the Hamming window or the symmetric cosine window.

[0094] Finally, to improve the auditory quality of the synthesized signal and eliminate spurious components introduced by spectral superposition, an adaptive noise suppression module can be applied to the complete time-domain signal. This module constructs a dynamic filter based on the energy variation curves of speech and environment frames, suppressing unnatural loudness spikes at frame gaps or at the end of speech segments. The output signal is then analog speech data containing both realistic human voice semantics and the ambient sound structure of the target geographical location.

[0095] The joint mixing process of speech and environment can be realized through an end-to-end model based on convolutional time-frequency transform. The speech frame input and the environment frame input are respectively connected to the time-frequency encoder. After multi-scale feature aggregation, energy modulation and spectrum compensation are performed in the feature layer, and then the decoder restores it into a mixed speech waveform.

[0096] Alternatively, a traditional digital signal processing approach can be adopted. First, a short-time Fourier transform module based on overlapping windowing is used to complete frame segmentation and frequency domain analysis. Then, a time-domain filter bank is used to achieve dynamic control between speech dominance and background suppression. Finally, the mixed waveform is reconstructed in the speech synthesis engine, supporting offline batch synthesis or real-time masquerading stream processing.

[0097] It can also use a nonlinear frequency domain fusion strategy to construct a weighted mask in the spectral energy distribution of each frame, and dynamically adjust the spectral fusion weights according to the speech clarity requirements, thereby improving the sense of environmental realism without sacrificing recognition accuracy, and adapting to audio security detection platforms with higher requirements for synthesized signal quality.

[0098] Example Description: In a financial customer service intelligent call scenario, when a user simulates being at a branch office via a remote voice channel, the original voice content from the customer's local microphone needs to be mixed and synthesized with real background sound extracted from that branch office. This mechanism achieves spectral-level fusion and temporal energy alignment, making the final synthesized signal resemble human voices in the acoustic environment of that branch office. This facilitates bypassing the acoustic positioning system's judgment and improves the credibility of the user's claimed location.

[0099] In telemedicine appointment systems, users may initiate voice requests remotely to obtain access to medical services or priority access within a specific geographical area. This mechanism precisely synthesizes the user's original voice with the ambient acoustic material of the target hospital's waiting area at the frame level, and modulates features such as loudness peaks and background rhythm to ensure that the voice data collected by the system exhibits highly consistent spatial contextual characteristics. This avoids the system's detection of anomalies in location information and enhances the credible expressiveness of the simulated environment.

[0100] This embodiment introduces a speech-activity-driven background sound energy modulation strategy and a frequency domain fusion mechanism, making the fusion of ambient sound and speech in terms of temporal structure and loudness distribution more natural. This significantly reduces forgery traces in synthesized speech due to energy fluctuations, background discontinuities, and phase conflicts. The synthesized signal achieves a balance between semantic fidelity and acoustic continuity, effectively simulating the acoustic scene of a specific target geographical area, thus improving the overall credibility and undetectability of geographical location spoofing. Compared to simple audio splicing methods, this fusion mechanism has higher adaptability and deception capabilities.

[0101] This invention relates to the field of speech processing technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for generating simulated environmental acoustic speech, comprising: acquiring and separating mixed speech data to generate original speech content and original environmental acoustic information; converting the original speech content into first text information; determining a target environmental acoustic label based on the original environmental acoustic label, the first text information, and target geographical location information; determining the target environmental acoustic information from a preset sound data set through label matching; adjusting the amplitude characteristics of the target environmental acoustic information to match the amplitude characteristics of the original environmental acoustic information; and synthesizing the original speech content and the adjusted target environmental acoustic information into simulated speech data. This invention introduces target geographical location information into the process of selecting environmental acoustic labels and generating acoustic information, enabling the synthesized environmental noise to possess acoustic characteristics consistent with the target location. Furthermore, by adjusting the amplitude characteristics, it achieves acoustic fusion between the simulated environment and the original environment, making the final generated simulated speech data more realistic in terms of acoustic distribution and geographical consistency, and more difficult to detect camouflage traces, thereby improving the acoustic camouflage effect of the speech content at the target geographical location.

[0102] In one embodiment, step S10 above includes:

[0103] S101 acquires mixed voice data through a multi-channel audio acquisition device;

[0104] S102, Perform noise suppression preprocessing on the mixed speech data to generate noise-reduced speech data;

[0105] S103, The noise-reduced speech data is input into a pre-trained sound source separation model, which includes a parallel processing branch for human voice extraction and an environmental acoustic feature extraction branch.

[0106] S104, output the original speech content through the human voice extraction branch, and output the original environmental acoustic information through the environmental acoustic feature extraction branch;

[0107] S105, based on the acoustic characteristic distribution of the target geographical environment, adjust the frequency response characteristics of the original environmental acoustic information through frequency band energy matching.

[0108] In this embodiment, mixed speech data is acquired using a multi-channel audio acquisition device to enhance the ability to capture signals from multiple sound sources in space, thereby enabling subsequent sound source separation operations to achieve higher resolution and spatial reconstruction accuracy. A multi-channel audio acquisition device typically includes at least two physically distributed microphone arrays, with each channel recording the raw sound pressure level signal under different angles, distances, and phase conditions relative to the main microphone. This structure can provide directional and source localization references for the signal separation model through time delay estimation, coherence characteristic analysis, and other methods.

[0109] Noise suppression preprocessing of mixed speech data aims to reduce the impact of ambient noise and irregular interference, thereby improving the input signal-to-noise ratio of subsequent sound source separation models. This processing is typically implemented based on spectral subtraction, Wiener filtering, adaptive spectral attenuation, or neural network noise estimation methods. The goal is to suppress random noise components while preserving speech and effective background acoustic features as much as possible, and the output signal should have good semantic intelligibility and acoustic structural integrity.

[0110] The core step in the separation operation is to input the denoised speech data into a pre-trained sound source separation model. Sound source separation models are typically implemented using deep learning architectures, such as time-frequency domain networks, convolutional gating networks, dual-path coding networks, or mask learning models based on attention mechanisms. This model includes two parallel branches: a human voice extraction branch and an environmental acoustic feature extraction branch. The human voice extraction branch aims to preserve speech components by learning the phoneme boundaries, dynamic features of articulation, and prosodic rhythm of human speech to restore the dominant speaker's speech flow. The environmental acoustic feature extraction branch focuses on non-speech regions and background energy regions, extracting structural environmental sound source information, including natural sounds, crowd activity, mechanical equipment operation, and spatial reverberation.

[0111] The output of the original speech content in the human voice extraction branch indicates that the model has successfully extracted the main speech signal, eliminated environmental interference and blurred background regions, and made the speech content directly usable for tasks such as speech recognition, speaker analysis, or semantic understanding. The output of the original environmental acoustic information in the environmental acoustic feature extraction branch demonstrates that the model, while maintaining the temporal integrity and structural features of the environmental sound, has eliminated the dominant speech signal, making the background sound a sound field expression with stable structure, complete spectrum, and clear energy distribution.

[0112] Adjusting the frequency response characteristics of the original environmental acoustic information based on the acoustic feature distribution of the target geographic environment is to achieve consistency between the ambient sound and the target location characteristics. The acoustic feature distribution is typically modeled from historical data, including the energy distribution spectrum of common frequency bands, background rhythm patterns, and spatial reverberation characteristics of the corresponding target geographic area over a specific time period. The matching process, based on the frequency band energy distribution, utilizes linear filter banks, equal-loudness correction curves, or data-driven dynamic frequency domain masking techniques to adjust the frequency energy structure of the original environmental acoustic information segment by segment. The final output frequency response adjustment result will more closely match the acoustic profile of the target geographic environment in terms of spectral structure, providing more reliable background support for subsequent replacement synthesis.

[0113] This embodiment effectively extracts a stable and clear speech backbone and structurally complete environmental background acoustic information from mixed speech through distributed multi-channel audio acquisition, preprocessing based on signal-to-noise ratio enhancement, and a dual-branch structure sound source separation mechanism. Furthermore, by combining the frequency band energy model of the target geographical environment with frequency response adjustment of the environmental sound, the environmental sound achieves geographical consistency in its spectral structure, thus providing accurate, realistic, and continuous acoustic material for subsequent synthetic simulation of acoustic scenes. This processing chain not only improves the accuracy of sound source separation but also significantly enhances the credibility of environmental sound reconstruction in location camouflage tasks, improving the deception capability and application robustness of synthesized speech.

[0114] In one embodiment, step S20 above includes:

[0115] S201, Perform speech activity detection on the original speech content and divide the original speech content into multiple speech segments;

[0116] S202, each speech segment is input into a pre-trained multi-dialect speech recognition model to generate corresponding intermediate text data;

[0117] S203, perform dialect consistency correction on the intermediate text data, and adjust the dialect vocabulary in the intermediate text data based on the target geographical location information to generate dialect-corrected text;

[0118] S204. Add timestamp annotations to the dialect-corrected text according to the timing characteristics of the original speech content to generate text data with timing tags.

[0119] S205. Aggregate all the text data with timing tags to generate the first text information.

[0120] In this embodiment, voice activity detection is performed on the original speech content, aiming to accurately identify the valid intervals containing language expressions from the continuously recorded speech stream, and avoid misinterpreting non-language paragraphs such as silence, breathing, coughing, etc. as semantic information sources. Voice activity detection is usually discriminated based on short-time energy, zero-crossing rate, spectrogram convolution pattern or end-to-end neural network. It not only needs to detect the existence of speech segments, but also should retain the boundary order and start / end annotations of the segments in the time dimension. The result output by voice activity detection directly determines the input granularity and accuracy of the subsequent speech recognition model. Therefore, the physical start / end boundaries of each speech segment must be structurally encoded.

[0121] Input each speech segment into a pre-trained multi-dialect speech recognition model, aiming to fully accommodate the accent variations in different language regions or different population backgrounds, and make the recognition system robust to non-Mandarin pronunciations, mixed dialects, speech rate differences, intonation fluctuations, etc. The speech recognition model usually used integrates a phoneme-level modeling structure, an acoustic attention mechanism and a language context decoder, and can adopt an end-to-end CTC structure or an RNN-Transducer architecture. The model needs to introduce real speech samples from multiple geographical regions in the pre-training stage to build a general recognition network that can cover mainstream dialects (such as Cantonese, Minnan dialect, Wu dialect, etc.) and variant pronunciations, so as to ensure that different regional pronunciations can be mapped to unified intermediate text data in the preliminary recognition stage.

[0122] Perform dialect consistency correction on the intermediate text data, aiming to solve the non-standard text structures generated by the multi-dialect model under the problems of pronunciation ambiguity or morpheme splitting, and normalize the words with similar pronunciations but significant semantic differences. This correction process not only needs to refer to pronunciation rules, but also needs to dynamically adjust the word usage probability in combination with the geographical language distribution rules. On this basis, introduce the target geographical location information to adjust the dialect vocabulary expression, in order to make the recognition result closer to the pragmatic habits of the target claimed area. For example, the same expression may use "come on" in the Sichuan-Chongqing area, while in South China it may be expressed as "go ahead", with the same meaning but different pragmatic environments. The adjustment process can be realized by rule mapping, statistical translation or reconstruction strategies based on geographical language models, and retain the historical annotations before and after the word adjustment.

[0123] Adding timestamps to the dialect-corrected text based on the temporal characteristics of the original speech content is crucial for maintaining strict alignment between the text structure and the speech source. Temporal feature extraction references the start and end frame numbers of the speech segments or the original sampling timestamps, and inserts time stamps at the character, word, or sentence level by comparing and identifying the output character boundaries or phoneme segmentation points. Timestamp annotation not only supports text-to-speech lookup but also provides a synchronization basis for subsequent synthesis stages.

[0124] Aggregating all time-tagged text data means concatenating the text output from multiple independent speech segments in chronological order to reconstruct continuous and complete language expression units. This aggregation process also needs to handle grammatical information such as sentence segmentation, pauses, and intonation transitions, ensuring no logical breaks or semantic jumps between segments. The final output structured text should support sentence-level retrieval, time-aligned playback, and geographic semantic comparison.

[0125] This embodiment not only achieves semantic extraction from unstructured speech by converting speech content into text information, but also enhances the adaptability between speech content and target geographic location by introducing geographic language correction and timestamp annotation mechanisms. Speech activity detection improves speech segment recognition accuracy, multi-dialect recognition models enhance language diversity coverage, and a geographic location-driven vocabulary adjustment strategy solves the problem of similar pronunciations but semantic inconsistencies. This results in generated text that more closely matches the user's claimed geographic region characteristics at the linguistic level, providing semantic support and temporal alignment for subsequent environmental acoustic synthesis and geographic location spoofing. This processing flow effectively improves the logical consistency and synthesis fidelity of the entire chain in geographic adaptation synthesis tasks.

[0126] In one embodiment, step S30 above includes:

[0127] S301, Perform spectrum segmentation on the original environmental acoustic information to generate multiple acoustic segments;

[0128] S302, Perform multi-level classification operations on each acoustic segment to generate a preliminary set of environmental labels;

[0129] S303, Select the corresponding sub-recognition model according to the label type in the preliminary environmental label set;

[0130] S304, perform acoustic feature extraction on the acoustic segment to generate a multi-dimensional feature vector, and input the multi-dimensional feature vector into the corresponding sub-recognition model to generate a fine-grained classification result;

[0131] S305: Merge all fine-grained classification results to generate refined environment labels;

[0132] S306, The refined environmental labels are correlated with the acoustic feature mapping table in the geographic knowledge base to screen out candidate labels that match the target geographic location information;

[0133] S307, Extract geographical keywords from the first text information;

[0134] S308, Construct structured input data containing candidate tags, geographic keywords, and target geographic location information;

[0135] S309, invoke the pre-trained geographic association language model to generate a label replacement strategy based on the structured input data and geographic knowledge base;

[0136] S310, according to the label replacement strategy, delete or replace conflicting items in the candidate labels, and generate a corrected candidate label set as the target environment acoustic label.

[0137] In this embodiment, spectral segmentation of the raw environmental acoustic information involves dividing the acoustic data into segments within a two-dimensional time-frequency space to extract fine-grained structures. This segmentation typically employs a short-time Fourier transform to divide the continuous signal into multiple time-frequency units, with each segment preserving the instantaneous intensity variations of the sound source across different frequency bands. This segmentation method not only maintains the changing trends of the speech frequency bands but also facilitates the processing of local features of each unit in the subsequent classification stage. The granularity of the spectral segmentation can be dynamically adjusted according to the signal sampling rate and analysis objectives, typically controlled within a window length of 20–40 milliseconds.

[0138] Performing multi-level classification on each acoustic segment aims to initially identify acoustic patterns in the environment in a hierarchical manner. The multi-level classification structure typically begins with broad category identification, such as natural sounds, mechanical noise, crowd noise, and traffic noise, and then refines the results into subcategories or smaller categories based on the previous level's findings. For example, after identifying "natural sounds," it might be further classified as "birdsong," "wind," or "insect chirping." Each level of classification can use convolutional neural networks, residual networks, or spectrogram-based Transformer networks as the basis for feature recognition, outputting a preliminary set of environmental labels and retaining classification confidence scores.

[0139] Selecting a corresponding sub-recognition model based on the label types in the initial environmental label set aims to improve recognition accuracy by calling local models with specialized recognition capabilities for different acoustic categories. Sub-recognition models typically employ enhanced training with fine-class samples; for example, bird sound recognition models focus on transient features with high frequency density and short duration, while crowd sound models emphasize continuous speech features. The model selection process is based on the mapping relationship between label types and can automatically activate the corresponding sub-network structure through the controller network, ensuring that the classification model matches the sound source type.

[0140] Acoustic feature extraction from acoustic segments involves converting the original waveform into a multidimensional feature vector with statistical descriptive capabilities. Extraction methods include using parameters such as Mel-frequency cepstral coefficients, spectral centroid, bandwidth, and zero-crossing rate. Alternatively, a self-supervised representation learning structure can be used to automatically extract the embedding vector. The extracted multidimensional feature vector is then input into a correspondence recognition model to generate fine-grained classification results. The output format is a probability distribution, category identifier, or audio event code. This classification result possesses stronger geographic resolution and can support the identification of specific species calls or localized sound sources.

[0141] Merging all fine-grained classification results involves temporal alignment and spatial aggregation of the label outputs from multiple segments to obtain a complete environmental semantic structure. The merging process must handle duplicate labels, low-confidence entries, and temporal conflicts, while preserving the temporal coverage and intensity statistics of the labels to generate refined environmental labels. This label set should have sufficient granularity and breadth to cover the main sound source information in the user's current audio environment.

[0142] The association verification between refined environmental labels and the acoustic feature mapping table in the geographic knowledge base aims to determine the logical consistency between labels and target geographic locations based on existing geographic distribution knowledge. The geographic knowledge base records the distribution of acoustic labels that may appear in different geographic regions during different seasons and time periods; for example, some bird calls only occur in specific regions and months. Association verification is performed through graph structure matching, rule-based semantic reasoning, or label-location-time triple matching to generate candidate labels that match the target geographic location information.

[0143] Extracting geographic keywords from the initial text information involves comparing the speech-recognized text with a geographic lexicon to identify text fields related to place names, landmarks, and locations. Geographic keywords include geographic entity names, administrative divisions, and facility types (such as "hospital," "square," and "subway station"). Named entity recognition models or dictionary-based matching methods can be used during extraction. These keywords will contribute to the contextual conditions for constructing subsequent language understanding models.

[0144] Constructing structured input data containing candidate tags, geographic keywords, and target geographic location information involves representing multi-source information in a unified data format with contextual hints, typically using key-value pairs such as JSON or XML. This structured data helps the language model understand geographic consistency and tag logic within the context and is passed as input to the language generation or matching / determination module.

[0145] Calling a pre-trained geographic association language model refers to using a language understanding structure trained on a large-scale corpus to determine the rationality of the current label set and generate label adjustment suggestions. The model structure can adopt T5, BART, or GPT architectures, incorporating geographic label descriptions, acoustic scene semantics, and spatial knowledge during training. This language model generates label replacement strategies based on the input structured data and geographic knowledge base, including deletion, merging, renaming, or substitution operations, and possesses natural language interpretable output capabilities.

[0146] Removing or replacing conflicting items in candidate labels according to the label replacement strategy involves performing adjustments generated by the model. Conflicting items can include labels that logically contradict the target geographic location information, labels with conflicting time periods, or labels with excessively low geographic distribution probabilities. Label adjustments must retain confidence levels, replacement sources, and adjustment records. The final output is a set of corrected candidate labels, forming the target environment acoustic labels.

[0147] This embodiment effectively identifies potential contradictions between ambient sound and geographic claims by verifying the semantic consistency between the original ambient acoustic tags and the target geographic location, thus improving the ability to analyze implicit geographic clues in ambient sound. The fusion of fine-grained tag recognition and geographic distribution mapping not only improves the accuracy of tag judgment but also enhances the reliability of location adaptation. Introducing geographic keywords into the text and modeling them together with acoustic tags allows the language model to assess the geographic rationality of tag combinations while understanding semantics. This supports the determination of whether there are tag contents in ambient sound that violate geographic authenticity, providing a more reliable tag foundation for subsequent simulated audio generation. This processing mechanism, through intelligent language model and multi-source tag fusion, effectively solves the limitation of lack of spatial adaptability in ambient sound annotation.

[0148] In one embodiment, step S40 above includes:

[0149] S401, Match the target environment acoustic labels with the acoustic labels in the preset sound data set to generate a candidate acoustic segment set;

[0150] S402, Extract the temporal energy distribution characteristics and spectral characteristics of the original environmental acoustic information;

[0151] S403, Perform acoustic parameter matching operation on each acoustic segment in the candidate acoustic segment set, and filter out acoustic segments whose temporal energy distribution characteristics and spectral characteristics are more than a preset similarity threshold with the original environmental acoustic information.

[0152] S404, Based on the temporal energy distribution characteristics of the original environmental acoustic information, the selected acoustic segments are subjected to amplitude modulation processing to generate modulated acoustic segments.

[0153] S405, the modulated acoustic segments are arranged and combined according to the temporal characteristics of the original environmental acoustic information to generate a continuous acoustic information stream;

[0154] S406, perform acoustic feature consistency verification on the continuous acoustic information stream. If the verification passes, the continuous acoustic information stream is used as the target environmental acoustic information.

[0155] In this embodiment, the target environment acoustic tags are matched with acoustic tags in a preset sound data set. This is achieved through methods such as keyword alignment, tag semantic embedding vector matching, or index retrieval based on tag classification structures. Audio segments matching the semantic tags are selected from a dataset with geographically distributed features. The preset sound data set stores structured acoustic data organized by geographical region, environment type, and typical sound sources. Each record includes a tag identifier and audio segment metadata, including sound source type, acquisition environment, duration, and acquisition device parameters. The matching process can be completed using methods such as cosine similarity, edit distance, or tree-structured tag path comparison, generating a candidate acoustic segment set as the object of subsequent parameter refinement operations.

[0156] Extracting the temporal energy distribution and spectral characteristics of raw environmental acoustic information involves transforming the physical properties of the original acoustic background into comparable acoustic descriptive indicators. Temporal energy distribution characteristics include the root mean square energy per unit time window, short-time energy envelope, mean energy, and dynamic range, reflecting the loudness variation pattern of the sound source over different time periods. Spectral characteristics include parameters such as dominant frequency, frequency band energy proportion, spectral centroid, and spectral slope, describing the distribution trend of sound along the frequency dimension. These features are typically extracted using short-time Fourier transform or Mel-frequency spectral analysis and represented in vector form.

[0157] For each segment in the candidate acoustic segment set, an acoustic parameter matching operation is performed. This involves comparing the candidate segment with the temporal energy distribution characteristics and spectral characteristics of the original environmental acoustic information one by one to determine their similarity in loudness variation patterns and frequency structure. Similarity calculation methods include dynamic time warping, multidimensional feature vector Euclidean distance calculation, or KL divergence matching. The matching results are filtered according to a preset similarity threshold, retaining only those segments that are similar to the original environmental acoustic information in terms of energy dynamics and spectral morphology. This step ensures that the target environmental acoustic information not only has semantic label consistency but also perceptual auditory similarity.

[0158] Based on the temporal energy distribution characteristics of the original environmental acoustic information, amplitude modulation is applied to the selected acoustic segments to achieve local loudness adaptation. During this process, the system performs frame-level gain adjustment on the candidate segments according to the energy profile of each frame of the original environmental acoustic information. Amplitude modulation employs dynamic linear scaling or adaptive adjustment based on the envelope function, ensuring that the modulated segments maintain consistency with the original environmental background in terms of macroscopic energy profile, thus preserving continuity and naturalness at the auditory level. Historical frame gain records must be retained during the amplitude modulation process to prevent abrupt changes that could cause a sense of discontinuity.

[0159] Arranging and combining modulated acoustic segments according to the temporal characteristics of the original environmental acoustic information reconstructs an acoustic stream consistent with the duration and rhythm of the original background environment. Temporal characteristics include the order in which segments appear on the timeline, adjacent intervals, and the length of silent periods. During the arrangement and combination process, the smooth transitions between segments need to be addressed, for example, by using crossfade-in / fade-out and transition frame interpolation to avoid non-linear alternations at connection points. The synthesized continuous acoustic information stream should be strictly aligned with the structure of the original environmental acoustic information in the time dimension.

[0160] Acoustic feature consistency verification of continuous acoustic information streams involves comparing the final spliced ​​result with the original environmental acoustic information at a holistic level to determine whether it meets similarity requirements in energy profile, spectral density, and loudness envelope. Verification methods include multi-scale acoustic similarity comparison, auditory model scoring, or calculating a consistency score based on automatic perception scoring metrics (such as PESQ and STOI). If the verification result exceeds a set threshold, the continuous acoustic stream is considered acceptable as target environmental acoustic information for subsequent processing.

[0161] This embodiment achieves dual fitting of the target environment's acoustic information at both the semantic and acoustic levels through acoustic data filtering guided by label semantics and matching based on physical acoustic features. This process ensures that the generated background sound not only matches the labels in content but also closely approximates the original acquisition environment in terms of energy fluctuations, frequency density, and rhythmic structure. Thus, while maintaining the user's actual speech content, it reconstructs a geographically plausible and perceptibly natural background sound environment. This processing mechanism effectively avoids the spatiotemporal discrepancies caused by using fixed background noise templates in existing methods, providing a reliable acoustic foundation for the spatial plausibility and realistic appearance of the subsequently synthesized simulated speech.

[0162] In one embodiment, step S50 above includes:

[0163] S501, the target environmental acoustic information and the original environmental acoustic information are processed by frame segmentation to generate multiple target short-time acoustic frames and multiple original short-time acoustic frames.

[0164] S502, extract the target temporal amplitude features of each target short-time acoustic frame, and extract the original temporal amplitude features of each original short-time acoustic frame;

[0165] S503, based on the original time-domain amplitude characteristics and the target time-domain amplitude characteristics, determine the gain adjustment coefficient for each target short-time acoustic frame;

[0166] S504, apply the corresponding gain adjustment coefficient to each target short-time acoustic frame to generate an amplitude-matched target short-time acoustic frame.

[0167] S505, compare the temporal amplitude features of the target short-time acoustic frame for amplitude matching with the corresponding original short-time acoustic frame;

[0168] S506, if the difference between the temporal amplitude feature of the amplitude-matched target short-time acoustic frame and the temporal amplitude feature of the corresponding original short-time acoustic frame is lower than a preset difference threshold, then the amplitude-matched target short-time acoustic frame is marked as a valid frame.

[0169] S507, all valid frames are overlapped and added together to generate target environment acoustic information.

[0170] In this embodiment, the target environmental acoustic information and the original environmental acoustic information are processed by framing separately to model the energy variation trend of the acoustic signal in the time domain with high resolution. Each segment of acoustic information is divided into short frames of fixed length, with the duration of each frame typically between 20ms and 50ms, ensuring sufficiently detailed energy estimation without introducing excessive processing delay. The framing method can be implemented based on a sliding window, allowing a certain proportion of overlap between frames to improve the ability to capture the characteristics of signal boundary regions.

[0171] Extracting the target temporal amplitude features of each short-time acoustic frame and the original temporal amplitude features of each original short-time acoustic frame involves converting each acoustic frame into several statistics representing its loudness and dynamic range. Commonly used amplitude features include root mean square energy (RMS), peak energy, and energy dynamic range, all of which reflect the changes in sound pressure intensity of the signal over a short period of time. The extraction method is usually based on the sum of squares of frame-level sampling points, which has good response speed and computational stability.

[0172] Based on the original and target temporal amplitude characteristics, the gain adjustment coefficient for each target short-time acoustic frame is determined by constructing a frame-level gain function to achieve dynamic alignment of the two amplitude trajectories. The gain coefficient can be calculated using a linear scaling factor, error back-calculation under the minimum mean square error (MSE) criterion, or a normalized offset strategy. The application of gain is limited by a set gain change rate window to prevent auditory discomfort caused by abrupt changes.

[0173] Applying a corresponding gain adjustment coefficient to each target short-time acoustic frame scales its amplitude range in the amplitude domain of each frame, thereby adjusting it to match the loudness level of the original environmental acoustic information. This processing does not change the frequency structure or duration of the frame, only affecting the envelope shape of the temporal amplitude, offering strong adaptability and high control granularity. The gain can be applied using a piecewise linear gain function or a smooth exponential adjustment function to enhance naturalness.

[0174] Comparing the temporal amplitude features of the target short-time acoustic frame for amplitude matching with the corresponding original short-time acoustic frame evaluates whether the modulation effect meets the matching conditions on a frame-by-frame basis. This comparison typically uses Euclidean distance, cosine similarity, or relative energy difference as difference indicators to quantify the deviation between amplitude envelopes. To improve accuracy, a weighted window function can be introduced during calculation to highlight the main energy regions within the frame.

[0175] If the amplitude characteristic difference between the adjusted target short-time acoustic frame and the original short-time acoustic frame is lower than a preset difference threshold, it is marked as a valid frame. This is a result filtering mechanism to ensure that the frames entering the synthesis stage have stability and reasonableness in terms of perceived intensity. The threshold setting is based on the subjectively acceptable range of difference, usually set within an energy deviation of no more than 2dB, to ensure the precision of the matching.

[0176] Overlapping and summing all valid frames reconstructs discrete frame data into continuous target environmental acoustic information. This overlapping and summing method can be achieved through windowing and accumulation, ensuring a smooth transition between adjacent frames and preventing distortion at breakpoints and boundaries. The reconstructed result exhibits stable and natural temporal energy fluctuations, making it suitable for subsequent speech synthesis processes.

[0177] This embodiment establishes a frame-level amplitude mapping relationship between the target environmental acoustic information and the original environmental acoustic information in the time domain, and introduces gain adjustment, error verification, and effective frame selection mechanisms. This allows for the matching of loudness dynamics and temporal energy patterns between two audio sources without altering the semantics and frequency structure of the audio content. This matching process significantly improves the natural integration of the target environmental acoustic information with the original environmental data, effectively eliminating problems commonly found in pseudo-synthesized background sounds, such as intensity gaps, energy jumps, or sudden increases in sound pressure levels over short periods.

[0178] In one embodiment, step S60 above includes:

[0179] S601, the original speech content and the adjusted target environmental acoustic information are processed by frame segmentation to generate speech short frames and environmental short frames respectively.

[0180] S602, based on the temporal characteristics of the original speech content, the speech short-time frame and the environment short-time frame are time-axis aligned to generate a synchronized speech frame and a synchronized environment frame.

[0181] S603, Based on the temporal energy distribution characteristics of the synchronized speech frame and the current energy level of the synchronized environment frame, determine the gain compensation coefficient of the synchronized environment frame.

[0182] S604, The synchronization environment frame is processed based on the gain compensation coefficient to generate a gain-balanced environment frame;

[0183] S605, the synchronous voice frame and the gain-balanced environmental frame are superimposed in the frequency domain to generate a hybrid frequency domain frame.

[0184] S606, Perform an inverse Fourier transform on the hybrid frequency domain frame to generate a hybrid time domain frame;

[0185] S607, The mixed time-domain frames are overlapped and added together to generate a continuous time-domain signal;

[0186] S608, perform adaptive noise suppression processing on the continuous time-domain signal to generate analog speech data.

[0187] In this embodiment, the original speech content and the adjusted target environmental acoustic information are processed by framing separately to address the inconsistency in frame granularity between the two types of audio in the temporal dimension, ensuring a unified time reference during subsequent fusion. Framing operations typically employ a fixed-length sliding window, with each frame length set at the tens of milliseconds level, and an inter-frame overlap mechanism can be introduced to preserve sufficient short-term continuity. After this processing, the speech signal and environmental acoustic information generate short-time speech frames and short-time environmental frames respectively, providing the basic units for time axis alignment and energy balancing operations.

[0188] Based on the temporal characteristics of the original speech content, short-time speech frames and short-time environment frames are time-aligned to ensure that both types of frames participate in the mixed processing synchronously within the same time period. Temporal characteristics may include the start and end times of speech frames, the distribution of active and silent segments of speech, etc. The time-axis alignment process is achieved through resampling, interpolation, or frame position correction techniques, so that the unfolding of environment frames in the time domain maintains the same length and step distribution as speech frames, generating synchronized speech frames and synchronized environment frames.

[0189] Based on the temporal energy distribution characteristics of the synchronized speech frame and the current energy level of the synchronized environment frame, a gain compensation coefficient is calculated to achieve background energy adaptation under speech dominance. The energy distribution of the speech frame can be modeled using root mean square energy or logarithmic energy, while the energy of the environment frame is determined by its envelope peak value or average sound pressure level. The gain compensation coefficient is dynamically set according to the difference between the two, so that the background sound pressure level is appropriately reduced when speech is present, and the background is maintained or enhanced when speech is absent, achieving auditory stability in camouflaged scenarios. The gain compensation strategy can employ a smooth transition function to avoid auditory abrupt changes.

[0190] Applying gain compensation coefficients to synchronized environment frames to generate gain-balanced environment frames is a specific implementation of adjusting background noise loudness at the single-frame level. The adjustment process is accomplished by multiplying the frame by the gain coefficient, without introducing spectral reconstruction or distortion, thus maintaining the frequency structure. This operation directly applies the control signal to the amplitude domain, making it suitable for background adaptation under various speech intensity conditions.

[0191] The frequency-domain superposition of synchronized speech frames and gain-balanced environmental frames aims to address the phase interference and energy imbalance issues caused by direct time-domain superposition during the synthesis process. Frequency-domain superposition first performs a Fast Fourier Transform on the two frames to obtain their respective amplitude and phase information, then performs weighted fusion in the amplitude domain, and finally reconstructs and synthesizes a hybrid frequency-domain frame. This method achieves component fusion while avoiding signal cancellation.

[0192] Performing an inverse Fourier transform on the mixed frequency domain frame to reconstruct the mixed time domain frame restores the frequency domain processing result to a synthesizable time domain signal. The inverse transform operation preserves spectral characteristics and phase structure, avoiding the introduction of additional frequency components or time domain distortion, and is a fundamental step in audio signal processing.

[0193] Generating a continuous time-domain signal by overlapping and adding mixed time-domain frames is a necessary means to achieve smooth transitions in multi-frame synthesis. Each frame undergoes window function weighting during synthesis, followed by accumulation along the time axis to eliminate discontinuities between frame boundaries. Overlapping and adding not only improves signal fluency but also reduces timbre abruptness during speech-to-background fusion.

[0194] Adaptive noise suppression processing is applied to continuous time-domain signals to further remove high-frequency noise or energy jitter that may remain at frequency band edges and at the junctions of speech and environment segments in the post-synthesis stage. This process uses dynamically estimated noise spectra as suppression templates and adjusts the gain of each frequency band according to the changes in speech energy density over different time periods, effectively improving the audible quality and consistency of the final simulated speech data.

[0195] Example Description: In a financial services scenario, when a user applies for remote account opening via telephone, the system first collects mixed voice data from the call. Combining this with the multi-channel acquisition configuration of the telephone call system, noise suppression preprocessing is performed on the mixed data to reduce background interference. This processed audio data is then input into a deployed parallel separation model for human voice and ambient sound to extract the user's voice content and the acoustic background features of their environment. The extracted voice signal is further fed into a multi-dialect speech recognition model, generating preliminary text data based on the semantic structure and sentence pattern of the speech. To ensure semantic consistency between the text and the target address, the system corrects specific dialect terms based on the language expression habits of the target geographical location and adds timestamps to the text according to the original temporal organization structure of the speech, ultimately generating structured text information as a representation of the speech semantic content.

[0196] Simultaneously, environmental acoustic information is divided into multiple acoustic segments through spectral segmentation, and its acoustic label content, including acoustic scene elements such as traffic noise, bird calls, and crowd activity, is identified through coarse classification and subclass recognition models. This preliminary identification result is then matched with geographic acoustic features stored in a localized acoustic knowledge base. Combined with location terms extracted from structured speech text and the user's claimed target location, a structured input with geographic semantic judgment dimensions is generated. This structured input is fed into a trained geolinguistic model, which evaluates the consistency logic between the current background acoustic labels and the target address. Based on the generated replacement strategy, candidate labels are adjusted, ultimately forming target environmental acoustic labels adapted to the target geographic location.

[0197] The system then retrieves corresponding acoustic samples from audio data resources based on the target label, matches and extracts background sound segments, and performs energy and rhythm consistency modulation on these acoustic samples based on the temporal energy and spectral characteristics of the user's original background environment. The modulated background sound segments are reordered on the timeline and then completely spliced ​​and synthesized, while verifying the consistency between their acoustic statistics and the original background.

[0198] The generated target environmental acoustic information needs to be further matched with the original background sound at the amplitude level. To this end, the system frames both the original and synthesized background sounds separately, extracts the energy distribution and dynamic range features of each frame, and processes the synthesized audio frame by frame by calculating gain adjustment coefficients to ensure that the final background closely matches the original state in terms of overall loudness and fluctuation characteristics. Only frames that pass the error verification with the original environmental features are included in the final synthesis.

[0199] Finally, the system performs speech-environment fusion with the adjusted target environment background. The fusion process first precisely aligns the speech frames and background frames on the timeline and calculates the background gain adjustment coefficients corresponding to the speech activity level in each frame, adapting to the energy differences between the speech and background. After synthesis in the frequency domain, the system performs an inverse Fourier transform and synthesizes the time-domain signal using an overlapping windowing method. Finally, a dynamic noise suppression module is added to optimize sound quality and coherence, outputting complete speech data that constitutes the background of the simulated target location. This process ensures the consistency of the user's speech semantics while adjusting the background information to highly match their claimed geographical location, increasing the difficulty of camouflage for subsequent verification systems that identify the user's true location through background acoustic content.

[0200] In healthcare services, remote diagnosis and treatment, chronic disease follow-up, and online check-ups are becoming increasingly common, with patients communicating with doctors or health management platforms via voice calls. To ensure the geographical compliance of voice interaction data during use, and to mitigate identification risks arising from sensitive diseases, privacy concerns, or inconsistencies in treatment areas, it is necessary to construct simulated voice data that reflects the environmental characteristics of the target medical institution. For example, if a user is at home but claims to be at a hospital outpatient clinic, the acoustic characteristics of the background noise in their call can be reconstructed without altering the voice content.

[0201] The process begins by collecting the user's voice input and separating the clear human voice from the surrounding background. The voice is then processed through speech recognition, dialect correction, and time stamping. Preliminary acoustic tags, such as call bells, crowd noise, and corridor echoes, are extracted from the background. Combining these with the environmental sound characteristics expected of the user's claimed location (e.g., the outpatient area of ​​a tertiary hospital), such as specific broadcast prompts, nurse call frequencies, and the rhythm of footsteps and gurney sounds, matching real sound source segments are selected from a pre-set sound database. These segments are then subjected to intensity adjustment and time splicing to construct a target environmental sound that matches the original background in energy distribution, spectral structure, and overall sound.

[0202] While ensuring natural sound and semantic integrity, the constructed ambient sound is precisely synchronized with the original human voice. Through frequency domain fusion and noise smoothing, a simulated speech data is output, making it sound completely consistent with the acoustic characteristics of the target medical location. This not only supports subsequent scenarios such as medical record compliance review, voice data geotagging, or voice model training, but also effectively hides the patient's real call location, protecting their medical privacy.

[0203] This embodiment achieves a natural transition between the background sound environment and the target geographical scene by mixing the original speech content with the adjusted target environmental acoustic information through time alignment, energy compensation, frequency domain fusion, and adaptive noise reduction. This process maintains the integrity of semantic expression. Compared to direct time-domain superposition, this method effectively eliminates common phase conflicts, energy abrupt changes, and boundary breaks in synthesized signals through multi-stage gain adjustment and spectrum fusion. It also avoids abrupt background noise in the dominant speech segments, improving the coherence and realism of the spoofed speech. The resulting simulated speech data possesses a reasonable acoustic hierarchy and is difficult to reverse engineer as pseudo-synthesized audio, providing reliable support for geolocation spoofing.

[0204] In one embodiment, an environmental acoustic simulated speech generation device is provided, which corresponds one-to-one with the environmental acoustic simulated speech generation method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the environmental acoustic simulation speech generation device of the present invention. The modules include a sound source separation module 10, a speech recognition module 20, an environmental tag recognition module 30, an acoustic retrieval and matching module 40, an acoustic modulation module 50, and a speech synthesis module 60. Detailed descriptions of each functional module are as follows:

[0205] The sound source separation module 10 is used to acquire mixed speech data and perform separation operations on the mixed speech data to generate original speech content and original environmental acoustic information;

[0206] The speech recognition module 20 is used to convert the original speech content into first text information;

[0207] The environmental label recognition module 30 is used to identify the original environmental acoustic label from the original environmental acoustic information, and to determine the target environmental acoustic label through the intelligent processing module based on the original environmental acoustic label, the first text information and the target geographical location information.

[0208] The acoustic retrieval and matching module 40 is used to extract the corresponding target environment acoustic information from a preset sound data set based on the target environment acoustic tag.

[0209] The acoustic modulation module 50 is used to adjust the amplitude characteristics of the target environmental acoustic information so that the amplitude characteristics of the target environmental acoustic information match the amplitude characteristics of the original environmental acoustic information.

[0210] The speech synthesis module 60 is used to synthesize the original speech content and the adjusted target environmental acoustic information into analog speech data.

[0211] In one embodiment, the sound source separation module 10 is specifically used for:

[0212] Acquire mixed speech data using a multi-channel audio acquisition device;

[0213] The mixed speech data is subjected to noise suppression preprocessing to generate noise-reduced speech data;

[0214] The denoised speech data is input into a pre-trained sound source separation model, which includes a parallel processing branch for human voice extraction and an environmental acoustic feature extraction branch.

[0215] The original speech content is output through the human voice extraction branch, and the original environmental acoustic information is output through the environmental acoustic feature extraction branch.

[0216] Based on the acoustic feature distribution of the target geographical environment, the frequency response characteristics of the original environmental acoustic information are adjusted by frequency band energy matching.

[0217] In one embodiment, the speech recognition module 20 is specifically used for:

[0218] Perform speech activity detection on the original speech content and divide the original speech content into multiple speech segments;

[0219] Each speech segment is input into a pre-trained multi-dialect speech recognition model to generate corresponding intermediate text data.

[0220] Dialect consistency correction is performed on the intermediate text data, and dialect vocabulary in the intermediate text data is adjusted based on the target geographic location information to generate dialect-corrected text;

[0221] Based on the temporal characteristics of the original speech content, timestamps are added to the dialect-corrected text to generate text data with temporal tags;

[0222] Aggregate all text data with time-series tags to generate the first text information.

[0223] In one embodiment, the environmental label recognition module 30 is specifically used for:

[0224] The original environmental acoustic information is spectrally segmented to generate multiple acoustic segments;

[0225] Perform multi-level classification operations on each acoustic segment to generate a preliminary set of environmental labels;

[0226] Based on the label types in the preliminary environmental label set, select the corresponding sub-recognition model;

[0227] Acoustic features are extracted from the acoustic segment to generate a multidimensional feature vector, and the multidimensional feature vector is input into the corresponding sub-recognition model to generate fine-grained classification results;

[0228] Merge all fine-grained classification results to generate refined environment labels;

[0229] The refined environmental labels are correlated with the acoustic feature mapping table in the geographic knowledge base to filter out candidate labels that match the target geographic location information.

[0230] Extract geographical keywords from the first text information;

[0231] Construct structured input data that includes candidate tags, geographic keywords, and target geographic location information;

[0232] A pre-trained geographic association language model is invoked to generate a label replacement strategy based on the structured input data and geographic knowledge base;

[0233] According to the label replacement strategy, conflicting items in the candidate labels are deleted or replaced to generate a corrected set of candidate labels as the target environment acoustic labels.

[0234] In one embodiment, the acoustic retrieval and matching module 40 is specifically used for:

[0235] The target environment acoustic tags are matched with acoustic tags in a preset sound data set to generate a candidate acoustic segment set;

[0236] Extract the temporal energy distribution characteristics and spectral characteristics of the original environmental acoustic information;

[0237] For each acoustic segment in the candidate acoustic segment set, perform an acoustic parameter matching operation to filter out acoustic segments whose temporal energy distribution characteristics and spectral characteristics are more than a preset similarity threshold with the original environmental acoustic information.

[0238] Based on the temporal energy distribution characteristics of the original environmental acoustic information, the selected acoustic segments are subjected to amplitude modulation processing to generate modulated acoustic segments.

[0239] The modulated acoustic segments are arranged and combined according to the temporal characteristics of the original environmental acoustic information to generate a continuous acoustic information stream;

[0240] The continuous acoustic information stream is subjected to acoustic feature consistency verification. If the verification is successful, the continuous acoustic information stream is used as the target environmental acoustic information.

[0241] In one embodiment, the acoustic modulation module 50 is specifically used for:

[0242] The target environmental acoustic information and the original environmental acoustic information are processed by framing to generate multiple target short-time acoustic frames and multiple original short-time acoustic frames.

[0243] Extract the target temporal amplitude features of each target short-time acoustic frame, and extract the original temporal amplitude features of each original short-time acoustic frame;

[0244] Based on the original temporal amplitude characteristics and the target temporal amplitude characteristics, the gain adjustment coefficient for each target short-time acoustic frame is determined;

[0245] Each target short-time acoustic frame is processed by applying the corresponding gain adjustment coefficient to generate an amplitude-matched target short-time acoustic frame.

[0246] The target short-time acoustic frame with amplitude matching is compared with the corresponding original short-time acoustic frame in the time domain amplitude feature.

[0247] If the difference between the temporal amplitude features of the target short-time acoustic frame that is amplitude matched and the temporal amplitude features of the corresponding original short-time acoustic frame is lower than a preset difference threshold, then the target short-time acoustic frame that is amplitude matched is marked as a valid frame.

[0248] All valid frames are overlapped and summed to generate acoustic information of the target environment.

[0249] In one embodiment, the speech synthesis module 60 is specifically used for:

[0250] The original speech content and the adjusted target environmental acoustic information are processed into frames to generate speech short frames and environmental short frames respectively.

[0251] Based on the temporal characteristics of the original speech content, the speech short frames and environment short frames are time-axis aligned to generate synchronized speech frames and synchronized environment frames.

[0252] Based on the temporal energy distribution characteristics of the synchronized speech frame and the current energy level of the synchronized environment frame, the gain compensation coefficient of the synchronized environment frame is determined.

[0253] The synchronization environment frame is processed based on the gain compensation coefficient to generate a gain-balanced environment frame.

[0254] The synchronized voice frame is superimposed with the gain-balanced environmental frame in the frequency domain to generate a hybrid frequency domain frame.

[0255] Perform an inverse Fourier transform on the hybrid frequency domain frame to generate a hybrid time domain frame;

[0256] The mixed time-domain frames are overlapped and added together to generate a continuous time-domain signal;

[0257] Adaptive noise suppression processing is applied to the continuous time-domain signal to generate analog speech data.

[0258] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side method for generating simulated environmental acoustic speech.

[0259] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the functions or steps of a user-side method for generating simulated ambient acoustic speech.

[0260] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0261] Acquire mixed speech data and perform a separation operation on the mixed speech data to generate original speech content and original environmental acoustic information;

[0262] The original speech content is converted into first text information;

[0263] The original environmental acoustic label is identified from the original environmental acoustic information, and the target environmental acoustic label is determined by the intelligent processing module based on the original environmental acoustic label, the first text information, and the target geographical location information.

[0264] Based on the target environment acoustic tag, extract the corresponding target environment acoustic information from a preset sound data set;

[0265] Adjust the amplitude characteristics of the target environmental acoustic information to match the amplitude characteristics of the original environmental acoustic information;

[0266] The original speech content is combined with the adjusted target environmental acoustic information to form analog speech data.

[0267] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0268] Acquire mixed speech data and perform a separation operation on the mixed speech data to generate original speech content and original environmental acoustic information;

[0269] The original speech content is converted into first text information;

[0270] The original environmental acoustic label is identified from the original environmental acoustic information, and the target environmental acoustic label is determined by the intelligent processing module based on the original environmental acoustic label, the first text information, and the target geographical location information.

[0271] Based on the target environment acoustic tag, extract the corresponding target environment acoustic information from a preset sound data set;

[0272] Adjust the amplitude characteristics of the target environmental acoustic information to match the amplitude characteristics of the original environmental acoustic information;

[0273] The original speech content is combined with the adjusted target environmental acoustic information to form analog speech data.

[0274] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0275] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0276] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0277] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for generating environmental acoustic simulation speech, characterized in that, Includes the following steps: Acquire mixed speech data and perform a separation operation on the mixed speech data to generate original speech content and original environmental acoustic information; The original speech content is converted into first text information; The original environmental acoustic label is identified from the original environmental acoustic information, and the target environmental acoustic label is determined by the intelligent processing module based on the original environmental acoustic label, the first text information, and the target geographical location information. Based on the target environment acoustic tag, extract the corresponding target environment acoustic information from a preset sound data set; Adjust the amplitude characteristics of the target environmental acoustic information to match the amplitude characteristics of the original environmental acoustic information; The original speech content is combined with the adjusted target environmental acoustic information to form analog speech data.

2. The environmental acoustic simulation speech generation method as described in claim 1, characterized in that, Acquire mixed speech data and perform a separation operation on the mixed speech data to generate original speech content and original environmental acoustic information, including: Acquire mixed speech data using a multi-channel audio acquisition device; The mixed speech data is subjected to noise suppression preprocessing to generate noise-reduced speech data; The denoised speech data is input into a pre-trained sound source separation model, which includes a parallel processing branch for human voice extraction and an environmental acoustic feature extraction branch. The original speech content is output through the human voice extraction branch, and the original environmental acoustic information is output through the environmental acoustic feature extraction branch. Based on the acoustic feature distribution of the target geographical environment, the frequency response characteristics of the original environmental acoustic information are adjusted by frequency band energy matching.

3. The environmental acoustic simulation speech generation method as described in claim 1, characterized in that, Converting the original speech content into first text information includes: Perform speech activity detection on the original speech content and divide the original speech content into multiple speech segments; Each speech segment is input into a pre-trained multi-dialect speech recognition model to generate corresponding intermediate text data. Dialect consistency correction is performed on the intermediate text data, and dialect vocabulary in the intermediate text data is adjusted based on the target geographic location information to generate dialect-corrected text; Based on the temporal characteristics of the original speech content, timestamps are added to the dialect-corrected text to generate text data with temporal tags; Aggregate all text data with time-series tags to generate the first text information.

4. The environmental acoustic simulation speech generation method as described in claim 1, characterized in that, Identifying the original environmental acoustic label from the original environmental acoustic information, and determining the target environmental acoustic label based on the original environmental acoustic label, the first text information, and the target geographical location information using an intelligent processing module, including: The original environmental acoustic information is spectrally segmented to generate multiple acoustic segments; Perform multi-level classification operations on each acoustic segment to generate a preliminary set of environmental labels; Based on the label types in the preliminary environmental label set, select the corresponding sub-recognition model; Acoustic features are extracted from the acoustic segment to generate a multidimensional feature vector, and the multidimensional feature vector is input into the corresponding sub-recognition model to generate fine-grained classification results; Merge all fine-grained classification results to generate refined environment labels; The refined environmental labels are correlated with the acoustic feature mapping table in the geographic knowledge base to filter out candidate labels that match the target geographic location information. Extract geographical keywords from the first text information; Construct structured input data that includes candidate tags, geographic keywords, and target geographic location information; A pre-trained geographic association language model is invoked to generate a label replacement strategy based on the structured input data and geographic knowledge base; According to the label replacement strategy, conflicting items in the candidate labels are deleted or replaced to generate a corrected set of candidate labels as the target environment acoustic labels.

5. The environmental acoustic simulation speech generation method as described in claim 1, characterized in that, Based on the target environment acoustic tag, extract the corresponding target environment acoustic information from a preset sound data set, including: The target environment acoustic tags are matched with acoustic tags in a preset sound data set to generate a candidate acoustic segment set; Extract the temporal energy distribution characteristics and spectral characteristics of the original environmental acoustic information; For each acoustic segment in the candidate acoustic segment set, perform an acoustic parameter matching operation to filter out acoustic segments whose temporal energy distribution characteristics and spectral characteristics are more similar to the original environmental acoustic information than a preset similarity threshold. Based on the temporal energy distribution characteristics of the original environmental acoustic information, the selected acoustic segments are subjected to amplitude modulation processing to generate modulated acoustic segments. The modulated acoustic segments are arranged and combined according to the temporal characteristics of the original environmental acoustic information to generate a continuous acoustic information stream; The continuous acoustic information stream is subjected to acoustic feature consistency verification. If the verification is successful, the continuous acoustic information stream is used as the target environmental acoustic information.

6. The environmental acoustic simulation speech generation method as described in claim 1, characterized in that, Adjusting the amplitude characteristics of the target environmental acoustic information to match the amplitude characteristics of the original environmental acoustic information includes: The target environmental acoustic information and the original environmental acoustic information are processed by framing to generate multiple target short-time acoustic frames and multiple original short-time acoustic frames. Extract the target temporal amplitude features of each target short-time acoustic frame, and extract the original temporal amplitude features of each original short-time acoustic frame; Based on the original temporal amplitude characteristics and the target temporal amplitude characteristics, the gain adjustment coefficient for each target short-time acoustic frame is determined; Each target short-time acoustic frame is processed by applying the corresponding gain adjustment coefficient to generate an amplitude-matched target short-time acoustic frame. The target short-time acoustic frame with amplitude matching is compared with the corresponding original short-time acoustic frame in the time domain amplitude feature. If the difference between the temporal amplitude features of the target short-time acoustic frame that is amplitude matched and the temporal amplitude features of the corresponding original short-time acoustic frame is lower than a preset difference threshold, then the target short-time acoustic frame that is amplitude matched is marked as a valid frame. All valid frames are overlapped and summed to generate acoustic information of the target environment.

7. The environmental acoustic simulation speech generation method as described in claim 1, characterized in that, The original speech content is combined with the adjusted target environmental acoustic information to form analog speech data, including: The original speech content and the adjusted target environmental acoustic information are processed into frames to generate speech short frames and environmental short frames respectively. Based on the temporal characteristics of the original speech content, the speech short frames and environment short frames are time-axis aligned to generate synchronized speech frames and synchronized environment frames. Based on the temporal energy distribution characteristics of the synchronized speech frame and the current energy level of the synchronized environment frame, the gain compensation coefficient of the synchronized environment frame is determined. The synchronization environment frame is processed based on the gain compensation coefficient to generate a gain-balanced environment frame. The synchronized voice frame is superimposed with the gain-balanced environmental frame in the frequency domain to generate a hybrid frequency domain frame. Perform an inverse Fourier transform on the hybrid frequency domain frame to generate a hybrid time domain frame; The mixed time-domain frames are overlapped and added together to generate a continuous time-domain signal; Adaptive noise suppression processing is applied to the continuous time-domain signal to generate analog speech data.

8. An environmental acoustic simulation speech generation device, characterized in that, The environmental acoustic simulation speech generation device includes: The sound source separation module is used to acquire mixed speech data and perform separation operations on the mixed speech data to generate original speech content and original environmental acoustic information; The speech recognition module is used to convert the original speech content into first text information; An environmental label recognition module is used to identify original environmental acoustic labels from the original environmental acoustic information, and to determine the target environmental acoustic label through an intelligent processing module based on the original environmental acoustic labels, the first text information, and the target geographical location information. The acoustic retrieval and matching module is used to extract the corresponding target environment acoustic information from a preset sound data set based on the target environment acoustic tag. An acoustic modulation module is used to adjust the amplitude characteristics of the target environmental acoustic information so that the amplitude characteristics of the target environmental acoustic information match the amplitude characteristics of the original environmental acoustic information. The speech synthesis module is used to synthesize the original speech content and the adjusted target environmental acoustic information into analog speech data.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and an environmental acoustic simulation speech generation program stored in the memory and executable on the processor, wherein the environmental acoustic simulation speech generation program, when executed by the processor, implements the steps of the environmental acoustic simulation speech generation method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores an environmental acoustic simulation speech generation program, which, when executed by a processor, implements the steps of the environmental acoustic simulation speech generation method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Feature information mining method and device and electronic equipment

    CN112614484A

  • Application scene simulation method and device, storage medium and electronic device

    CN113971338A