Ambient sound feature-based authenticity analysis method, apparatus and device, and medium

By performing speech separation and feature extraction on the original speech data and combining knowledge information with user statement information for comprehensive analysis, the problem of insufficient understanding of environmental sounds in the existing technology is solved, and accurate judgment of the authenticity of the speaker's statement is achieved, thereby improving the accuracy and robustness of the judgment.

CN120612960AActive Publication Date: 2025-09-09PING AN TECH (BEIJING) CO LTD

Patent Information

Application Number
CN202510844794.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-09
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

Existing technologies lack a deep structural understanding of environmental sounds and cross-modal correlation analysis capabilities in location verification, resulting in the inability to effectively identify the contradiction between the acoustic environment and semantic content in forged scenes, weak anti-spoofing capabilities and low verification accuracy.

Method used

By obtaining the original voice data for voice separation, generating environmental sound data and pure voice data, extracting the environmental sound features and retrieving related knowledge information, performing voice content recognition on the pure voice data, generating conversation text data, and inputting this data and user declaration information into the analysis model for comprehensive analysis.

Benefits of technology

It achieves accurate judgment in multi-source heterogeneous environmental information, improves the accuracy and robustness of discrimination in complex scenarios, and enhances the ability to judge the authenticity of the speaker's statement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612960A_ABST
    Figure CN120612960A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical treatment and health and the like, and discloses an authenticity analysis method, device and equipment based on environmental sound characteristics and a medium. Performing voice separation processing on the original voice data to generate environment voice data and pure voice data, analyzing the environment voice data and retrieving knowledge information associated with the environment voice data, and identifying the content of the pure voice data to generate dialogue text data, and inputting the environment sound features, the knowledge information, the dialogue text data and the user declaration information into an analysis model, and outputting a authenticity analysis result. According to the method, the environment sound data and the voice content are separated, the available features of the environment sound data and the voice content are extracted respectively, and the background knowledge and the user declaration information are combined to perform fusion reasoning in the unified analysis model, so that the accuracy of authenticity judgment and the adaptability to complex scenes can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing technology, and in particular to a method, device, equipment and storage medium for authenticity analysis based on environmental sound features. Background Art

[0002] Although some existing location verification technologies have attempted to incorporate environmental information as auxiliary features, these methods mostly remain at the shallow stage of feature extraction and matching, lacking the ability to deeply structure environmental data and analyze cross-modal associations. In practical applications, environmental sounds are often highly diverse and heterogeneous, such as birdsong in the background, human dialects, and traffic sirens. These sounds can be contradictory or ambiguous, making it difficult for existing technologies to cope with this complexity. In particular, when there are multiple conflicting environmental signals, existing methods lack a systematic conflict detection and interpretation mechanism, and are generally unable to make accurate judgments.

[0003] In the fintech sector, scenarios such as remote account opening, online loan applications, and identity verification are becoming increasingly common. After users submit their declared information via voice channels, platforms urgently need to determine whether they are actually in the claimed location to prevent fraud caused by location spoofing. However, current mainstream technologies still rely primarily on basic features such as IP addresses, device location, or voice signals. They lack effective utilization of ambient sound during calls, and are particularly unable to combine voice content with background information to make reliable inferences, resulting in limited fraud detection capabilities.

[0004] In healthcare, for scenarios like remote consultations, voice recording of medical histories, and online emergency dispatch, if the location information reported by the patient or caller doesn't match the actual environment, it can lead to dispatch delays or resource mismatches. Existing technologies primarily rely on voice recognition and general positioning methods, but lack an effective mechanism that can leverage ambient acoustic signals to aid in determining a patient's true location. This inability to accurately determine a patient's location is particularly problematic in emergency situations where GPS or other high-precision positioning methods are lacking.

[0005] In summary, existing technologies have significant deficiencies in multi-source environmental information fusion analysis, conflict information resolution, and spatial consistency inference, making it difficult to meet the intelligent verification needs of application scenarios such as financial risk control and medical dispatch that have high requirements for location authenticity. Summary of the Invention

[0006] The main purpose of the present invention is to provide an authenticity analysis method, device, equipment and storage medium based on environmental sound characteristics, aiming to solve the technical problems that the existing verification method relies on a single data source and lacks collaborative analysis of environmental information and semantic logic, resulting in the inability to effectively identify the contradiction between the acoustic environment and semantic content in the counterfeit scene, weak anti-deception ability and low verification accuracy.

[0007] To achieve the above object, the present invention provides a method for authenticity analysis based on environmental sound features, comprising:

[0008] Obtaining raw voice data including environmental sounds;

[0009] Performing voice separation processing on the original voice data to generate ambient sound data and pure voice data;

[0010] Analyzing the environmental sound data, extracting environmental sound features, and retrieving knowledge information associated with the environmental sound features;

[0011] Performing voice content recognition on the clean voice data to generate dialogue text data;

[0012] The environmental sound features, knowledge information, dialogue text data and user declaration information are input into an analysis model, and an authenticity analysis result is output.

[0013] Furthermore, to achieve the above-mentioned purpose, the present invention provides an authenticity analysis device based on environmental sound characteristics, comprising:

[0014] The voice acquisition module is used to obtain the original voice data including the environmental sound;

[0015] A speech separation module is used to perform speech separation processing on the original speech data to generate environmental sound data and pure speech data;

[0016] An environmental perception analysis module, configured to analyze the environmental sound data, extract environmental sound features, and retrieve knowledge information associated with the environmental sound features;

[0017] A speech recognition module, configured to perform speech content recognition on the clean speech data and generate dialogue text data;

[0018] The authenticity analysis module is used to input the environmental sound characteristics, knowledge information, dialogue text data and user declaration information into the analysis model and output the authenticity analysis results.

[0019] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and an authenticity analysis program based on environmental sound features stored in the memory and runnable on the processor. When the authenticity analysis program based on environmental sound features is executed by the processor, the steps of the authenticity analysis method based on environmental sound features as described above are implemented.

[0020] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which is stored an authenticity analysis program based on environmental sound features, and when the authenticity analysis program based on environmental sound features is executed by a processor, the steps of the authenticity analysis method based on environmental sound features as described above are implemented.

[0021] Beneficial effects: The present invention relates to the field of speech processing technology and can be applied to business scenarios such as financial technology and medical health. It discloses a method, device, equipment and medium for authenticity analysis based on environmental sound features, including: obtaining original speech data containing environmental sounds; performing speech separation on the original speech data to generate environmental sound data and pure speech data; analyzing the environmental sound data to generate environmental sound features, and retrieving knowledge information associated with the environmental sound features; performing speech content recognition on the pure speech data to generate dialogue text data; inputting the environmental sound features, knowledge information, dialogue text data and user declaration information into the analysis model to generate an authenticity judgment result. The present invention separates and processes the environmental sounds and pure speech in the original speech data, extracts the environmental features and semantic content respectively, and combines the knowledge information and user declaration information for unified modeling and analysis. In the presence of multi-source heterogeneous environmental information, it can perform reasoning based on semantic, geographical and temporal dimensions, thereby achieving accurate judgment of the authenticity of the speaker's statement and improving the accuracy and robustness of discrimination in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0023] Figure 1 A schematic diagram of an application environment of an authenticity analysis method based on environmental sound features according to an embodiment of the present invention;

[0024] Figure 2 1. It is a flow chart of an embodiment of a method for authenticity analysis based on environmental sound features of the present invention;

[0025] Figure 3 Schematic diagram of functional modules of a preferred embodiment of the authenticity analysis device based on environmental sound features of the present invention;

[0026] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0027] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0029] The authenticity analysis method based on environmental sound characteristics provided by the embodiment of the present invention can be applied in Figure 1 In an application environment, a user terminal communicates with a server terminal via a network. The server terminal can obtain raw speech data containing ambient sound from the user terminal; perform speech separation on the raw speech data to generate ambient sound data and clean speech data; analyze the ambient sound data to generate ambient sound features and retrieve knowledge information associated with the ambient sound features; perform speech content recognition on the clean speech data to generate conversation text data; and input the ambient sound features, knowledge information, conversation text data, and user statement information into an analysis model to generate an authenticity judgment result. The present invention separates the ambient sound and clean speech in the raw speech data, extracts the environmental features and semantic content respectively, and combines the knowledge information with the user statement information for unified modeling and analysis. This enables reasoning based on semantic, geographical, and temporal dimensions in the presence of multi-source heterogeneous environmental information, thereby accurately judging the authenticity of the speaker's statement and improving the accuracy and robustness of judgment in complex scenarios. The user terminal can include, but is not limited to, various personal computers, laptops, smartphones, tablet computers, and portable wearable devices. The server terminal can be implemented as a standalone server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0030] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of the authenticity analysis method based on environmental sound features provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0031] like Figure 2 As shown, the authenticity analysis method based on environmental sound features proposed in the present invention includes the following steps:

[0032] S10, obtaining original voice data including environmental sounds;

[0033] In this embodiment, acquiring raw speech data containing ambient sounds essentially involves capturing an audio signal that exhibits a mixture of speech and ambient background. The goal of this process isn't simply to capture the content of the speech, but rather to intentionally preserve information from naturally occurring auxiliary sound sources in the environment. These ambient sounds can include background conversations, distant horns, traffic noise, animal calls common in a specific area, and the sounds of construction equipment. The inclusion of these sounds isn't an error or redundant information in the noise filtering process; rather, they are actively incorporated into the data and form a crucial component of the data structure for subsequent analysis.

[0034] When implementing this operation, attention should be paid to the broadband response coverage of the collected signal to ensure that different characteristic sounds, such as high-frequency bands such as alarms and bird calls, mid-frequency bands such as human voices, and low-frequency bands such as mechanical vibrations, are effectively recorded. In addition, the acquisition of this audio data should preserve temporal continuity and intensity dynamics, and waveform features that may carry indicative meaning in the audio should not be smoothed or excessively noise-reduced.

[0035] Because environmental sounds have regional, temporal, and probabilistic characteristics, the collection process should avoid interfering with or changing the natural sound field as much as possible. For example, spaces with strong interference sources or background music should be avoided to avoid masking the spatial distribution structure of natural environmental sounds. In some scenarios, it is also possible to optimize the collection integrity and clarity of environmental sounds by setting physical structure parameters (such as recording distance or reception direction), thereby providing more auxiliary information for subsequent semantic recognition and regional feature judgment.

[0036] The acquired audio data must be raw, meaning it is waveform data that is directly saved without going through semantic extraction, sound source separation, or content filtering. It not only carries verbal content but also provides temporal continuity, sound field integrity, and the foundation for expressing unprocessed information, making it an indispensable raw input for multi-level analysis.

[0037] Sound acquisition logic can be configured to enhance the complete recording of ambient sounds. For example, data acquisition can be prioritized in open or semi-open spaces to increase the probability of natural sound inclusion. Alternatively, audio dynamic range control can be configured to store data only when the overall sound pressure level is within a threshold range, thus avoiding inappropriate data in extremely quiet or noisy conditions. Time tags, location tags, or environmental classification tags can also be attached before or after acquisition, allowing the audio data to incorporate richer contextual information when it is subsequently used for speech separation, content recognition, and authenticity inference. For different environmental contexts, the acquisition duration can be adjusted to ensure a recordable time window for typical background sound events, such as extending the acquisition duration to cover periodic sound signals (such as scheduled broadcasts or fixed-point prompts). In some cases, a concurrent acquisition mechanism can be used to simultaneously record multiple sound source directions in the same scene, expanding the sound source range and sound field structure depth of the environmental characteristics.

[0038] Example: In healthcare, when a user submits a medical condition description via a remote terminal, the system simultaneously records the ambient sound while receiving the voice input. Information naturally included in the audio, such as elevator voice prompts, background medical equipment prompts, and medical staff conversations, can serve as key evidence to determine whether the user is in a hospital environment.

[0039] In fintech scenarios, when users upload voice data for remote account opening or identity verification, the system can analyze the raw audio for regionally specific sound signatures (such as specific dialects or trading floor announcements) to help determine whether their statements are consistent with the actual environment. This approach eliminates the need for ambient sound as a distraction in the recognition task and instead becomes a crucial component in building a trustworthy analytics system.

[0040] This embodiment preserves the ambient sound components in the original speech data, not only providing context for subsequent speech recognition and semantic understanding, but also providing auxiliary judgment conditions for more complex area recognition, behavior verification, or scene matching. This data acquisition method preserves physical parameters such as temporal structure, spectral composition, and sound field structure, allowing downstream analysis modules to fully explore the non-semantic dimension information in the audio and perform tasks such as spatial positioning and authenticity judgment based on these structured or unstructured signals, thereby improving the generalization and anomaly recognition capabilities of the entire process in complex situations.

[0041] S20, performing voice separation processing on the original voice data to generate ambient sound data and pure voice data;

[0042] In this embodiment, speech separation processing is performed on raw speech data, with the goal of effectively distinguishing and reconstructing portions of the recorded composite audio signal that possess distinct sound source characteristics. Raw speech data typically consists of multiple sounds, including the primary speech target (i.e., the speaker's voice) and accompanying environmental sound components (such as traffic, natural sounds, or other conversations). These sounds exhibit significant differences in spectrum, energy, and temporal structure, making them amenable to separation using computational models.

[0043] The processing method typically includes four stages: sound source modeling, feature extraction, mask generation, and signal reconstruction. First, the raw speech data is input into the acoustic modeling framework. A neural network model with speech separation capabilities is used in an unsupervised or semi-supervised manner to learn the internal sound source representation of the mixed signal. This representation process uses the time-frequency graph of the raw speech data as input and, through a hierarchical convolutional structure, extracts a time-frequency domain representation that includes features such as speech dynamics, formant structure, and background noise texture.

[0044] After extracting the representations, the sound source features that may belong to the background environment can be clustered, labeled, or enhanced to construct a sound source mask matrix, which indicates the distribution of signals belonging to the main speaker or the ambient sound in specific frequency bands. This mask matrix can be in binary form to forcibly block signals in non-target frequency bands, or in probabilistic form to indicate the degree of mixing and signal ratio.

[0045] Next, the frequency domain representation of the original speech data is subjected to a frequency-band-by-frequency operation on the generated mask matrix to reconstruct two audio signals representing different sound source channels. Finally, these frequency domain signals are subjected to an inverse Fourier transform (IFT) or an inverse short-time Fourier transform (ISFT) to generate time domain audio signals, which are output as ambient sound data and clean speech data, respectively.

[0046] During the separation process, the integrity of speech boundaries and the continuity of environmental structures should be maintained to prevent speech fragmentation or excessive background cleaning caused by masking operations. The separation model should also be able to generalize to unknown environmental sounds to adapt to speech input data from different devices, scenarios, or regions.

[0047] Speech separation can be achieved using structured neural network models, such as a separation model based on a convolutional time domain network (e.g., Conv-TasNet), which learns the mapping relationship between the mixed speech and the separation target through end-to-end training. Alternatively, time-frequency domain methods based on mask learning can be used, such as performing a short-time Fourier transform on the original speech signal and inputting it into a deep neural network (e.g., Dual-Path RNN) to extract the soft mask matrix corresponding to the sound source, followed by an inverse transform to reconstruct the separated audio track. Attention mechanisms can also be combined with sound source representation clustering strategies to model the changing background sound field in speech, further improving the separation model's ability to discriminate complex background sounds.

[0048] In model input design, background silence segments, spectral comparison segments, or known background samples can be added in sync with data acquisition to help the model establish the characteristic boundaries of the sound source. At the output stage, an energy constraint balancing strategy can be applied to the separated audio tracks to prevent weak ambient sounds or marginal speech from being completely cut off.

[0049] Example: In healthcare scenarios, remote voice communications between doctors and patients are often accompanied by background noise, such as ward announcements, instrument prompts, or other patient conversations. Through speech separation processing, the doctor-patient conversation content can be clearly extracted for semantic analysis, while the ambient sound is retained as an auxiliary clue to determine whether the communication actually occurred in a medical environment.

[0050] In FinTech businesses, the audio uploaded by customers during remote identity verification is often accompanied by counter noise, crowd noise, or background broadcasts. Voice separation preserves the high-quality customer voice for system recognition, while independently separating out ambient sound to serve as auxiliary information to verify that the voice was collected at the specific financial service site. By separating and processing voice signals, the system's ability to interpret complex input data and its decision-making reliability are further enhanced.

[0051] This embodiment effectively separates the ambient sounds interwoven within the raw speech data from the pure speech signal. This not only provides clearer, less-interfered primary speech input for subsequent speech recognition, improving the accuracy of text generation, but also allows the separated ambient sound data to be independently used for sound source identification, regional feature analysis, or spatial background modeling, enhancing the overall analysis process's ability to understand the speaking environment. This processing mechanism not only decouples semantics from context but also expands the depth and precision of multi-dimensional utilization of input audio content.

[0052] S30, analyzing the environmental sound data, extracting environmental sound features, and retrieving knowledge information associated with the environmental sound features;

[0053] In this embodiment, the goal of analyzing ambient sound data is to extract features with geographic, semantic, or scene-recognition value, and to access associated external knowledge resources based on these features. Ambient sound data typically contains audio patterns associated with specific spatiotemporal locations, such as bird calls, human dialects, vehicle horns, and mechanical operations. These patterns possess quantifiable acoustic structures that can be used to construct inference pathways between sounds and locations, species, or events.

[0054] First, acoustic event detection is performed on the ambient sound data. This involves identifying and segmenting acoustic segments with independent meaning in the temporal dimension. This process typically combines metrics such as short-term energy, spectral variation, and signal-to-noise ratio. Alternatively, a trained acoustic event detection model can be used to extract a set of candidate acoustic events. Each acoustic event includes elements such as start and end time, frequency structure, and energy parameters.

[0055] The acoustic events are then classified into multiple categories. Some of these acoustic events are biological acoustic events, and the corresponding bird, insect, or animal species can be identified through species classification models. The model is usually based on convolutional neural networks or transfer learning methods. During the training process, global or regional bird song datasets are introduced to identify and output specific species labels, regional characteristics, frequency of occurrence, and other indicators. Another part of the acoustic events are non-biological acoustic events, such as ambulance sirens, subway passing, campus broadcasts, etc., which can be classified and identified through voiceprint template matching, MFCC feature clustering, time series alignment, and other methods.

[0056] The above classification results are integrated into a unified set of environmental sound features. This set retains information such as sound source type, classification confidence, time location, etc. in the data structure, and can be used as structured input for downstream analysis. Based on the feature set, external retrieval is further performed through knowledge graphs or embedded indexing systems. Knowledge information usually contains established entity relationships such as sound source-location, sound source-time, and sound source-ecological chain. For example, if the detected bird song signal matches "white-crowned bulbul", relevant knowledge about its occurrence in urban green belts, specific latitude ranges, and high activity in the early morning can be retrieved; if "ambulance sirens" are detected, corresponding information about its occurrence in urban hospital-dense areas, specific road types, time period patterns, etc. can be retrieved.

[0057] Ultimately, the generated environmental sound features and their associated knowledge information form the basis for understanding the current audio environment structure and provide interpretable support for subsequent comprehensive judgments.

[0058] Acoustic event recognition can be performed using a multi-model collaborative approach. For example, in the first stage, the YAMNet model is used to perform coarse-grained event detection on environmental sound data and extract preliminary acoustic segments. In the second stage, the bird classification model and the non-biological event soundprint matching model are respectively called to perform refined classification operations on the candidate segments. For the classification of bioacoustic events, based on bird sound resources with rich species annotations and geographical distribution characteristics, regional adaptation training can be performed on sound recognition models with pre-training capabilities. This allows the model to further enhance its ability to distinguish regional species distribution characteristics on the basis of recognizing different types of bird sounds in audio, thereby achieving more accurate bioacoustic event classification and feature extraction in different geographical contexts. In non-biological sound recognition, multiple MFCC template-based recognizers can also be constructed, targeting typical urban environments such as transportation, construction, and campuses.

[0059] Knowledge retrieval can rely on a locally constructed environmental sound source knowledge graph, implemented in RDF format or an embedded vector library. For each type of sound source, this graph indexes information such as temporal distribution, spatial distribution, and ecological association chains. Alternatively, a pre-trained language model can be used to input sound source labels and perform a retrieval-augmented generation (RAG) process. This process extracts trusted, relevant knowledge from literature and city databases, outputting structured spatial distribution data.

[0060] Example description: In a healthcare scenario, a remote diagnosis and treatment platform can use the hospital background broadcast sounds, pager prompt sounds, elevator arrival sounds, etc. that appear in the voice uploaded by patients to be identified as environmental sound features. Combined with the sound source distribution information of the hospital's specific environment in the knowledge base, it can assist in determining whether the voice collection actually occurred in a medical setting, thereby improving the credibility of remote medical interactions.

[0061] In financial business scenarios, when customers upload voice content through mobile terminals, the backend system can detect dialects, bank call system voice prompts, printer operating sounds, etc. appearing in the background. By comparing them with the environmental characteristics of known financial business outlets and combining the geographical layout knowledge of financial institutions, it can improve the ability to judge the authenticity of customer behavior locations, which helps prevent remote fraud and identity forgery risks.

[0062] This embodiment systematically analyzes ambient sound data to extract spatiotemporal and temporal tag information from audio clips. This approach, combined with structured knowledge resources, allows for in-depth analysis of the consistency between the environment and the speaker's statements. Compared to single-source sound detection or spectrum analysis alone, this processing approach can simultaneously detect and interpret multiple types of acoustic events, effectively improving understanding of the audio environment and enhancing the reliability of the decision-making basis of downstream judgment modules.

[0063] S40, performing voice content recognition on the clean voice data to generate dialogue text data;

[0064] In this embodiment, pure speech data is used as input and first subjected to frame-level segmentation and window function weighting to form a set of continuous but finite-length audio frames. This process can be accomplished using a sliding window of fixed duration, for example, setting the frame length to 25 milliseconds and the frame shift to 10 milliseconds. A window function is applied to each frame to reduce spectral leakage effects and improve signal localization in frequency domain analysis. Feature vectors reflecting acoustic properties are extracted from each frame. A commonly used representation is the Mel-frequency cepstral coefficient, which is obtained by performing a short-time Fourier transform, Mel filter bank processing, logarithmic compression, and then performing a discrete cosine transform on the speech signal. It can effectively simulate the distribution characteristics of the spectrum perceived by the human ear. These feature vectors are fed as input into the speech recognition model for modeling.

[0065] Speech recognition models can use a multi-layer acoustic modeling network to model the contextual relationships in acoustic feature sequences through temporal modeling structures (such as convolutional neural networks or structures with attention mechanisms). This model outputs a probability distribution sequence of phonemes or subword units, which can be used to generate preliminary text content through decoding algorithms such as optimal path search based on conditional random fields or connectionist temporal classification (CTC). To improve the readability and semantic integrity of the text content, this initial output needs to be input into a semantic layer language model for correction. The language model can capture syntactic structure, semantic associations, and contextual coherence, enabling automatic error correction, grammar completion, and format standardization, ultimately forming grammatically complete and clearly expressed conversational text data.

[0066] In implementation, a deep model based on an attention mechanism can be used to model clean speech data. For example, frame-level features can be fed into a bidirectional encoder network, and contextual information can be reweighted using positional encoding and an attention layer. Phoneme decoding can be combined with language priors and acoustic likelihood for improved robustness. Conversational text generation can utilize an end-to-end speech recognition architecture, embedding acoustic and language modeling into a unified network. Multi-task training can simultaneously optimize recognition accuracy and language consistency. A post-processing language model can also be used to correct the initially recognized text, using a context-based neural network generator to identify timing errors, improper word forms, and grammatical flaws.

[0067] When adapting to different languages ​​or scenarios with significant pronunciation differences, the model can be trained through transfer learning or parameter adjustment to accommodate a variety of speech input features. For example, fine-tuning can be performed on speech data with accent variations, specific dialects, or mixed languages, ensuring that the model maintains good recognition capabilities even when phoneme boundaries are blurred and syntactic structures vary.

[0068] Example: In a healthcare scenario, a user remotely describes their medical condition via a mobile device. This speech is often accompanied by ambient noise and interference from accompanying personnel. The system first removes background sound from the overall speech data, retaining only the user's subjective expression as pure speech data, and then performs speech content recognition on this data. This step, through sequence modeling and language structure restoration of the pure speech data, accurately generates conversational text from expressions mixed with accents or medical terminology. This allows the patient's symptom description to be standardized, providing high-quality structured input for doctors or consultation systems.

[0069] In FinTech scenarios, users submit claim statements or location verification information via voice. The system extracts clean voice data and then performs speech recognition on it. This step accurately identifies key elements in the user's pronunciation, such as time, location, and event sequence. This is particularly true in situations involving dialects, uneven speech speeds, or misplaced information delivery. Language modeling can restore semantic order and complete missing vocabulary, generating conversational text data for subsequent verification and analysis. This supports verification of the consistency between the claim and background information in complex scenarios.

[0070] This embodiment uses multi-stage processing of clean speech data, from low-level acoustic feature extraction to high-level semantic calibration for language model modification. This ensures a clean background while accurately extracting speech content, effectively avoiding recognition bias caused by environmental interference, and improving the accuracy and semantic consistency of conversational text. Combining this temporal modeling structure with a language model significantly enhances the adaptability and generalization capabilities of the speech recognition system in natural contexts.

[0071] S50, inputting the environmental sound features, knowledge information, conversation text data and user declaration information into an analysis model, and outputting an authenticity analysis result.

[0072] In this embodiment, the operation of integrating environmental sound features, knowledge information, conversation text data, and user statement information into the input analysis model is intended to complete deep joint modeling of multi-source data to achieve reasoning analysis of potential consistency or contradiction between information. Environmental sound features are generally composed of acoustic label information extracted in the previous step, such as a multidimensional vector set representing a specific biological species or non-biological event. These features are usually sensitive to geographical distribution and have temporal occurrence regularity. Knowledge information comes from structured data corresponding to the above-mentioned environmental features, which can be specifically expressed as the corresponding probability of a certain type of sound and a geographical location or time period, known occurrence areas or restrictions, etc. This information is mapped into a semantic map or knowledge vector through a pre-trained embedding method.

[0073] Conversational text data originates from the clean speech content recognition step and contains subjective user expressions. It is semantically represented through language understanding mechanisms, forming semantic vectors that reflect spatiotemporal orientation, behavioral descriptions, or event states. User declaration information, on the other hand, consists of subjective statements submitted by users related to task objectives. These statements may include geographic location, time point, or identity attributes. This information must be standardized and embedded into a structured and readable numerical representation. Vectorization is often achieved using a geographic coordinate embedder, location encoding network, or spatiotemporal indexing model.

[0074] The four types of data described above must be temporally aligned and spatially normalized before input to ensure consistency across scales and analysis intervals. This can include synchronizing event granularity using sliding window techniques, unifying spatial reference systems using geographic coordinate mapping or regional labels, or adjusting the correspondence between statements and sound events based on semantic annotation. The aligned data form a unified input set, where subvectors represent environmental features, knowledge semantics, text semantics, and statement coordinates, respectively.

[0075] After this collection is input into the analysis model, the multimodal attention mechanism first calculates the correlation matrix between different modalities, extracts suspicious conflicting or highly consistent feature areas, and then integrates them into a unified multimodal representation vector through a fusion module. During the reasoning phase, this vector evaluates whether there are semantic / geographic / temporal conflicts or whether the information exhibits a high degree of coordination and consistency through structured reasoning units containing hierarchical logical judgment paths. The final model outputs the authenticity analysis results, which are expressed as structured labels (such as true or false, credibility scores) and interpretable information supporting the logical reasons for the results.

[0076] In different implementations, the multimodal fusion method can be selected based on actual system resources and application scenarios. An attention-guided soft fusion strategy can be used to integrate information from different sources in a weighted manner, or a graph neural network structure can be constructed to construct a feature node graph and propagate information based on node relationships. The encoding method for user declaration information can also be adjusted as needed, such as using city label embedding for city-level judgments, using latitude and longitude encoding in fine-grained location scenarios and aligning with geographic nodes in the knowledge graph. The processing of conversation text data can introduce language context enhancement mechanisms, such as introducing a historical conversation window or an external timeline structure to assist in identifying the spatiotemporal attributes of the subject. The vectorization of environmental sound features can be based on a label set frequency histogram, an event duration vector, or an event confidence distribution map. Different encoding methods will affect the subsequent semantic alignment effect.

[0077] Example: In a fintech business, a user submitted an address via voice to assist with risk verification. The system captured background sound characteristics from the user's surroundings and extracted a bird call signature unique to Southeast Asia. The user also indicated during the call that they were located in Europe and submitted a location statement that indicated a European city. The system aligned these inputs in time and space, identifying a significant spatial discrepancy between the bird call source and the European location. Combined with semantic analysis, the system model outputted a low confidence authenticity analysis result, along with an explanation of the discrepancy.

[0078] In a healthcare scenario, a user in a remote consultation claims to be resting at home and simultaneously states, "It's been very noisy nearby, a fire truck passing by." The system analyzes the pure voice content and the background environment, identifies the repeated sounds of fire truck sirens, and retrieves from knowledge information that there have been no fire calls to the user's stated location in the past three hours. The multimodal fusion analysis system determines a discrepancy between the user's voice description and the actual statement, outputs an authenticity analysis result of inconsistency, and identifies the cause of the temporal anomaly. This mechanism provides support for determining environmental consistency and user status authenticity during remote consultations.

[0079] This embodiment can conduct a comprehensive analysis of the temporal and spatial or semantic conflicts implied in different modalities by jointly modeling environmental sound features, knowledge information, conversation text data, and user declaration information. This can not only identify explicit inconsistent behaviors (such as the contradiction between the declared area and the species distribution), but also capture fine-grained reasoning relationships between cross-modalities (such as the inconsistency between the declaration time and the time period of the ambulance sound). This mechanism enables the system to derive complex relationships from multi-dimensional clues, significantly improving the accuracy and robustness of authenticity judgments, and is particularly suitable for processing information environments with ambiguous, contradictory, and complex semantic interweaving in the real world.

[0080] The present invention relates to the field of speech processing technology and can be applied to business scenarios such as financial technology and healthcare. A method, apparatus, device, and medium for authenticity analysis based on environmental sound features are disclosed, including: obtaining original speech data containing environmental sound; performing speech separation on the original speech data to generate environmental sound data and clean speech data; analyzing the environmental sound data to generate environmental sound features and retrieving knowledge information associated with the environmental sound features; performing speech content recognition on the clean speech data to generate conversation text data; and inputting the environmental sound features, knowledge information, conversation text data, and user statement information into an analysis model to generate an authenticity judgment result. The present invention separates the environmental sound and clean speech in the original speech data, extracts the environmental features and semantic content respectively, and combines the knowledge information with the user statement information for unified modeling and analysis. In the presence of multi-source heterogeneous environmental information, the present invention can perform reasoning based on semantic, geographical, and temporal dimensions, thereby accurately judging the authenticity of the speaker's statement and improving the accuracy and robustness of judgment in complex scenarios.

[0081] In one embodiment, the above step S10 includes:

[0082] S101, collecting an original voice signal containing ambient sound through a microphone module of a terminal device;

[0083] S102, performing ambient sound audio segment selective gain processing on the original speech signal to generate original speech data containing enhanced ambient sound;

[0084] S103, converting the original voice data into a standardized audio format file;

[0085] S104: Store the standardized audio format file in a storage area, and record the acquisition time and acquisition device identifier of the original voice data.

[0086] In this embodiment, the acquisition of raw speech data containing ambient sounds is intended to construct a realistic, multi-dimensional acoustic input source, ensuring a sufficient basis for expressing environmental elements in subsequent speech analysis. This process can be broken down into multiple, sequentially linked processing steps.

[0087] First, capturing the original voice signal through the terminal device is a prerequisite for obtaining sound information. The sound signal should have wide coverage, realistic noise structure, and obvious environmental differences. Collection does not pre-set the collection location, device form, or scene restrictions. The only requirement is that the terminal device has audio collection capabilities and can simultaneously capture the human voice and environmental background sounds. During the collection process, the original voice signal contains the user's voice content and the surrounding natural environmental sounds such as traffic noise, crowd conversations, bird calls, and mechanical sounds, reflecting the mixed nature of voice data in real-world scenarios.

[0088] To improve the recognizability of environmental elements in subsequent analysis, the above-mentioned signals need to be subjected to frequency-band selective gain processing. Gain processing refers to the implementation of amplitude enhancement of the frequency region containing environmental sound characteristics while preserving the overall speech structure. Environmental sounds are usually concentrated in the non-speech main frequency band area, such as mechanical sounds in the low frequency band, birdsong or ambulance sirens in the high frequency band. The processing method can use a dynamic spectrum separation algorithm to perform weighted adjustment on the spectrum of the original signal, making the environmental information more prominent in the spectral energy and enhancing its perception ability in the subsequent separation and recognition steps. This type of gain processing emphasizes the maintenance of acoustic structure, does not introduce new speech content or artificial sound sources, and only makes purposeful adjustments to the signal energy distribution.

[0089] Processed audio data must be converted to a unified standard to ensure consistent technical representation across different devices, acquisition formats, or sampling frequencies. This conversion process includes operations such as sampling rate unification (e.g., 48kHz), bit depth adjustment (e.g., 16-bit), and encoding format standardization (e.g., WAV, FLAC), ensuring that the audio data is compatible with subsequent signal processing and model input requirements. This conversion also provides a consistent foundation for horizontal comparison and feature mining of environmental sounds.

[0090] The final voice data is stored in a structured audio database, along with the acquisition time and device identification information. Time information supports subsequent time window analysis, such as determining whether specific events occur repeatedly within a specific time period. Device identification tracks potential differences in device type and acquisition location, supporting post-processing compensation strategies for hardware capability differences. This recording process emphasizes the structural binding of metadata and audio content to ensure complete source and context for data use and reconstruction.

[0091] This embodiment can significantly improve the recognizability and accuracy of environmental elements in subsequent separation and identification by synchronously acquiring original voice data containing human voices and environmental sounds, and performing gain processing on the characteristic frequency bands of environmental sounds. The standardized storage mechanism ensures that audio maintains consistent expression across devices and platforms, while the acquisition time and device identification records provide complete data chain support for subsequent retrospective analysis. This mechanism ensures that the input voice data has a comprehensive, stable structural foundation that can be used for multi-dimensional analysis, effectively supporting the input quality of authenticity assessment in downstream tasks.

[0092] In one embodiment, the above step S20 includes:

[0093] S201, loading a pre-trained speech separation model based on a deep neural network;

[0094] S202, extracting time-frequency domain voiceprint features and background noise energy distribution features of the original speech data through the speech separation model;

[0095] S203: generating a frequency band mask matrix according to the frequency band energy ratio of the time-frequency domain voiceprint feature and the background noise energy distribution feature, and using the frequency band mask matrix as a sound source separation threshold;

[0096] S204, performing frequency band masking processing on the frequency domain signal of the original speech data based on the sound source separation threshold by using the speech separation model;

[0097] S205, performing inverse short-time Fourier transform on the masked frequency domain signal to generate a time domain reconstructed signal;

[0098] S206 , separating the time domain reconstructed signal to generate ambient sound data and pure voice data.

[0099] In this embodiment, the purpose of performing speech separation processing on the original speech data is to explicitly decouple the human voice information contained therein from the environmental background sound, thereby facilitating the subsequent independent recognition and reasoning of the human voice content and the environmental context. In this process, it is first necessary to load a neural network model with strong generalization capabilities for multi-channel component analysis of the speech source. The model should be pre-trained on a wide range of audio corpus so that it can recognize the spectral differences between different sound source features and have the ability to learn complex sound source mixture structures. The pre-trained model usually adopts an encoding-mask-decoding structure, including a time-frequency transformation module, a context modeling module and a demixing output module, which can process long time series and high-dimensional spectral features.

[0100] After the raw speech data enters the model, it is converted into a frequency domain representation, and its time-frequency spectrum distribution is obtained through a short-time Fourier transform. The model first extracts the time-frequency domain voiceprint features from the spectrum. This refers to the spectral structure corresponding to the speaker's speech content, which is typically concentrated in the range of 300Hz to 3.4kHz and has relatively continuous energy peak variations. Simultaneously, the energy distribution characteristics of background noise are extracted, including but not limited to non-verbal signals such as traffic sounds, animal calls, crowd noise, and equipment operation. These spectral characteristics often differ significantly from those of human voice components in terms of energy fluctuations and frequency band distribution. The extraction process of these two features can be carried out through a multi-channel self-attention mechanism to enhance the model's ability to model spectral position and temporal changes.

[0101] To achieve efficient separation, the system calculates the energy ratio of each frequency band based on the two aforementioned features, thereby constructing a frequency band mask matrix. This matrix indicates whether each spectral unit should be attributed to a human voice source or an ambient sound source in the current time window. The value is typically between 0 and 1, representing the weight or probability of energy retention. The resulting frequency band mask matrix can be considered a dynamic sound source separation threshold, incorporating multiple parameters such as energy contrast between sound sources, spectral shape differences, and temporal consistency.

[0102] The speech separation model then masks the frequency domain signal of the original speech data on a band-by-band basis using the masking matrix. This suppresses non-speech components in the human voice frequency band and weakens the human voice signal in the ambient sound frequency band, retaining the signal portion that best represents the characteristics of the target sound source. Masking does not destroy the original temporal structure of the signal, but reconstructs the components in the spectral domain. The masked spectral data is then restored to the time domain signal through an inverse short-time Fourier transform, resulting in a speech waveform that has been selectively retained and suppressed through frequency bands, resulting in a higher purity of the target source components.

[0103] Finally, sound source separation is performed based on the reconstructed waveform signal. This is typically further divided into two time-domain audio streams through methods such as frequency band clustering, phase reestimation, or independent component analysis: one segment is the ambient sound data, which retains background noise and non-semantic sound sources, and the other segment is the pure speech data, which is used for subsequent speech content understanding and intent recognition. This process forms a complete closed loop in terms of algorithm structure and computational path, ensuring the stability and generalizability of signal processing.

[0104] This embodiment loads a neural network model with complex sound source modeling capabilities and constructs a frequency band mask matrix with joint time-frequency domain features, which can accurately decompose the human voice and environmental sound in the original mixed speech into independent signal channels that do not interfere with each other. The use of the mask matrix makes the separation process adjustable and interpretable, while retaining the temporal structure and subjective listening experience in the original speech. The final output environmental sound data has high-fidelity background information, and the pure voice data provides a clean input signal for semantic recognition, thereby improving the accuracy and robustness of subsequent processes such as position judgment, dialogue analysis, and environmental semantic reasoning. This processing mechanism not only enhances the physical feasibility of sound source separation, but also provides structured input conditions for downstream models to build multi-source collaborative reasoning.

[0105] In one embodiment, the above step S30 includes:

[0106] S301, performing acoustic event detection on the environmental sound data to generate an acoustic event set;

[0107] S302, performing species classification and identification on the bioacoustic events in the acoustic event set to generate a bioacoustic feature set;

[0108] S303, performing voiceprint pattern matching on the non-biological acoustic events in the acoustic event set to generate a mechanical acoustic feature set;

[0109] S304, merging the bioacoustic feature set and the mechanical acoustic feature set into an environmental sound feature set;

[0110] S305: Retrieve spatial distribution information associated with the environmental sound feature set from the geographic knowledge graph as the knowledge information.

[0111] In this embodiment, in order to extract environmental sound features with geographic directivity and semantic labeling capabilities, acoustic event detection processing is first performed on the environmental sound data. Acoustic event detection refers to dividing a continuous audio stream into event segments containing independent environmental sound components, and marking each segment with its start and end boundaries and preliminary category attributes in the time domain. The determination of acoustic events can be based on sudden changes in time-frequency energy, changes in spectral shape, or through trained supervision models to complete boundary detection and coarse-grained classification, such as sudden chirping, intermittent traffic sounds, continuous background sounds, etc. The output of this process is a structured set of acoustic events, each event with a timestamp, spectral structure and preliminary type label, which serves as the input basis for subsequent classification and recognition.

[0112] In the set of acoustic events, if an event exhibits bioacoustic properties, such as high frequency concentration, obvious syllable structure, and regular temporal rhythm, it can be classified as a bioacoustic event. For such events, species classification and identification are further performed. The classification and identification process can be completed through a multi-class label recognition model. The model structure includes modules such as temporal convolutional networks and spectral attention mechanisms, and supports transfer training of regional corpora to adapt it to common bird or animal species in different geographical environments. The bioacoustic feature set output by this step can include fields such as species name labels, phonetic structure representations, and regional co-occurrence probabilities, providing a basic index for subsequent associated geographic information.

[0113] Events lacking biorhythmic characteristics are classified as non-biological acoustic events, such as mechanical equipment sounds, emergency vehicle sirens, and crowd noise. Features from these events are extracted through voiceprint pattern matching. The matching process is based on the event's spectral envelope and harmonic structure, as well as the similarity between it and pre-set templates or voiceprint vectors in historical records. The output is a set of mechanical acoustic features, including indicators such as event type (e.g., alarm, train passing), voiceprint match confidence, and matching source label.

[0114] The bioacoustic feature set and the mechanical acoustic feature set are then structurally merged to produce a unified set of environmental sound feature sets. This set, based on events, includes event type labels, occurrence times, spectrum descriptions, and category probability distributions, forming a high-level audio representation with both information density and structural versatility.

[0115] After the environmental sound feature set is generated, its label fields (such as species labels, event types) are used as search keys to access the structured geographic knowledge graph system. The graph is constructed based on a predefined knowledge triple structure, such as "White-crowned Bulbul - distributed in - South China", "Ambulance sirens - frequently appear in - central cities". Each node represents a geographic entity or environmental event, and each edge represents a semantic relationship. Through matching calculations between label vectors and knowledge graph entity nodes (such as cosine similarity, embedding vector nearest neighbor search, etc.), the spatial distribution information associated with the current environmental sound feature set is obtained to form knowledge information. This spatial distribution information can not only include the names of administrative regions or ecological belts, but also be associated with their co-occurrence probabilities and credibility levels in specific time periods and specific contexts.

[0116] This embodiment can accurately extract the directional information of sound sources from different sources by deconstructing the environmental sound data into specific acoustic events and executing respective adaptive recognition algorithms on the biological and non-biological events. By converting the set of acoustic events into a set of environmental sound features and further matching it with the geographic knowledge graph, a structured interpretation of the geographic semantics in the recording environment is achieved. This processing not only improves the ability to interpret environmental data, but also provides an input basis based on geographic knowledge support for subsequent location judgment, semantic comparison and multimodal cross-validation. This mechanism achieves effective mapping to structural semantics while maintaining the non-structural characteristics of environmental data, thereby enhancing the stability and scalability of the system in complex spatial semantic analysis.

[0117] In one embodiment, the above step S40 includes:

[0118] S401, performing frame and window processing on the clean voice data to generate a voice signal after frame and window processing;

[0119] S402, extracting Mel-frequency cepstral coefficient features of the speech signal after the frame and window processing to generate an acoustic feature vector;

[0120] S403, performing phoneme sequence decoding on the acoustic feature vector using a pre-trained speech recognition model to generate initial text data;

[0121] S404: Perform grammar correction on the initial text data based on a grammar correction model to generate standardized dialogue text data.

[0122] In this embodiment, pure speech data is first processed by framing and windowing. This process divides the continuous speech signal into short, fixed-length signal frames, each of which is locally stable in time, typically between 20 and 30 milliseconds. During the frame division process, a window function (such as a Hamming window or a Blackman window) is applied for weighting to reduce inter-frame signal discontinuities, thereby reducing spectral leakage during the short-time Fourier transform (SFT) process. The processed speech signal is structured as multiple frames with time-domain envelope constraints, which facilitates subsequent extraction of acoustic features.

[0123] After obtaining the framed and windowed speech signal, its acoustic feature vector needs to be extracted. In this process, Mel-frequency cepstral coefficient features are used as an acoustic representation method. Mel-frequency cepstral coefficients (MFCCs) simulate the nonlinear perception mechanism of the human auditory system and have strong robustness in the field of speech recognition. The specific implementation involves performing a short-time Fourier transform (STFT) on each frame signal, calculating the power spectrum, then performing frequency band compression through a Mel filter bank, and finally extracting characteristic coefficients of several dimensions through a discrete cosine transform (DCT) to form an acoustic feature vector. This feature vector effectively preserves the pronunciation structure of the speech and is an important input for subsequent recognition models.

[0124] Next, the acoustic feature vector is input into a pre-trained speech recognition model for phoneme sequence decoding. The model can adopt an end-to-end structure, such as a deep neural network based on CTC (Connectionist Temporal Classification) or Transducer architecture. This type of model supports mapping feature vector sequences of variable length into speech phonemes or subword-level text sequences. The model can include multiple layers of convolution, recurrent structures, and attention mechanisms to extract temporal correlations and integrate contextual information. The initial text data output is a direct transcription without language rules and may contain spelling, formatting, or word order errors.

[0125] To improve the usability of recognized text, the initial text data needs to be grammatically corrected to produce more structured conversation content. This process is based on a language model, which can be an N-gram statistical model or a Transformer-based deep neural language model, such as BERT or GPT variants. The language model performs syntactic analysis, part-of-speech repair, and contextual reorganization on the initial text to identify and correct grammatical errors, inappropriate word usage, and word order confusion, ultimately outputting standardized conversation text data. The resulting text is highly readable and semantically clear, making it easy to integrate semantic analysis with multimodal data such as environmental information and statement information.

[0126] This embodiment can convert the speech signal into a clearly structured acoustic feature vector by performing frame windowing and Mel-frequency cepstral coefficient extraction on the pure speech data, effectively preserving the acoustic characteristics of the speech and eliminating time domain interference. By using the speech recognition model to perform phoneme decoding on the acoustic feature vector, high-precision speech transcription can be achieved. At the same time, the initial text is post-processed in combination with the grammar correction model to further optimize the recognition quality and language fluency. This processing flow realizes the conversion from unstructured speech signals to analyzable text data, providing a stable information foundation for subsequent semantic recognition, context association and fact verification, and has strong adaptability in the task of comparing user claims with actual context.

[0127] In one embodiment, the above step S50 includes:

[0128] S501, performing data alignment processing on the environmental sound features, knowledge information, conversation text data, and user declaration information to generate a spatiotemporally aligned input data set;

[0129] S502, encoding the environmental sound features in the spatiotemporally aligned input data set into an environmental feature vector;

[0130] S503, encoding the knowledge information in the spatiotemporally aligned input data set into a geographic knowledge vector;

[0131] S504, encoding the conversation text data in the spatiotemporally aligned input data set into a semantic vector;

[0132] S505, encoding the user declaration information in the spatiotemporally aligned input data set into a declaration position vector;

[0133] S506, fusing the environmental feature vector, the geographic knowledge vector, the semantic vector, and the declared location vector to generate a multimodal fusion feature vector;

[0134] S507: Input the multimodal fusion feature vector into a pre-trained location anomaly analysis model, and output the geographic location authenticity analysis result and analysis reasons.

[0135] In this embodiment, when processing multi-source input information for authenticity analysis, it is necessary to first perform unified data alignment processing on environmental sound features, knowledge information, dialogue text data and user declaration information. The purpose of data alignment is to eliminate the differences in time scale and spatial coordinate system between different modal data, so that they are comparable and fusible. In the time domain, each data item can be sliced, normalized or interpolated according to a unified time window; in the spatial domain, a unified coordinate mapping method (such as geographic grid coding or administrative block standards) can be used to map geographic information to a consistent spatial reference framework. The spatiotemporal aligned input data set obtained after alignment has structural consistency, providing a processable input format for subsequent encoding.

[0136] When encoding the environmental sound features in the input dataset into environmental feature vectors, either shallow or deep feature encoding methods can be used. Environmental sound features may include information such as identified species labels, mechanical sound types, and their frequency of occurrence. These features are converted into numerical vectors through an embedding layer or statistical encoding model. Such vectors typically have a fixed dimension and preserve similarity structures between categories. For example, a semantic space can be established between bird species through hierarchical taxonomic relationships, and sound types can be modeled through sound source analogies.

[0137] The encoding of this knowledge is based on a geographic knowledge graph, which contains mappings between known environmental sound labels and geographic regions. For example, certain bird calls occur only in specific areas of the south, or certain mechanical sounds occur only in specific traffic environments. During encoding, graph embedding methods can be combined to represent spatial distribution information and acoustic labels, generating a geographic knowledge vector. This vector can express geographic biases and their relationships with environmental factors.

[0138] After alignment, conversational text data needs to be further encoded into semantic vectors, which represent the implicit meaning and contextual information of the speaker's speech. Text encoding methods can be implemented based on pre-trained language models, such as BERT-like models or lightweight sentence embedding networks, to extract sentence-level semantic representations. The vector output is typically a point in high-dimensional space, capturing information such as location description, scene depiction, and temporal semantics in the text, providing linguistic evidence for subsequent judgment.

[0139] User declarations also need to be converted into actionable vector form. Declarations typically include location claims (e.g., "I'm in Shanghai") and temporal context (e.g., "This morning at 10:00"). Natural language parsing and entity alignment methods can be used to parse these claims into structured geographic locations and time periods. An encoder is then used to represent these claims as declared location vectors, which incorporate spatial coordinate features and temporal range characteristics, facilitating comparison with contextual information.

[0140] The four aforementioned vectors are further fused into a multimodal fused feature vector. This fusion process can be achieved through concatenation, weighted averaging, or an attention mechanism, allowing multi-source data to complement each other in a unified vector space. The resulting fused feature representation simultaneously incorporates geographic, semantic, environmental, and subjective information, enabling cross-modal consistency and conflict detection.

[0141] Finally, the multimodal fusion feature vector is input into a pre-trained location anomaly analysis model. The model's internal structure can be a multi-layer perceptron, a graph attention network, or a fine-tuned large-scale language model architecture, equipped with logical reasoning and inconsistency detection capabilities. The model performs contradiction analysis, association reasoning, and anomaly labeling on the input vector, outputting a location authenticity analysis result and generating an interpretable justification for the analysis. This justification can include conflict labels, reasoning paths, or example explanations to enhance the transparency and credibility of the analysis results.

[0142] This embodiment aligns a variety of heterogeneous information, including environmental sound features, knowledge graph information, and text semantic data with user statements in a unified spatiotemporal manner, and encodes them into unified format vectors. It then uses a multimodal fusion strategy to build a complete information expression structure, which can achieve a systematic judgment of the conflicting relationships between various types of information. After the fusion vector is input into an anomaly analysis model with reasoning capabilities, it can further identify spatial deviations, semantic contradictions, and temporal inconsistencies between information, thereby outputting authenticity analysis results with high interpretability and accuracy. This structure effectively solves the problem of difficulty in interactive comparison of data of different modalities, and improves the ability to judge the credibility of user statements in complex contexts.

[0143] In one embodiment, the above step S507 includes:

[0144] S5071, analyzing the spatial correlation between the environmental feature vector and the geographic knowledge vector in the multimodal fusion feature vector through the logical conflict detection module in the location anomaly analysis model, and generating a credibility score of the geographic conflict point;

[0145] S5072: Detecting the temporal distribution consistency of the environment feature vector and the semantic vector in the multimodal fusion feature vector through the temporal regularity analysis module in the location anomaly analysis model, and generating a time window deviation coefficient;

[0146] S5073: Generating intermediate analysis data including spatial conflict heat and temporal anomaly level based on the geographical conflict point credibility score and the time window deviation coefficient;

[0147] S5074: Based on the intermediate analysis data, determine the abnormal probability of the user's declared information through an abnormality determination decision tree, and generate a geographic location authenticity analysis result;

[0148] S5075: Generate a visual analysis reason text based on the anomaly probability, the geographical conflict point credibility score, and the time window deviation coefficient.

[0149] In this embodiment, a multimodal fusion feature vector is used as input data to integrate coded information from different sources. The environmental feature vector and the geographic knowledge vector respectively represent the acoustic labels extracted from the environmental sounds and their geographical relationships with the regional knowledge graph, which together constitute a spatial information subset. The logical conflict detection module analyzes the spatial semantic compatibility between these two types of vectors to identify their deviations in dimensions such as geographical range, habitat, and probability of occurrence. The model calculates a conflict score based on spatial embedding similarity, label distribution overlap, or regional semantic co-occurrence, and derives a geographic conflict point credibility score for judging the credibility of the information. This score measures the intensity of the conflict between the environmental content and the user's claimed area.

[0150] In the temporal dimension, the temporal features in the environmental feature vector and the semantic vector are fed into the temporal pattern analysis module for comparison. Events in environmental sounds often exhibit specific temporal distribution patterns, such as certain bird calls that occur only in the early morning or during certain seasons, or certain mechanical sounds like traffic noise that exhibit statistical differences between weekdays and weekends. Semantic vectors contain temporal information from the conversation text, such as specific time descriptions, vocabulary for daily routines, or clues to the sequence of events. The module uses time window overlap, event-time co-occurrence statistics, or time series matching algorithms to calculate the degree of temporal alignment between the two and output a time window deviation coefficient as a quantitative result.

[0151] The credibility scores of geographic conflict points and the time window deviation coefficient are input into the structured mapping module to generate intermediate analytical data describing the conflict status of information. This intermediate data clearly identifies two indicators: spatial conflict intensity and temporal anomaly level. The former indicates the degree of spatial inconsistency between the user's declared location and the objective environmental information, while the latter reflects significant shifts or misalignments in temporal usage. This labeled conflict level facilitates subsequent classification and judgment within a rule-driven mechanism.

[0152] The anomaly determination decision tree receives intermediate analysis data as input and makes inferences based on pre-set branching conditions. For example, the decision tree may output a high-risk label when both spatial conflict intensity and temporal anomaly level are high. Each path is designed based on real-world data validation and is both interpretable and stable. The inference result is an anomaly probability, which expresses the degree to which the current data combination does not conform to normal declared behavior. It also outputs the geographic location authenticity analysis result as the overall judgment conclusion.

[0153] Combining this inference information, the model further generates visual analysis justification text. This text is constructed based on a structured data interpretation template, combining elements such as anomaly probability, spatial conflict level, temporal deviation, and referenced environmental feature labels into natural language expressions. The output text can include a description of the specific conflict point, the cause of the data inconsistency, referenced evidence, or the predicted logical path, helping reviewers or users understand the basis for the judgment and enhancing system transparency and trust.

[0154] Example: In fintech, to prevent remote identity forgery and location fraud, when a bank customer remotely applies for high-risk financial services (such as online large loan activation), the system automatically collects background voice data from a location in Guangdong Province. This voice data, acquired through the intelligent customer service access module, contains the applicant's statement and surrounding ambient sounds. The collected voice data is first fed into the ambient sound enhancement module, where the system uses selective frequency band gain to enhance the non-speech background components, making environmental features such as bird song and traffic sounds more clearly distinguishable during subsequent processing. The audio is converted to a standardized format and stored with the acquisition time and access device ID, creating a traceable data file. The audio data is then fed into a speech separation model, which extracts the time-frequency voiceprint and background noise energy distribution, constructs a frequency band mask matrix, and separates speech from non-speech signals. The system generates two independent audio streams: one containing ambient sounds such as street noise, crowd conversation, and a mixture of dialects; the other contains the pure voice data of the user's statement. Ambient sounds were fed into the acoustic event detection system, which classified them and found them to include suspected bird calls and specific non-natural human voices. The system further identified species within the bioacoustic events, identifying characteristic soundprints of bird species endemic to southern China. Among the non-biological features, it identified a high-confidence match between Cantonese tonal patterns and the local ambulance siren frequency band. Multi-source labels were summarized into a set of ambient sound features, which were then searched against a built-in geographic knowledge graph. The system returned information indicating that this soundprint population primarily occurred in coastal cities in Guangdong. Simultaneously, after framing, windowing, and feature extraction, the user's speech content was processed using a speech recognition model to generate preliminary text. This text was then grammatically corrected using a language model to produce structured conversational text containing semantic clues such as "I'm at the beach now" and "There were a lot of people in the square this morning." This set of ambient sound features, geographic knowledge, and conversational text data were then fed into the analysis model along with the user's declared location information. The system first unified the time and geographic dimensions, then vectorized each type of information and performed multimodal fusion. This fused feature was then fed into the location anomaly analysis model. The system's logical conflict module found that the bird and dialect voiceprints strongly supported the claimed location, resulting in a low spatial conflict score. However, the temporal pattern analysis module detected a significant discrepancy between the ambient sound and the user's statement and the different time periods. Intermediate analysis data indicated a low spatial conflict intensity and a medium temporal anomaly level, resulting in a 23% probability of anomaly. The model ultimately outputted a "credible authenticity" statement, along with the rationale: the ambient sound had a high correlation with the claimed location, but the semantic timeline of the conversation differed significantly from the background sound, recommending a secondary verification. This mechanism enables financial institutions to automatically verify the geographic authenticity of high-risk remote activities without disrupting business processes.

[0155] In a healthcare service scenario, a remote chronic disease patient conducted daily follow-up visits with the hospital's health management center via a voice platform. During these visits, the system automatically collected ambient audio and conversation data for location consistency monitoring and health behavior authenticity assessment. In one voice recording, the patient claimed to be "currently taking a walk around the neighborhood." The system detected weak background broadcasts and short human voices in the raw voice signal. To improve data quality, the system performed gain processing on the audio to enhance the ambient noise, marked the acquisition time, and encoded the file on the terminal device before saving it. The audio was processed using a pre-trained speech separation model, constructing frequency band masks based on the time-frequency structure and outputting two independent data streams. The ambient audio component included the tone of local news from the community radio, the sounds of elderly people chatting, and occasional vehicle horns; the speech component retained only the patient's proactive statements. After acoustic event detection, the system identified a mention of an announcement related to the "Shinan Community Health Station" in the broadcast segment and a human voice segment with Shandong dialect characteristics in the non-biometric voiceprint. While bioacoustic features were not clearly evident, mechanical acoustic features closely matched the tone from the residential parking lot's sound system. Through a graph search, the system classified these features as belonging to the Shinan District activity sound domain, consistent with the patient's claimed location. After speech recognition, the conversational speech was converted into structured text, including content such as "Just got back from a workout. There's a free clinic announcement downstairs today." The system extracted semantic tags related to health activities and timelines. Ambient sound features, spatial knowledge tags, semantic conversation content, and user registered address information were quantized and integrated into the analysis model, which performed spatial and semantic logic checks to determine whether there were any geographic or temporal conflicts between these dimensions. The final system analysis concluded: "The claim is credible," and the analyzed text indicated: "The announcement content corresponds to the location, and the speech semantics and temporal distribution are consistent, supporting the current geographic claim." This process leverages environmental acoustic cues to enhance service authenticity verification in healthcare interaction scenarios without compromising the patient's interactive experience, improving the credibility of remote health data and assisting in closed-loop control of chronic disease follow-up and health behavior management.

[0156] This embodiment structures multimodal data such as environmental feature vectors, geographic knowledge vectors, semantic vectors, and declaration information, and relies on logical conflict detection and temporal law analysis to verify the inherent consistency of information. It also introduces structured mapping and explainable reasoning mechanisms, which can systematically identify anomalies such as semantic deviations and regional inconsistencies between multi-source information. By generating credibility scores and deviation quantification coefficients, a standardized expression of spatial and temporal conflicts is achieved, and authenticity is output probabilistically under a decision tree mechanism. The resulting visual analysis reason not only provides judgment results, but also has the ability to trace data hierarchically and explain logical paths, effectively improving the level of intelligent analysis of the authenticity of user claims.

[0157] In one embodiment, a device for authenticity analysis based on environmental sound features is provided, and the device for authenticity analysis based on environmental sound features corresponds one-to-one to the authenticity analysis method based on environmental sound features in the above embodiment. Figure 3 , Figure 3 This is a functional module diagram of a preferred embodiment of the authenticity analysis device based on environmental sound features of the present invention. It includes a speech acquisition module 10, a speech separation module 20, an environmental perception and analysis module 30, a speech recognition module 40, and an authenticity analysis module 50. Each functional module is described in detail below:

[0158] The voice collection module 10 is used to obtain original voice data including environmental sounds;

[0159] The speech separation module 20 is used to perform speech separation processing on the original speech data to generate environmental sound data and pure speech data;

[0160] An environmental perception analysis module 30 is configured to analyze the environmental sound data, extract environmental sound features, and retrieve knowledge information associated with the environmental sound features;

[0161] A speech recognition module 40 is used to perform speech content recognition on the clean speech data to generate dialogue text data;

[0162] The authenticity analysis module 50 is used to input the environmental sound features, knowledge information, dialogue text data and user declaration information into the analysis model and output the authenticity analysis results.

[0163] In one embodiment, the voice collection module 10 is specifically configured to:

[0164] The original voice signal including the ambient sound is collected through the microphone module of the terminal device;

[0165] Performing ambient sound audio segment selective gain processing on the original speech signal to generate original speech data containing enhanced ambient sound;

[0166] Converting the original voice data into a standardized audio format file;

[0167] The standardized audio format file is stored in a storage area, and the acquisition time and acquisition device identifier of the original voice data are recorded.

[0168] In one embodiment, the speech separation module 20 is specifically configured to:

[0169] Load a pre-trained deep neural network-based speech separation model;

[0170] Extracting the time-frequency domain voiceprint features and background noise energy distribution features of the original speech data through the speech separation model;

[0171] Generating a frequency band mask matrix according to the frequency band energy ratio of the time-frequency domain voiceprint feature and the background noise energy distribution feature, and using the frequency band mask matrix as a sound source separation threshold;

[0172] Performing frequency band masking processing on the frequency domain signal of the original speech data based on the sound source separation threshold by the speech separation model;

[0173] Performing inverse short-time Fourier transform on the masked frequency domain signal to generate a time domain reconstructed signal;

[0174] The time domain reconstructed signal is separated to generate ambient sound data and pure speech data.

[0175] In one embodiment, the environment perception analysis module 30 is specifically configured to:

[0176] Performing acoustic event detection on the environmental sound data to generate an acoustic event set;

[0177] Performing species classification and identification on the bioacoustic events in the acoustic event set to generate a bioacoustic feature set;

[0178] Performing voiceprint pattern matching on non-biological acoustic events in the acoustic event set to generate a mechanical acoustic feature set;

[0179] Combining the bioacoustic feature set and the mechanical acoustic feature set into an environmental sound feature set;

[0180] The spatial distribution information associated with the set of environmental sound features is retrieved from the geographic knowledge graph as the knowledge information.

[0181] In one embodiment, the speech recognition module 40 is specifically configured to:

[0182] Performing frame and window processing on the clean voice data to generate a frame and window processed voice signal;

[0183] Extracting Mel-frequency cepstral coefficient features of the speech signal after the frame and window processing to generate an acoustic feature vector;

[0184] Performing phoneme sequence decoding on the acoustic feature vector using a pre-trained speech recognition model to generate initial text data;

[0185] The initial text data is grammatically corrected based on a grammar correction model to generate standardized dialogue text data.

[0186] In one embodiment, the authenticity analysis module 50 is specifically configured to:

[0187] Performing data alignment processing on the environmental sound features, knowledge information, conversation text data, and user declaration information to generate a spatiotemporally aligned input data set;

[0188] encoding the ambient sound features in the spatiotemporally aligned input data set into an ambient feature vector;

[0189] Encoding the knowledge information in the spatiotemporally aligned input data set into a geographic knowledge vector;

[0190] Encoding the conversation text data in the spatiotemporally aligned input data set into semantic vectors;

[0191] encoding user declaration information in the spatiotemporally aligned input data set into a declaration position vector;

[0192] fusing the environmental feature vector, the geographic knowledge vector, the semantic vector, and the declared location vector to generate a multimodal fusion feature vector;

[0193] The multimodal fusion feature vector is input into a pre-trained location anomaly analysis model, and the geographic location authenticity analysis result and analysis reason are output.

[0194] In one embodiment, the authenticity analysis module 50 is specifically configured to:

[0195] By using the logical conflict detection module in the location anomaly analysis model, the spatial correlation between the environmental feature vector and the geographic knowledge vector in the multimodal fusion feature vector is analyzed to generate a credibility score of the geographic conflict point;

[0196] The time regularity analysis module in the position anomaly analysis model is used to detect the time distribution consistency of the environmental feature vector and the semantic vector in the multimodal fusion feature vector, and generate a time window deviation coefficient;

[0197] Generating intermediate analysis data including spatial conflict heat and temporal anomaly level based on the credibility score of the geographical conflict point and the time window deviation coefficient;

[0198] Based on the intermediate analysis data, determining the abnormal probability of the user's declared information through an abnormality determination decision tree, and generating a geographic location authenticity analysis result;

[0199] A visual analysis reason text is generated based on the anomaly probability, the geographical contradiction point credibility score and the time window deviation coefficient.

[0200] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a method for authenticity analysis based on environmental sound characteristics.

[0201] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a method for authenticity analysis based on environmental sound characteristics.

[0202] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0203] Obtaining raw voice data including environmental sounds;

[0204] Performing voice separation processing on the original voice data to generate ambient sound data and pure voice data;

[0205] Analyzing the environmental sound data, extracting environmental sound features, and retrieving knowledge information associated with the environmental sound features;

[0206] Performing voice content recognition on the clean voice data to generate dialogue text data;

[0207] The environmental sound features, knowledge information, dialogue text data and user declaration information are input into an analysis model, and an authenticity analysis result is output.

[0208] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0209] Obtaining raw voice data including environmental sounds;

[0210] Performing voice separation processing on the original voice data to generate ambient sound data and pure voice data;

[0211] Analyzing the environmental sound data, extracting environmental sound features, and retrieving knowledge information associated with the environmental sound features;

[0212] Performing voice content recognition on the clean voice data to generate dialogue text data;

[0213] The environmental sound features, knowledge information, dialogue text data and user declaration information are input into an analysis model, and an authenticity analysis result is output.

[0214] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0215] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0216] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0217] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A method for authenticity analysis based on environmental sound characteristics, characterized in that: The following steps are involved: Obtaining raw voice data including environmental sounds; Performing voice separation processing on the original voice data to generate ambient sound data and pure voice data; Analyzing the environmental sound data, extracting environmental sound features, and retrieving knowledge information associated with the environmental sound features; Performing voice content recognition on the clean voice data to generate dialogue text data; The environmental sound features, knowledge information, dialogue text data and user declaration information are input into an analysis model, and an authenticity analysis result is output.

2. The authenticity analysis method based on environmental sound features according to claim 1, characterized in that: Get raw voice data containing environmental sounds, including: The original voice signal including the ambient sound is collected through the microphone module of the terminal device; Performing ambient sound audio segment selective gain processing on the original speech signal to generate original speech data containing enhanced ambient sound; Converting the original voice data into a standardized audio format file; The standardized audio format file is stored in a storage area, and the acquisition time and acquisition device identifier of the original voice data are recorded.

3. The authenticity analysis method based on environmental sound features according to claim 1, characterized in that: Performing voice separation processing on the original voice data to generate ambient sound data and pure voice data, including: Load a pre-trained deep neural network-based speech separation model; Extracting the time-frequency domain voiceprint features and background noise energy distribution features of the original speech data through the speech separation model; Generating a frequency band mask matrix according to the frequency band energy ratio of the time-frequency domain voiceprint feature and the background noise energy distribution feature, and using the frequency band mask matrix as a sound source separation threshold; Performing frequency band masking processing on the frequency domain signal of the original speech data based on the sound source separation threshold by the speech separation model; Performing inverse short-time Fourier transform on the masked frequency domain signal to generate a time domain reconstructed signal; The time domain reconstructed signal is separated to generate ambient sound data and pure speech data.

4. The authenticity analysis method based on environmental sound features according to claim 1, characterized in that: Analyzing the environmental sound data, extracting environmental sound features, and retrieving knowledge information associated with the environmental sound features, including: Performing acoustic event detection on the environmental sound data to generate an acoustic event set; Performing species classification and identification on the bioacoustic events in the acoustic event set to generate a bioacoustic feature set; Performing voiceprint pattern matching on non-biological acoustic events in the acoustic event set to generate a mechanical acoustic feature set; Combining the bioacoustic feature set and the mechanical acoustic feature set into an environmental sound feature set; The spatial distribution information associated with the set of environmental sound features is retrieved from the geographic knowledge graph as the knowledge information.

5. The authenticity analysis method based on environmental sound features according to claim 1, characterized in that: Performing voice content recognition on the clean voice data to generate conversation text data includes: Performing frame and window processing on the clean voice data to generate a frame and window processed voice signal; Extracting Mel-frequency cepstral coefficient features of the speech signal after the frame and window processing to generate an acoustic feature vector; Performing phoneme sequence decoding on the acoustic feature vector using a pre-trained speech recognition model to generate initial text data; The initial text data is grammatically corrected based on a grammar correction model to generate standardized dialogue text data.

6. The authenticity analysis method based on environmental sound features according to claim 1, characterized in that: The environmental sound features, knowledge information, conversation text data, and user declaration information are input into an analysis model, and authenticity analysis results are output, including: Performing data alignment processing on the environmental sound features, knowledge information, conversation text data, and user declaration information to generate a spatiotemporally aligned input data set; encoding the ambient sound features in the spatiotemporally aligned input data set into an ambient feature vector; Encoding the knowledge information in the spatiotemporally aligned input data set into a geographic knowledge vector; Encoding the conversation text data in the spatiotemporally aligned input data set into semantic vectors; encoding user declaration information in the spatiotemporally aligned input data set into a declaration position vector; fusing the environmental feature vector, the geographic knowledge vector, the semantic vector, and the declared location vector to generate a multimodal fusion feature vector; The multimodal fusion feature vector is input into a pre-trained location anomaly analysis model, and the geographic location authenticity analysis result and analysis reason are output.

7. The authenticity analysis method based on environmental sound features according to claim 6, characterized in that: The multimodal fusion feature vector is input into a pre-trained location anomaly analysis model, and the location authenticity analysis result and analysis reasons are output, including: By using the logical conflict detection module in the location anomaly analysis model, the spatial correlation between the environmental feature vector and the geographic knowledge vector in the multimodal fusion feature vector is analyzed to generate a credibility score of the geographic conflict point; By using the time regularity analysis module in the position anomaly analysis model, the time distribution consistency of the environment feature vector and the semantic vector in the multimodal fusion feature vector is detected to generate a time window deviation coefficient; Generating intermediate analysis data including spatial conflict heat and temporal anomaly level based on the credibility score of the geographical conflict point and the time window deviation coefficient; Based on the intermediate analysis data, determining the abnormal probability of the user's declared information through an abnormality determination decision tree, and generating a geographic location authenticity analysis result; A visual analysis reason text is generated based on the anomaly probability, the geographical contradiction point credibility score and the time window deviation coefficient.

8. An authenticity analysis device based on environmental sound characteristics, characterized in that: The authenticity analysis device based on environmental sound features includes: The voice acquisition module is used to obtain the original voice data including the environmental sound; A speech separation module is used to perform speech separation processing on the original speech data to generate ambient sound data and pure speech data; An environmental perception analysis module, configured to analyze the environmental sound data, extract environmental sound features, and retrieve knowledge information associated with the environmental sound features; A speech recognition module, configured to perform speech content recognition on the clean speech data and generate dialogue text data; The authenticity analysis module is used to input the environmental sound characteristics, knowledge information, dialogue text data and user declaration information into the analysis model and output the authenticity analysis results.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and an authenticity analysis program based on environmental sound features stored in the memory and capable of running on the processor. When the authenticity analysis program based on environmental sound features is executed by the processor, the steps of the authenticity analysis method based on environmental sound features as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores an authenticity analysis program based on environmental sound features, which, when executed by a processor, implements the steps of the authenticity analysis method based on environmental sound features as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Authentication method and device based on knowledge graph and voiceprint recognition, equipment and medium

    CN111883140A

  • Medical business data processing method and device, computer equipment and storage medium

    CN116245539A

  • Voiceprint recognition method and device, electronic equipment and medium

    CN116631411A

  • Speech synthesis model training method and device based on common sense reasoning and synthesis method

    CN117238275A

  • Voice separation enhancement method and system in multi-sound-source and noise environment

    CN117238311A

Cited By

  • Interference source identification method and system based on vehicle-mounted mobile monitoring data analysis

    CN121485838A

  • Subway station multi-modal data noise processing method, computer equipment and program product

    CN122333304A

  • Information verification system, information verification method, and information verification program

    JP7843984B1