Abnormal sound detection method and device
By jointly training the abnormal sound synthesis model and the description model, and using duality loss to generate abnormal sound training samples, the problem of low recognition accuracy caused by lack of samples is solved, and high-quality abnormal sound recognition is achieved.
Patent Information
- Application Number
- CN202511676844.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-02-27
AI Technical Summary
In the absence of abnormal sound samples, existing technologies struggle to train high-quality abnormal sound recognition models, resulting in low recognition accuracy. Furthermore, manual annotation is costly, inefficient, subjective, and subject to significant labeling noise.
By acquiring sound description text and jointly training an abnormal sound synthesis model and a description model, a dual consistency loss is constructed to generate abnormal sound training samples, thereby achieving bidirectional constraints on semantics and sound and generating sound samples that conform to abnormal characteristics.
In the absence of real abnormal samples, a large-scale, controllable abnormal sound training dataset was constructed, which improved the quality of abnormal sound generation and the accuracy of the recognition model, reduced noise generation, and enhanced the semantic consistency and generalization ability of the model.
Smart Images

Figure CN121583284A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of sound recognition technology, and more specifically, to an abnormal sound detection method and apparatus. Background Technology
[0002] Anomaly detection utilizes artificial intelligence and acoustic analysis techniques. Typically, anomaly detection relies on supervised training with a large number of labeled anomaly sound samples. However, in real-world production and operational scenarios, anomaly sounds often exhibit diversity, randomness, and scarcity, making it difficult to obtain sufficient high-quality samples. Furthermore, manually labeling anomaly sound types and features is not only costly and inefficient but also suffers from strong subjectivity and high label noise. Therefore, how to train a high-quality anomaly detection model in the absence of sufficient anomaly sound samples is a pressing problem that needs to be solved. Summary of the Invention
[0003] In view of this, this application provides an abnormal sound detection method and apparatus to solve the problem of low accuracy of the recognition model due to the small number of abnormal sound samples.
[0004] Specifically, this application is implemented through the following technical solution: In a first aspect, embodiments of this application provide an abnormal sound detection method, including: Obtain a first sound description text; the first sound description text is used to describe the environmental sound in the target work scenario; the first sound description text contains first descriptive information for abnormal sounds; The first sound description text is input into the abnormal sound synthesis model to obtain the first sound segment corresponding to the first sound description text; the first sound segment contains the abnormal sound corresponding to the first description information. The first sound segment is input into the abnormal sound description model to obtain the second sound description text corresponding to the first sound segment; the second sound description text contains second description information for the abnormal sound in the first sound segment. Based on the first sound description text and the second sound description text, the abnormal sound synthesis model and the abnormal sound description model are jointly trained, and abnormal sound training samples are generated based on the trained abnormal sound synthesis model. The abnormal sound detection model is trained using the abnormal sound training samples to obtain a trained abnormal sound detection model; the trained abnormal sound detection model is used to detect abnormal sounds in the target work scene.
[0005] Optionally, the step of jointly training the abnormal sound synthesis model and the abnormal sound description model based on the first sound description text and the second sound description text includes: Based on the difference information between the first sound description text and the second sound description text, a first dual consistency loss is constructed. Based on the first duality consistency loss, the abnormal sound synthesis model and the abnormal sound description model are jointly trained.
[0006] Optionally, the method further includes: Acquire a second sound segment; the second sound segment contains abnormal sounds from the target work scenario; The second sound segment is input into the abnormal sound description model to obtain the third sound description text corresponding to the second sound segment; The third sound description text is input into the abnormal sound synthesis model to obtain the third sound segment corresponding to the third sound description text; Based on the difference information between the second sound segment and the third sound segment, a second dual consistency loss is constructed; The step of jointly training the abnormal sound synthesis model and the abnormal sound description model based on the first sound description text and the second sound description text includes: The abnormal sound synthesis model and the abnormal sound description model are jointly trained based on the first dual consistency loss and the second dual consistency loss.
[0007] Optionally, the step of jointly training the abnormal sound synthesis model and the abnormal sound description model based on the first sound description text and the second sound description text includes: Based on the first dual consistency loss, the second dual consistency loss, the first weight corresponding to the first dual consistency loss, and the second weight corresponding to the second dual consistency loss, a joint loss is constructed. The abnormal sound synthesis model and the abnormal sound description model are jointly trained using the joint loss.
[0008] Optionally, the method further includes: Obtain a fourth sound segment and any abnormal sound type label; the fourth sound segment contains normal sounds under the target work scenario. The fourth sound segment and the abnormal sound type label are input into the abnormal sound description model to obtain the fourth sound description text that matches the fourth sound segment and the abnormal sound type label; the fourth sound description text contains the description information corresponding to the abnormal sound type label; The fourth sound description text is used as the first sound description text.
[0009] Optionally, the abnormal sound type label includes at least one of the following: Friction, impact, fracture, specific frequency range, explosion.
[0010] Optionally, the method further includes: Obtain the detection results output by the trained abnormal sound detection model; Based on the detection results, abnormal information of the target work scenario is determined; the abnormal information includes at least one of abnormal type, abnormal location, and abnormal object.
[0011] Secondly, this application also provides an abnormal sound detection device, comprising: The acquisition module is used to acquire a first sound description text; the first sound description text is used to describe the environmental sound in the target work scenario; the first sound description text contains first descriptive information for abnormal sounds. The first generation module is used to input the first sound description text into the abnormal sound synthesis model to obtain a first sound segment corresponding to the first sound description text; the first sound segment contains the abnormal sound corresponding to the first description information. The second generation module is used to input the first sound segment into the abnormal sound description model to obtain the second sound description text corresponding to the first sound segment; the second sound description text contains second description information for the abnormal sound in the first sound segment. The first training module is used to jointly train the abnormal sound synthesis model and the abnormal sound description model based on the first sound description text and the second sound description text, and generate abnormal sound training samples based on the trained abnormal sound synthesis model. The second training module is used to train the abnormal sound detection model using the abnormal sound training samples to obtain a trained abnormal sound detection model; the trained abnormal sound detection model is used to detect abnormal sounds in the target work scene.
[0012] Optionally, the first training module is used for: Based on the difference information between the first sound description text and the second sound description text, a first dual consistency loss is constructed. Based on the first duality consistency loss, the abnormal sound synthesis model and the abnormal sound description model are jointly trained.
[0013] Optionally, the apparatus further includes a third generation module for: Acquire a second sound segment; the second sound segment contains abnormal sounds from the target work scenario; The second sound segment is input into the abnormal sound description model to obtain the third sound description text corresponding to the second sound segment; The third sound description text is input into the abnormal sound synthesis model to obtain the third sound segment corresponding to the third sound description text; Based on the difference information between the second sound segment and the third sound segment, a second dual consistency loss is constructed; The first training module is used for: The abnormal sound synthesis model and the abnormal sound description model are jointly trained based on the first dual consistency loss and the second dual consistency loss.
[0014] Optionally, the first training module is used for: Based on the first dual consistency loss, the second dual consistency loss, the first weight corresponding to the first dual consistency loss, and the second weight corresponding to the second dual consistency loss, a joint loss is constructed. The abnormal sound synthesis model and the abnormal sound description model are jointly trained using the joint loss.
[0015] Optionally, the apparatus further includes a fourth generation module for: Obtain a fourth sound segment and any abnormal sound type label; the fourth sound segment contains normal sounds under the target work scenario. The fourth sound segment and the abnormal sound type label are input into the abnormal sound description model to obtain the fourth sound description text that matches the fourth sound segment and the abnormal sound type label; the fourth sound description text contains the description information corresponding to the abnormal sound type label; The fourth sound description text is used as the first sound description text.
[0016] Optionally, the abnormal sound type label includes at least one of the following: Friction, impact, fracture, specific frequency range, explosion.
[0017] Optionally, the device further includes an analysis module for: Obtain the detection results output by the trained abnormal sound detection model; Based on the detection results, abnormal information of the target work scenario is determined; the abnormal information includes at least one of abnormal type, abnormal location, and abnormal object.
[0018] Thirdly, embodiments of this application also provide a computer device, which includes a processor and a memory. The memory stores machine-readable instructions executable by the processor. The processor is used to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, they perform the steps of the first aspect above, or any possible implementation of the first aspect.
[0019] Fourthly, optional embodiments of this application also provide a computer-readable storage medium storing a computer program that, when run, performs the steps of the first aspect or any possible implementation of the first aspect.
[0020] The abnormal sound detection method and apparatus provided in this application can form a dual consistency loss based on sound description text and sound segments, and jointly train the abnormal sound synthesis model and the abnormal sound description model, realizing bidirectional constraints of semantics and sound. Thus, in the absence of real abnormal samples, a large-scale and controllable abnormal sound training dataset can be constructed. At the same time, the dual consistency loss is constructed based on the difference information between the first sound description text and the second sound description text generated by the model, so that the model maintains semantic consistency in both the synthesis and description directions, which helps to generate sound samples that are more consistent with abnormal characteristics and improves the quality of abnormal sound generation. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating an abnormal sound detection method according to an exemplary embodiment of this application; Figure 2 This is a schematic diagram illustrating the abnormal sound model training process in an exemplary embodiment of this application; Figure 3 This is a schematic diagram of an abnormal sound detection device shown in an exemplary embodiment of this application; Figure 4 This is a schematic diagram of a computer device illustrated in an exemplary embodiment of this application. Detailed Implementation
[0022] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0023] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0024] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0025] Research has found that during the training process of abnormal sound detection models, abnormal sounds often exhibit diversity, randomness, and scarcity, making it difficult to obtain enough high-quality abnormal sound samples. At the same time, manually labeling abnormal sound types and features is not only costly and inefficient, but also suffers from strong subjectivity and high label noise.
[0026] In view of this, embodiments of this application provide an abnormal sound detection method that can form a dual consistency loss based on sound description text and sound segments, and jointly train an abnormal sound synthesis model and an abnormal sound description model, realizing bidirectional constraints on semantics and sound. Thus, in the absence of real abnormal samples, a large-scale and controllable abnormal sound training dataset can be constructed. At the same time, a dual consistency loss is constructed based on the difference information between the first sound description text and the second sound description text generated by the model, so that the model maintains semantic consistency in both synthesis and description directions, which helps to generate sound samples that are more consistent with abnormal characteristics and improves the quality of abnormal sound generation.
[0027] The deficiencies of the existing technical solutions are the result of the inventor's practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed in this application below should be considered as the inventor's contributions to this application.
[0028] To facilitate understanding of this embodiment, the application scenario of the abnormal sound detection method disclosed in this application embodiment will first be introduced. The execution subject of the abnormal sound detection method provided in this application embodiment can be a computer device. In some possible implementations, the abnormal sound detection method can be implemented by a processor calling computer-readable instructions stored in memory.
[0029] See Figure 1 The diagram shown is a flowchart illustrating an abnormal sound detection method according to an exemplary embodiment of this application. The method includes steps S101-S105, wherein: S101. Obtain the first sound description text; the first sound description text is used to describe the environmental sound in the target work scenario; the first sound description text contains first description information for abnormal sounds.
[0030] The aforementioned target operation scenarios may include industrial equipment scenarios, security monitoring scenarios, medical monitoring scenarios, etc., and ambient sound can refer to the sound in the aforementioned scenarios.
[0031] The aforementioned first sound description text may include textual information such as main activities, abnormal situations, time characteristics, context inferences, and summaries. For example, the first sound description text may be as follows: Key activity: The primary activity identified in the audio is the operation of a large, electrically driven device, most likely a refrigerator or freezer. A continuous hum indicates that the compressor is running continuously, a typical characteristic of such equipment.
[0032] Anomaly: No significant anomalies were detected in the audio. The only perceptible phenomenon was a slight high-frequency noise (hiss), which could originate from noise artifacts from internal electronic components of the recording equipment or appliance. However, this noise remained consistent throughout the recording and did not interfere with the main low-frequency hum, so it can be considered part of normal operating sound rather than an anomaly.
[0033] Temporal characteristics: The audio begins with a stable low-frequency hum, with constant pitch and volume, without fluctuations or cadence changes. At a certain moment, the hum and hiss abruptly cease completely, immediately followed by a loud synthesized electronic sound. This synthesized sound is a pure low-frequency square wave with a center frequency of approximately 110Hz, and continues until the end of the recording. The transition from the hum to the synthesized sound is very abrupt, without fading or lingering sound.
[0034] Scenario inference: The recording appears to have taken place in a small, acoustically sound-absorbing space, possibly a kitchen or equipment room. The quiet, isolated environment, devoid of other audible sounds, indicates that the recording environment was well-controlled. The presence of hum and hiss suggests that the microphone was positioned very close to the sound source, directly capturing equipment operation sounds rather than ambient reflections.
[0035] In summary, the audio clip primarily records the steady operating sounds of household appliances (especially refrigerators or freezers), characterized by continuous low-frequency noise. The recording then abruptly switches to a loud, synthesized electronic sound, creating a strong contrast and abrupt effect. This sudden change may be used to depict a situation where daily operation is interrupted by some kind of emergency electronic alarm. The overall environment is controlled and quiet, emphasizing the contrast between the mechanical, everyday sounds and the human alarm.
[0036] The first sound description text may contain descriptive text about a certain sound. The object described by this descriptive text may be derived from an actual sound, from another piece of text, or from a combination of sound and text. This descriptive text may be generated by a trained model or determined and provided manually.
[0037] For example, a first sound description text can be determined using a real-world ambient sound with unusual sounds.
[0038] Alternatively, a simple text describing a sound can be used to generate a more complex text describing the sound, resulting in the first sound description text.
[0039] Alternatively, a normal sound segment and the characteristic information of the abnormal sound can be used to obtain a first sound description text describing the abnormal sound.
[0040] In some possible implementations, a fourth sound segment and any abnormal sound type tag can be obtained; the fourth sound segment contains normal sounds in the target work scenario; the fourth sound segment and the abnormal sound type tag are input into an abnormal sound description model to obtain a fourth sound description text that matches the fourth sound segment and the abnormal sound type tag; the fourth sound description text contains description information corresponding to the abnormal sound type tag; the fourth sound description text is used as the first sound description text.
[0041] The fourth sound segment can contain normal sounds within the target work scenario, such as sounds emitted by industrial equipment during normal operation, sounds from a security monitoring scenario under safe conditions, or sounds from a medical monitoring scenario under normal conditions. The aforementioned abnormal sound type tags can be semantic tags describing the characteristics of abnormal sounds, such as friction, impact, breakage, a specific frequency range, or explosion.
[0042] The aforementioned abnormal sound description model can be considered a multimodal large model, a type of artificial intelligence model capable of simultaneously processing and understanding multiple information modalities (such as text, images, audio, video, and speech). Unlike traditional models that only process a single modality (such as language or vision), multimodal large models map data from different modalities into the same semantic space through a unified representation and alignment mechanism, achieving cross-modal understanding and generation. For example, it can generate images based on text descriptions, generate text descriptions based on audio, images, or other content, or perform semantic analysis based on video content.
[0043] In this way, based on the normal sound and abnormal sound type labels, the required abnormal sound description information can be obtained, improving the realism of the target operation scenario in the description information and increasing the coverage of abnormal sound type description information.
[0044] For example, in the target operation scenario of industrial equipment, the normal operation sound of the factory assembly line equipment can be used as the fourth sound segment, the abnormal attribute label can be metal friction, and the first sound description text generated by the abnormal sound description model can include "periodic sharp friction sound mixed in with stable low-frequency operation sound", etc.
[0045] S102. Input the first sound description text into the abnormal sound synthesis model to obtain the first sound segment corresponding to the first sound description text; the first sound segment contains the abnormal sound corresponding to the first description information.
[0046] In this step, the abnormal sound synthesis model can generate sound segments that match the input text. This abnormal sound synthesis model can be a single-modal large model or a multi-modal large model.
[0047] For example, an anomalous sound synthesis model can be based on a vocoder architecture of a diffusion process, taking a semantic vector output by a text encoder as input and outputting the Mel spectrum of the target anomalous sound.
[0048] S103. Input the first sound segment into the abnormal sound description model to obtain the second sound description text corresponding to the first sound segment; the second sound description text contains second description information for the abnormal sound in the first sound segment.
[0049] In this step, the abnormal sound description model can be the same model as the abnormal sound description model used to determine the first sound description text, or it can be a different model.
[0050] For example, an abnormal sound description model can adopt a text generation structure based on audio Transformer to map sound features to semantic space, thereby achieving a two-way semantic correspondence from audio to text.
[0051] S104. Based on the first sound description text and the second sound description text, jointly train the abnormal sound synthesis model and the abnormal sound description model, and generate abnormal sound training samples based on the trained abnormal sound synthesis model.
[0052] The task of generating a first sound segment from the first sound description text is the inverse of the task of generating a second sound description text from the first sound segment. Based on the inverse relationship between the two tasks, joint training with mutual supervision can be performed to improve the accuracy of the model and the quality of the generated abnormal sounds.
[0053] In some possible implementations, a first dual consistency loss can be constructed based on the difference information between the first sound description text and the second sound description text; based on the first dual consistency loss, the abnormal sound synthesis model and the abnormal sound description model can be jointly trained.
[0054] In this step, the differences between the first sound description text and the second sound description text can be determined. These differences may include text semantic similarity, cross-modal semantic consistency, etc. Based on these differences, a first dual consistency loss can be constructed. By minimizing this loss, the parameters of the abnormal sound synthesis model and the abnormal sound description model can be optimized, so that the two models can be trained under mutual constraints.
[0055] In one possible implementation, the first dual consistency loss may include semantic similarity-based loss, reconstruction error-based loss, contrastive learning-based loss, distribution consistency-based loss, etc.
[0056] See Figure 2 The diagram shown illustrates the training process of an abnormal sound model according to an exemplary embodiment of this application. In this process, normal sound samples and abnormal attribute labels can be input into the abnormal sound description model to obtain a first sound description text; a first sound segment corresponding to the first sound description text can be generated through an abnormal sound synthesis model; the first sound segment can be input into the abnormal sound description model to obtain a second sound description text; a first duality consistency loss can be constructed using the first and second sound description texts; and the abnormal sound synthesis model and the abnormal sound description model can be jointly trained using the first duality consistency loss.
[0057] In one possible implementation, a second sound segment can also be obtained, which may contain abnormal sounds from the target work scenario. The second sound segment can be input into the abnormal sound description model to obtain a third sound description text corresponding to the second sound segment; the third sound description text can be input into the abnormal sound synthesis model to obtain a third sound segment corresponding to the third sound description text; and a second dual consistency loss can be constructed based on the difference information between the second sound segment and the third sound segment.
[0058] At this point, joint training can be performed by combining the first dual consistency loss and the second dual consistency loss.
[0059] The second audio segment mentioned above can be real abnormal audio data. The first dual consistency loss can construct text-audio-text consistency constraints, while the second dual consistency loss can construct audio-text-audio reverse consistency constraints, achieving bidirectional consistency training, which can further enhance the semantic mapping accuracy and interoperability between audio and text.
[0060] When jointly training using the first dual consistency loss and the second dual consistency loss, different weights can be assigned to them to adjust the training direction. For example, a joint loss can be constructed based on the first dual consistency loss, the second dual consistency loss, the first weight corresponding to the first dual consistency loss, and the second weight corresponding to the second dual consistency loss; then, the joint loss can be used to jointly train the abnormal sound synthesis model and the abnormal sound description model.
[0061] When constructing the joint loss, in addition to the first dual consistency loss and the second dual consistency loss, at least one of the following can be combined: semantic preservation loss, spectrogram reconstruction loss, cross-modal alignment loss, contrastive learning loss, adversarial loss, and distribution consistency loss, in order to further improve the training stability, semantic consistency, and authenticity of the generated samples of the abnormal sound synthesis model and the abnormal sound description model.
[0062] For example, the joint loss can be expressed as: ; in, For joint losses, As weight, For the first duality consistency loss, This is the second duality consistency loss. Other losses.
[0063] Once training is complete, the trained abnormal sound synthesis model can be used to generate the required abnormal sound training samples, and training of the abnormal sound detection model can begin.
[0064] In this way, a large amount of anomalous data can be generated using the trained anomalous sound synthesis model, solving the problems of difficult anomalous sound data collection and insufficient samples. At the same time, the strong consistency constraint introduced by the dual learning framework ensures the quality of the anomalous sound synthesis model. During the anomalous description preparation stage, normal sound samples and anomalous attribute labels are referenced, making the first sound description text more consistent with the actual situation and effectively reducing noise generation.
[0065] S105. The abnormal sound detection model is trained using the abnormal sound training samples to obtain a trained abnormal sound detection model; the trained abnormal sound detection model is used to detect abnormal sounds in the target work scene.
[0066] In this step, abnormal sound training samples can be input into the abnormal sound detection model to be trained. By defining a loss function, the abnormal sound detection model can continuously adjust its parameters through repeated iterations, learn the rules for distinguishing between normal and abnormal sounds, and improve the accuracy and generalization ability of the model through verification and optimization.
[0067] Because it can generate synthetic data covering a variety of abnormal attributes, anomaly sound detection models trained with this data can learn more robust feature representations and have better generalization detection capabilities for unknown types of anomalies. Users can control the type and intensity of generated abnormal sounds by modifying the anomaly description text, achieving "on-demand generation" and facilitating the training of diagnostic models for specific faults.
[0068] Among them, the abnormal sound detection models mentioned above can be deep learning models, recurrent neural networks, hybrid models, models based on self-supervised and representation learning, contrastive learning models, time-series prediction or distribution modeling models, modern large-scale models and multimodal models, etc.
[0069] In some possible implementations, the detection results output by the trained abnormal sound detection model can be obtained; based on the detection results, the abnormal information of the target work scene can be determined; the abnormal information includes at least one of abnormal type, abnormal location, and abnormal object.
[0070] In this step, the abnormal sound detection model outputs detection results that may include the presence of abnormal sounds, the time or frame location of the abnormal sound occurrence, the category probability or feature distribution of the abnormal sound, and relevant information about the target work scene. After obtaining the detection results, specific anomaly information can be determined through feature analysis or post-processing algorithms, thereby enabling safety monitoring based on the abnormal sound detection information.
[0071] In some possible implementations, video, images, or point cloud information of the target work scene can also be combined to determine abnormal information of the target work scene.
[0072] Based on the same inventive concept, this application also provides an abnormal sound detection device corresponding to the abnormal sound detection method. Since the principle of the device in this application is similar to the abnormal sound detection method described above in this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0073] See Figure 3 The diagram shown is a schematic representation of an abnormal sound detection device according to an exemplary embodiment of this application. The device includes: The acquisition module 310 is used to acquire a first sound description text; the first sound description text is used to describe the environmental sound in the target work scenario; the first sound description text contains first descriptive information for abnormal sounds. The first generation module 320 is used to input the first sound description text into the abnormal sound synthesis model to obtain a first sound segment corresponding to the first sound description text; the first sound segment contains the abnormal sound corresponding to the first description information. The second generation module 330 is used to input the first sound segment into the abnormal sound description model to obtain the second sound description text corresponding to the first sound segment; the second sound description text contains second description information for the abnormal sound in the first sound segment. The first training module 340 is used to jointly train the abnormal sound synthesis model and the abnormal sound description model based on the first sound description text and the second sound description text, and generate abnormal sound training samples based on the trained abnormal sound synthesis model. The second training module 350 is used to train the abnormal sound detection model using the abnormal sound training samples to obtain a trained abnormal sound detection model; the trained abnormal sound detection model is used to detect abnormal sounds in the target work scene.
[0074] Optionally, the first training module 340 is used for: Based on the difference information between the first sound description text and the second sound description text, a first dual consistency loss is constructed. Based on the first duality consistency loss, the abnormal sound synthesis model and the abnormal sound description model are jointly trained.
[0075] Optionally, the device further includes a third generation module 360, used for: Acquire a second sound segment; the second sound segment contains abnormal sounds from the target work scenario; The second sound segment is input into the abnormal sound description model to obtain the third sound description text corresponding to the second sound segment; The third sound description text is input into the abnormal sound synthesis model to obtain the third sound segment corresponding to the third sound description text; Based on the difference information between the second sound segment and the third sound segment, a second dual consistency loss is constructed; The first training module 340 is used for: The abnormal sound synthesis model and the abnormal sound description model are jointly trained based on the first dual consistency loss and the second dual consistency loss.
[0076] Optionally, the first training module 340 is used for: Based on the first dual consistency loss, the second dual consistency loss, the first weight corresponding to the first dual consistency loss, and the second weight corresponding to the second dual consistency loss, a joint loss is constructed. The abnormal sound synthesis model and the abnormal sound description model are jointly trained using the joint loss.
[0077] Optionally, the device further includes a fourth generation module 370, used for: Obtain a fourth sound segment and any abnormal sound type label; the fourth sound segment contains normal sounds under the target work scenario. The fourth sound segment and the abnormal sound type label are input into the abnormal sound description model to obtain the fourth sound description text that matches the fourth sound segment and the abnormal sound type label; the fourth sound description text contains the description information corresponding to the abnormal sound type label; The fourth sound description text is used as the first sound description text.
[0078] Optionally, the abnormal sound type label includes at least one of the following: Friction, impact, fracture, specific frequency range, explosion.
[0079] Optionally, the device further includes an analysis module 380, used for: Obtain the detection results output by the trained abnormal sound detection model; Based on the detection results, abnormal information of the target work scenario is determined; the abnormal information includes at least one of abnormal type, abnormal location, and abnormal object.
[0080] The abnormal sound detection device provided in this application embodiment can form a dual consistency loss based on sound description text and sound segments, and jointly train the abnormal sound synthesis model and the abnormal sound description model, realizing bidirectional constraints of semantics and sound. Thus, in the absence of real abnormal samples, a large-scale and controllable abnormal sound training dataset can be constructed. At the same time, the dual consistency loss is constructed based on the difference information between the first sound description text and the second sound description text generated by the model, so that the model maintains semantic consistency in both synthesis and description directions, which helps to generate sound samples that are more consistent with abnormal characteristics and improves the quality of abnormal sound generation.
[0081] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.
[0082] This application also provides a computer device, such as... Figure 4 The diagram shown is a schematic representation of a computer device structure according to an exemplary embodiment of this application. The computer device includes: Processor 41 and memory 42; the memory 42 stores machine-readable instructions executable by the processor 41, and the processor 41 executes the machine-readable instructions stored in the memory 42. When the machine-readable instructions are executed by the processor 41, the processor 41 performs the following steps: Obtain a first sound description text; the first sound description text is used to describe the environmental sound in the target work scenario; the first sound description text contains first descriptive information for abnormal sounds; The first sound description text is input into the abnormal sound synthesis model to obtain the first sound segment corresponding to the first sound description text; the first sound segment contains the abnormal sound corresponding to the first description information. The first sound segment is input into the abnormal sound description model to obtain the second sound description text corresponding to the first sound segment; the second sound description text contains second description information for the abnormal sound in the first sound segment. Based on the first sound description text and the second sound description text, the abnormal sound synthesis model and the abnormal sound description model are jointly trained, and abnormal sound training samples are generated based on the trained abnormal sound synthesis model. The abnormal sound detection model is trained using the abnormal sound training samples to obtain a trained abnormal sound detection model; the trained abnormal sound detection model is used to detect abnormal sounds in the target work scene.
[0083] Optionally, the step of jointly training the abnormal sound synthesis model and the abnormal sound description model based on the first sound description text and the second sound description text includes: Based on the difference information between the first sound description text and the second sound description text, a first dual consistency loss is constructed. Based on the first duality consistency loss, the abnormal sound synthesis model and the abnormal sound description model are jointly trained.
[0084] Optionally, processor 41 is also used to perform: Acquire a second sound segment; the second sound segment contains abnormal sounds from the target work scenario; The second sound segment is input into the abnormal sound description model to obtain the third sound description text corresponding to the second sound segment; The third sound description text is input into the abnormal sound synthesis model to obtain the third sound segment corresponding to the third sound description text; Based on the difference information between the second sound segment and the third sound segment, a second dual consistency loss is constructed; The step of jointly training the abnormal sound synthesis model and the abnormal sound description model based on the first sound description text and the second sound description text includes: The abnormal sound synthesis model and the abnormal sound description model are jointly trained based on the first dual consistency loss and the second dual consistency loss.
[0085] Optionally, the step of jointly training the abnormal sound synthesis model and the abnormal sound description model based on the first sound description text and the second sound description text includes: Based on the first dual consistency loss, the second dual consistency loss, the first weight corresponding to the first dual consistency loss, and the second weight corresponding to the second dual consistency loss, a joint loss is constructed. The abnormal sound synthesis model and the abnormal sound description model are jointly trained using the joint loss.
[0086] Optionally, processor 41 is also used to perform: Obtain a fourth sound segment and any abnormal sound type label; the fourth sound segment contains normal sounds under the target work scenario. The fourth sound segment and the abnormal sound type label are input into the abnormal sound description model to obtain the fourth sound description text that matches the fourth sound segment and the abnormal sound type label; the fourth sound description text contains the description information corresponding to the abnormal sound type label; The fourth sound description text is used as the first sound description text.
[0087] Optionally, the abnormal sound type label includes at least one of the following: Friction, impact, fracture, specific frequency range, explosion.
[0088] Optionally, processor 41 is also used to perform: Obtain the detection results output by the trained abnormal sound detection model; Based on the detection results, abnormal information of the target work scenario is determined; the abnormal information includes at least one of abnormal type, abnormal location, and abnormal object.
[0089] The aforementioned memory 42 includes a main memory 421 and an external memory 422; the main memory 421, also known as internal memory, is used to temporarily store the computational data in the processor 41, as well as the data exchanged with external memory 422 such as a hard disk. The processor 41 exchanges data with the external memory 422 through the main memory 421.
[0090] The specific execution process of the above instructions can be referred to the steps of the abnormal sound detection method described in the embodiments of this application, and will not be repeated here.
[0091] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0092] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the abnormal sound detection method described in the above-described method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0093] This application also provides a computer program product, including a computer program / instruction, which, when executed by the computer program / instruction processor, implements the abnormal sound detection method provided in the various embodiments of this application.
[0094] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0095] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0096] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0097] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0098] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0099] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0100] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for detecting abnormal sounds, characterized in that, The method includes: Obtain a first sound description text; the first sound description text is used to describe the environmental sound in the target work scenario; the first sound description text contains first descriptive information for abnormal sounds; The first sound description text is input into the abnormal sound synthesis model to obtain the first sound segment corresponding to the first sound description text; the first sound segment contains the abnormal sound corresponding to the first description information. The first sound segment is input into the abnormal sound description model to obtain the second sound description text corresponding to the first sound segment; the second sound description text contains second description information for the abnormal sound in the first sound segment. Based on the first sound description text and the second sound description text, the abnormal sound synthesis model and the abnormal sound description model are jointly trained, and abnormal sound training samples are generated based on the trained abnormal sound synthesis model. The abnormal sound detection model is trained using the abnormal sound training samples to obtain a trained abnormal sound detection model; the trained abnormal sound detection model is used to detect abnormal sounds in the target work scene.
2. The method according to claim 1, characterized in that, The step of jointly training the abnormal sound synthesis model and the abnormal sound description model based on the first sound description text and the second sound description text includes: Based on the difference information between the first sound description text and the second sound description text, a first dual consistency loss is constructed. Based on the first duality consistency loss, the abnormal sound synthesis model and the abnormal sound description model are jointly trained.
3. The method according to claim 2, characterized in that, The method further includes: Acquire a second sound segment; the second sound segment contains abnormal sounds from the target work scenario; The second sound segment is input into the abnormal sound description model to obtain the third sound description text corresponding to the second sound segment; The third sound description text is input into the abnormal sound synthesis model to obtain the third sound segment corresponding to the third sound description text; Based on the difference information between the second sound segment and the third sound segment, a second dual consistency loss is constructed; The step of jointly training the abnormal sound synthesis model and the abnormal sound description model based on the first sound description text and the second sound description text includes: The abnormal sound synthesis model and the abnormal sound description model are jointly trained based on the first dual consistency loss and the second dual consistency loss.
4. The method according to claim 3, characterized in that, The step of jointly training the abnormal sound synthesis model and the abnormal sound description model based on the first sound description text and the second sound description text includes: Based on the first dual consistency loss, the second dual consistency loss, the first weight corresponding to the first dual consistency loss, and the second weight corresponding to the second dual consistency loss, a joint loss is constructed. The abnormal sound synthesis model and the abnormal sound description model are jointly trained using the joint loss.
5. The method according to claim 1, characterized in that, The method further includes: Obtain a fourth sound segment and any abnormal sound type label; the fourth sound segment contains normal sounds under the target work scenario. The fourth sound segment and the abnormal sound type label are input into the abnormal sound description model to obtain the fourth sound description text that matches the fourth sound segment and the abnormal sound type label; the fourth sound description text contains the description information corresponding to the abnormal sound type label; The fourth sound description text is used as the first sound description text.
6. The method according to claim 5, characterized in that, The abnormal sound type label includes at least one of the following: Friction, impact, fracture, specific frequency range, explosion.
7. The method according to claim 1, characterized in that, The method further includes: Obtain the detection results output by the trained abnormal sound detection model; Based on the detection results, abnormal information of the target work scenario is determined; the abnormal information includes at least one of abnormal type, abnormal location, and abnormal object.
8. An abnormal sound detection device, characterized in that, The device includes: The acquisition module is used to acquire a first sound description text; the first sound description text is used to describe the environmental sound in the target work scenario; the first sound description text contains first descriptive information for abnormal sounds. The first generation module is used to input the first sound description text into the abnormal sound synthesis model to obtain a first sound segment corresponding to the first sound description text; the first sound segment contains the abnormal sound corresponding to the first description information. The second generation module is used to input the first sound segment into the abnormal sound description model to obtain the second sound description text corresponding to the first sound segment; the second sound description text contains second description information for the abnormal sound in the first sound segment. The first training module is used to jointly train the abnormal sound synthesis model and the abnormal sound description model based on the first sound description text and the second sound description text, and generate abnormal sound training samples based on the trained abnormal sound synthesis model. The second training module is used to train the abnormal sound detection model using the abnormal sound training samples to obtain a trained abnormal sound detection model; the trained abnormal sound detection model is used to detect abnormal sounds in the target work scene.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7.