Sound event detection data synthesis and sound event detection model training method

By generating audio data through semantic prompts and utilizing large-scale language models and audio generation models, the problem of data acquisition difficulties in training sound event detection models is solved, enabling efficient and low-cost generation of diverse audio data and improving the efficiency and generalization ability of model training.

CN120877699APending Publication Date: 2025-10-31SHANGHAI NORMAL UNIVERSITY +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510684973.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

The high-quality data that sound event detection models rely on for training is difficult to obtain cost-effectively, and the limited data resources restrict the model's generalization ability.

Method used

Synthetic audio data is generated through semantic prompts. A large-scale language model is used to generate structured semantic prompts to guide the audio generation model to synthesize audio data that matches the semantic description information. The labels of the synthesized audio data are automatically determined, and semantic constraint rules are combined to ensure the diversity and accuracy of the generated data.

Benefits of technology

It enables the efficient and low-cost generation of diverse, high-quality audio data, breaking through the efficiency bottleneck of traditional manual annotation, significantly improving the efficiency of data construction and large-scale production capabilities, and providing an efficient and reliable data source for training sound event detection models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877699A_ABST
    Figure CN120877699A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic prompt-based sound event detection data synthesis and sound event detection model training method. The sound event detection task is converted into the semantic description information, and the structured semantic prompt instruction which accurately reflects the target sound event characteristics is generated in combination with the semantic constraint rule, so that automatic mapping from semantic description to instruction generation is realized, and the manual intervention cost is reduced. And inputting the structured semantic prompt instruction into the audio generation model, and synthesizing the audio data in a large scale, thereby reducing the data acquisition cost and improving the sample diversity and expandability. The structured semantic prompt instruction can guide the model to synthesize audios of various sound event types in batches, and is automatically generated by a large language model to ensure that the synthesized audios are strictly aligned with instruction semantics. When the label is generated, the sample event type can be obtained without manual labeling, and an efficient and reliable data source is provided for sound event detection model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and in particular to a method for synthesizing sound event detection data based on semantic prompts and training a sound event detection model. Background Technology

[0002] With the rapid development of machine learning and neural network technologies and the rise of artificial intelligence applications, intelligent voice technology has gradually penetrated into all aspects of people's daily lives, covering fields such as Sound Event Detection (SED), speech recognition, and speech enhancement. Among these, Sound Event Detection (SED), as one of the research hotspots in the field of acoustics in recent years, aims to identify and locate specific sound events from environmental audio, such as alarms, voices, dog barks, and the sound of breaking glass. Compared to traditional speech or music recognition, Sound Event Detection faces more complex challenges—its input is usually multi-source, unstructured, and rich in background noise from real-world audio scenarios, which places higher demands on the model's generalization ability and robustness. Currently, this technology has been widely applied in smart homes, autonomous driving, environmental monitoring, and many other fields.

[0003] In recent years, breakthroughs in deep learning have brought a new paradigm to sound event detection. Models based on architectures such as Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and Convolutional Recurrent Neural Networks (CRNNs) have been widely used. Meanwhile, the introduction of attention mechanisms and transformers has further enhanced the models' ability to model temporal features. Furthermore, the integration of semi-supervised learning, weakly supervised learning, and self-supervised learning methods has effectively alleviated the problem of insufficient training data.

[0004] In the process of model and algorithm iteration, high-quality training data has always been the core driving force for performance improvement. Taking the Detection and Classification of Acoustic Scenes and Events (DCASE) as an example, its standardized dataset provides researchers with a unified evaluation platform through a dual annotation system that includes strong labels (precisely annotating the start and end times of events) and weak labels (only annotating the type of sound event). However, the collection and annotation of large-scale data in real-world scenarios faces significant challenges: high labor costs, label consistency problems, and a lack of rare event samples make obtaining diverse, balanced, and realistic training data at low cost a technical bottleneck, directly restricting further improvements in model performance.

[0005] The rise of generative models has provided new pathways to solving data challenges. Audio generation techniques based on diffusion models, generative adversarial networks (GANs), and variational autoencoders (VAEs) can synthesize high-fidelity, semantically consistent audio data. Meanwhile, large-scale language models such as ChatGPT and GPT-4, through natural language prompts, have achieved semantically driven content generation, offering the possibility of cross-modal fusion for data generation.

[0006] In summary, the field of sound event detection still faces key challenges: the high-quality data upon which sound event detection models rely for training is difficult to obtain cost-effectively, and the limited availability of data resources restricts the generalization ability of the models. Summary of the Invention

[0007] This application provides a method, apparatus, device, and medium for synthesizing sound event detection data and training a sound event detection model based on semantic prompts, which solves the problem that the high-quality data on which existing sound event detection models rely for training is difficult to obtain economically and efficiently, and the limited generalization ability of the model is restricted due to the limited data resources.

[0008] Firstly, this application provides a method for synthesizing sound event detection data based on semantic prompts, the method comprising:

[0009] Obtain semantic description information for a sound event detection task; wherein the semantic description information includes at least one sound event type and semantic constraint rules pre-configured for the at least one sound event type;

[0010] Using a pre-trained large-scale language model, structured semantic prompts are generated based on the semantic description information, corresponding to each of the at least one sound event type; wherein, the structured semantic prompts for any sound event type are used to guide the generation of the corresponding type of audio.

[0011] Using a pre-trained audio generation model, synthetic audio data that matches the semantic description information is generated based on each of the structured semantic prompts.

[0012] For each of the synthesized audio data, the sound event type label corresponding to the synthesized audio data is determined based on the structured semantic prompt instruction corresponding to the synthesized audio data.

[0013] Secondly, this application also provides a method for training a sound event detection model, the method comprising:

[0014] Obtain any audio sample in the sample set and its corresponding real sound event detection label; wherein, the audio sample includes strongly labeled audio samples and weakly labeled audio samples, the weakly labeled audio samples include synthesized audio data obtained based on any of the sound event detection data synthesis methods, the real sound event detection label corresponding to the strongly labeled audio sample includes a real sound event type label and a real time detection label, and the real sound event detection label corresponding to the weakly labeled audio sample includes a real sound event type label;

[0015] Based on the audio samples, the predicted sound event detection results corresponding to the audio samples are obtained using the original sound event detection model.

[0016] Based on the predicted sound event detection results and the actual sound event detection labels, the original sound event detection model is trained to obtain a trained sound event detection model.

[0017] Thirdly, this application also provides a sound event detection data synthesis apparatus based on semantic prompts, the apparatus comprising:

[0018] An acquisition unit is used to acquire semantic description information of a sound event detection task; wherein, the semantic description information includes at least one sound event type and semantic constraint rules pre-configured for the at least one sound event type;

[0019] The structured semantic generation unit is used to generate structured semantic prompts corresponding to the at least one sound event type based on the semantic description information using a pre-trained large-scale language model; wherein, the structured semantic prompts for any sound event type are used to guide the generation of audio of the corresponding type;

[0020] The audio generation unit is used to generate synthetic audio data that matches the semantic description information based on each of the structured semantic prompt instructions, using a pre-trained audio generation model.

[0021] The tag generation unit is used to determine the sound event type tag corresponding to each synthesized audio data based on the structured semantic prompt instruction corresponding to the synthesized audio data.

[0022] Fourthly, this application also provides a training device for a sound event detection model, the device comprising:

[0023] The acquisition module is used to acquire any audio sample in the sample set and its corresponding real sound event detection label; wherein, the audio sample includes strongly labeled audio samples and weakly labeled audio samples, the weakly labeled audio samples include synthesized audio data obtained based on the sound event detection data synthesis method described above, the real sound event detection label corresponding to the strongly labeled audio sample includes a real sound event type label and a real time detection label, and the real sound event detection label corresponding to the weakly labeled audio sample includes a real sound event type label;

[0024] The processing module is used to obtain the predicted sound event detection result corresponding to the audio sample based on the audio sample using the original sound event detection model;

[0025] The training module is used to train the original sound event detection model based on the predicted sound event detection results and the real sound event detection labels to obtain the trained sound event detection model.

[0026] Fifthly, this application provides a computer device including a processor, which executes a computer program stored in a memory to implement the steps of the semantic prompt-based sound event detection data synthesis method described above, or to implement the steps of the sound event detection model training method described above.

[0027] Sixthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the semantic prompt-based sound event detection data synthesis method described above, or implements the steps of the sound event detection model training method described above.

[0028] The beneficial effects of this application are as follows:

[0029] 1. By transforming sound event detection tasks into semantic description information and combining them with semantic constraint rules, the generated structured semantic prompts accurately reflect the characteristics of the target sound events, ensuring a high degree of match between the synthesized audio data and task requirements, thus improving the data's ability to simulate real-world scenarios. Simultaneously, it achieves automated mapping from semantic description to instruction generation, replacing the traditional method of relying on manual design of audio generation rules, reducing manual intervention costs, and making it suitable for batch processing of various types of sound events.

[0030] 2. By inputting structured semantic prompts into the audio generation model, efficient and controllable audio data synthesis can be achieved. Compared with traditional field collection and manual recording, this significantly reduces data acquisition costs and substantially improves the diversity and scalability of sample generation.

[0031] 3. Structured semantic prompts directly guide the audio generation model, ensuring that the generated audio content strictly follows the event types and constraint rules in the semantic description information, and realizing the batch synthesis of multiple sound event types in the semantic description information.

[0032] 4. Structured semantic prompts are automatically generated by a large language model, containing clear event type identifiers, and the synthesized audio data is strictly aligned with the semantics of the instructions. This not only enables the acquisition of a large amount of high-quality data required for training sound event detection models, reducing acquisition costs, but also allows for the direct acquisition of sample event categories during label generation without manual annotation, breaking through the efficiency bottleneck of traditional manual annotation, significantly improving the efficiency of data construction and scalable production capabilities, and providing an efficient and reliable data source for training sound event detection models. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 A schematic diagram illustrating the process of synthesizing sound event detection data based on semantic prompts, provided in an embodiment of this application;

[0035] Figure 2 A schematic diagram illustrating the specific semantic prompt-based sound event detection data synthesis process provided in this application embodiment;

[0036] Figure 3 A schematic diagram illustrating the training process of a sound event detection model provided in an embodiment of this application;

[0037] Figure 4A schematic diagram of a sound event detection data synthesis device based on semantic prompting provided in this application embodiment;

[0038] Figure 5 A schematic diagram of a training device for a sound event detection model provided in an embodiment of this application;

[0039] Figure 6 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of this application. Detailed Implementation

[0040] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0041] In order to obtain a large amount of high-quality data for training sound event detection models and reduce the cost of obtaining high-quality data, this application provides a method, apparatus, device and medium for synthesizing sound event detection data based on semantic prompts and training sound event detection models.

[0042] Example 1:

[0043] This application provides a method for synthesizing sound event detection data based on semantic prompts. Figure 1 A schematic diagram of a sound event detection data synthesis process based on semantic prompts provided in this application embodiment is shown. The process includes:

[0044] S101: Obtain semantic description information for the sound event detection task; wherein, the semantic description information includes at least one sound event type and semantic constraint rules pre-configured for the at least one sound event type.

[0045] In this application, the semantic prompt-based sound event detection data synthesis method is applied to a computer device, which can be a smart terminal, such as a computer or a robot, or a server, such as an application server or a business server.

[0046] Training sound event detection models requires a large amount of audio data that closely matches real-world application scenarios. However, data collection in real-world scenarios is costly and has limited coverage. For example, collecting traffic noise data under specific weather conditions or during specific time periods often requires significant investment of manpower, resources, and time, and obtaining complete scene samples is difficult. Therefore, this application incorporates semantic description information for sound event detection tasks. This semantic description information provides clear data synthesis directions and specifications for computer devices, enabling the synthesized data to more accurately match training requirements.

[0047] The semantic description information includes at least one sound event type and pre-configured semantic constraint rules for that sound event type. Specifically, the sound event type in the semantic description information is the core element of data synthesis, determining the main content of the final generated audio. For example, in a security monitoring scenario, the sound event type might be the sound of breaking glass or picking a lock; in a smart home scenario, it might be the sound of flowing water or a gas leak alarm. Based on the determined sound event type and considering the diverse needs of the scenario, corresponding semantic constraint rules are pre-configured to standardize the generated content of subsequent semantic prompts, ensuring that the generated audio data possesses both semantic control and diversity. For example, in the dimension of event sound source characteristics, the volume range and repetition of a sound event are specified; in the dimension of background environment type and interference level, the background environment used in synthesizing the audio and the degree of interference of background noise on the target sound are determined; in the dimension of audio data volume, the number of audio data points generated for each sound event type is specified. These rules act like audio production standards, guiding computer equipment to synthesize high-quality, diverse audio data.

[0048] It should be noted that the semantic description information includes one or more sound event types. That is, the computer device can generate audio data for a certain sound event type, such as only containing "snoring" or "Infantcry", or it can generate audio data in batches for multiple sound event types, such as containing both "snoring" and "Infantcry".

[0049] S102: Using a pre-trained large-scale language model, based on the semantic description information, generate structured semantic prompts corresponding to each of the at least one sound event type; wherein, the structured semantic prompts for any sound event type are used to guide the generation of audio of the corresponding type.

[0050] After obtaining semantic description information based on the above embodiments, this semantic description information can be input into a pre-trained large-scale language model, such as ChatGPT or GPT-4. The pre-trained large-scale language model processes the input semantic description information to generate structured semantic prompts corresponding to at least one type of sound event in the semantic description information. These structured semantic prompts describe the required sound events and their scene features in a structured format. The prompt content can automatically combine background types and semantic elements for different sound events, thereby generating semantic descriptions that are grammatically correct, logically coherent, and rich in scene details. These prompts, expressed in natural language, cover a wide range of scene and sound event combinations, possess semantic controllability and diversity, and can directly drive subsequent audio data synthesis.

[0051] It should be noted that this large-scale language model is a large language model pre-trained on massive amounts of text and has natural language processing capabilities. The specific training process will not be elaborated here.

[0052] This structured semantic prompting instruction generation mechanism can efficiently produce hundreds or thousands of text descriptions with high semantic consistency, grammatical rationality, and control accuracy, greatly saving the cost of manually designing audio descriptions and enhancing control capabilities and sample diversity in the data generation process.

[0053] S103: Using a pre-trained audio generation model, synthesized audio data that matches the semantic description information is generated according to each of the structured semantic prompts.

[0054] After obtaining the structured semantic prompts based on the above embodiments, the obtained structured semantic prompts can be input into an audio generation model, such as AudioGEN, which has natural language-driven audio generation capabilities. By processing the input structured semantic prompts through this audio generation model, synthesized audio data that matches the semantic description information can be obtained. This audio generation model has the ability to convert natural language descriptions into corresponding audio data, can better restore the sound event features in the natural language description, and can batch process structured semantic prompts and save the audio results.

[0055] For example, let P be the set containing various structured semantic prompt instructions, and let p be any structured semantic prompt instruction. i Input ∈P into a pre-trained audio generation model, and output synthesized audio data x. i The audio data is stored in .wav file format, thus forming a set X containing various synthesized audio data. This audio data set X = {x1, x2, ..., x...} N}, where N represents the total number of sound event types, x NThis represents the synthesized audio data corresponding to the Nth sound event type.

[0056] In one possible implementation, the audio generation model supports English natural language input, meaning it supports structured semantic prompts in English. For example, "Snoring: A person snoring loudly in a quiet room."

[0057] The audio generation model can synthesize corresponding audio segments based on the sound event type and background information in the structured semantic prompts, ensuring that the target sound event is clearly expressed and the background environment is reasonable. The generated audio does not rely on real recording equipment, and its content is highly consistent with the semantic prompts. It can construct the audio data required for training in a structured and efficient manner in batches, making it suitable for covering scarce scenarios or increasing the number of task samples.

[0058] It should be noted that the training process of this audio generation model can refer to existing technologies, but will not be elaborated on here.

[0059] S104: For each of the synthesized audio data, determine the sound event type label corresponding to the synthesized audio data based on the structured semantic prompt instruction corresponding to the synthesized audio data.

[0060] In training sound event detection models, the audio data used for training typically corresponds to specific labels, with the label system including both strong and weak labels. Strong labels encompass both sound event type labels and time detection labels, while weak labels only include sound event type labels. After obtaining each synthesized audio data through the aforementioned embodiments, the automatic generation of weak labels can be achieved using a pre-defined rule mapping mechanism of structured semantic prompts. Specifically, for each synthesized audio data, the sound event type label corresponding to that audio data is accurately determined based on its corresponding structured semantic prompt. Since the structured semantic prompts are automatically generated by a large language model, they are essentially a structured description of the target sound event, containing a clear event type identifier. The synthesized audio data, as the physical implementation of the instruction-driven audio generation model, must have its content strictly aligned with the semantics of the structured semantic prompts. This characteristic allows for direct acquisition of the event category corresponding to each sample during label generation without relying on manual annotation, fundamentally breaking through the efficiency bottleneck of traditional manual annotation, significantly improving the efficiency of data construction and scalable production capabilities, and providing a more efficient and reliable data source support for training sound event detection models.

[0061] In one possible implementation, determining the sound event type label corresponding to each synthesized audio data based on the structured semantic prompt instruction corresponding to the synthesized audio data includes:

[0062] For each synthesized audio data, the sound event type label corresponding to the synthesized audio data is determined according to the sound event type identifier at a preset position in the structured semantic prompt instruction corresponding to the synthesized audio data.

[0063] Because structured semantic prompts follow specific format specifications during generation, their internal descriptions of sound event types are typically confined to fixed fields or specific locations (e.g., at the beginning, key phrases, or end of the instruction). These pre-defined event type identifiers, semantically calibrated by a large language model, possess clear classification orientation. Therefore, by locating and reading the identifier content at these pre-defined locations, they can be directly mapped to the sound event type labels corresponding to the synthesized audio data. For example, for each piece of synthesized audio data, the corresponding structured semantic prompt is parsed first, and the sound event type identifier at the pre-defined location is extracted. This method of extracting identifiers based on pre-defined locations, leveraging the standardized design of the instruction format, achieves the regularization and automation of label generation logic, further improving the accuracy and processing efficiency of label allocation, and avoiding label errors caused by semantic ambiguity of instructions or differences in human interpretation.

[0064] For example, since structured semantic prompts are generated under structural control through a large-scale language model, a fixed position within the structured semantic prompt can be controlled to represent the sound event type. Based on this structure, rule extraction can be used to automatically identify the sound event type name as a label, constructing a weak label set Y = {y} for each synthesized audio data. i},in:

[0065] y i =extract_event(p i )

[0066] The `extract_event()` function is a tag extraction function based on regular expressions or keyword matching. It can extract the sound event type as a weak tag from a fixed position in a structured semantic prompt instruction. For example, inputting the structured semantic prompt instruction "Snoring:A person snoringloudlyin a quiet room." can automatically extract the sound event type "Snoring," and then match it with the corresponding synthesized audio data x. i Construct training data pairs (x i ,y i These are used together to support the subsequent model training process.

[0067] The beneficial effects of this application are as follows:

[0068] 1. By transforming sound event detection tasks into semantic description information and combining them with semantic constraint rules, the generated structured semantic prompts accurately reflect the characteristics of the target sound events, ensuring a high degree of match between the synthesized audio data and task requirements, thus improving the data's ability to simulate real-world scenarios. Simultaneously, it achieves automated mapping from semantic description to instruction generation, replacing the traditional method of relying on manual design of audio generation rules, reducing manual intervention costs, and making it suitable for batch processing of various types of sound events.

[0069] 2. By inputting structured semantic prompts into the audio generation model, efficient and controllable audio data synthesis can be achieved. Compared with traditional field collection and manual recording, this significantly reduces data acquisition costs and substantially improves the diversity and scalability of sample generation.

[0070] 3. Structured semantic prompts directly guide the audio generation model, ensuring that the generated audio content strictly follows the event types and constraint rules in the semantic description information, and realizing the batch synthesis of multiple sound event types in the semantic description information.

[0071] 4. Structured semantic prompts are automatically generated by a large language model, containing clear event type identifiers, and the synthesized audio data is strictly aligned with the semantics of the instructions. This not only enables the acquisition of a large amount of high-quality data required for training sound event detection models, reducing acquisition costs, but also allows for the direct acquisition of sample event categories during label generation without manual annotation, breaking through the efficiency bottleneck of traditional manual annotation, significantly improving the efficiency of data construction and scalable production capabilities, and providing an efficient and reliable data source for training sound event detection models.

[0072] Example 2:

[0073] To ensure that the generated audio data meets the training requirements of the sound event detection task, based on the above embodiments, in this application, the semantic constraint rules include, but are not limited to, one or more of the following: event sound source characteristics, background sound environment type, background sound interference level, scene spatial attributes, temporal structure features, audio data volume, and clarity.

[0074] For event sound source characteristics, this describes the physical and dynamic attributes of the event sound source, such as whether the volume is prominent and whether the event sound source repeats. For background sound environment type, this defines the background sound introduction strategy, type combination, and temporal relationship. Specifically, it needs to be clarified whether to introduce background sound and what type of background sound to use (such as human voices, music, appliance operation sounds, etc.). Taking a home environment as an example, it can be specified that when synthesizing the sound of flowing water in the kitchen scene, background sounds of appliances such as microwave oven beeping and refrigerator running can be introduced; in the living room scene, elements such as television playback and air conditioner operation can be added. At the same time, it can also be clarified whether the background sound overlaps with the event sound source, so as to construct a more realistic acoustic scene. For background sound interference level, this describes the degree of influence of the background sound on the target sound source. Different levels of interference intensity can be set, such as light, moderate, and heavy, corresponding to different background sound intensity ranges and component combinations. For example, in an airport security scenario, when the target sound is the metal detector alarm, mild interference only includes conversations of passengers in the distance; moderate interference adds announcements and the sound of luggage cart wheels; and severe interference incorporates the roar of aircraft taking off and landing. Simultaneously, to avoid background noise interfering with the accuracy of the model's training labels, all background noise is excluded from the labeling scope, ensuring the model focuses on learning the features of the target sound event. For scene spatial attributes, the scene environment is defined, such as whether it is indoors or outdoors, and the basic level of environmental noise, such as quiet or noisy, giving the synthesized audio spatial acoustic characteristics consistent with the actual scene. For temporal structure features, the dynamic performance of the sound event over time is standardized, such as whether the event source appears periodically and the duration of the event source, to make the synthesized audio more closely resemble the temporal evolution of the actual event. For audio data volume, the amount of audio data required to generate the sound event is specified. For clarity, the sound events in the generated audio must be clearly distinguishable to provide high-quality data samples for subsequent model training.

[0075] By combining one or more of the above data, a complete semantic constraint rule system is formed, providing comprehensive and detailed guidance for data synthesis.

[0076] In one example, the semantic constraint rules may originate from one or more of the following: a general semantic constraint rule set and a personalized semantic constraint rule set. The general semantic constraint rule set includes unified semantic constraint rules applicable to all sound event types, while the personalized semantic constraint rule set includes exclusive semantic constraint rules applicable to specific sound event types. The existence of the general semantic constraint rule set greatly improves the basic efficiency and standardization of data synthesis. It covers universal rules for all sound event types, ensuring that data synthesis follows these unified standards regardless of the sound event type. This avoids the tedious process of repeatedly formulating basic rules and ensures the consistency of basic parameters in synthesized data from different types of sound events, providing a stable basic data environment for subsequent model training. The personalized semantic constraint rule set, on the other hand, is a customized specification for specific sound event types based on different industries, scenarios, and specific research needs.

[0077] For example, a set of sound event types is constructed based on the types of sound events of interest in the sound event detection task. For instance, C = {c1, c2, ..., c...} N}, where C represents the set of sound event types, N is the total number of sound event types, and c N This represents the Nth sound event type, the specific type of which is determined based on the actual task requirements. A corresponding set of semantic generation rule constraints is configured for this set of sound event types. This stage aims to provide a unified semantic framework and generation requirements for the subsequent construction of structured semantic prompts and audio generation processes. The two types of semantic generation rule constraint sets are explained below:

[0078] I. A universal semantic constraint rule set, meaning that all sound event types in the set of sound event types use the same semantic constraint rules. This universal semantic constraint rule set can be defined as: R = {r1, r2, ..., r...} M}, where r m Let R represent the m-th uniform semantic constraint rule, for example: R = {event sound is clearly identifiable, volume is prominent, background sound is indoor human speech, number of generated rules is 100 per category}. This set of general semantic constraint rules is suitable for generating structured semantic prompts with consistent quality requirements and expression style, and can be batch reused in the generation process of all sound event types.

[0079] II. Personalized semantic constraint rule set: This refers to setting specific semantic constraint rules for particular sound event types. This personalized semantic constraint rule set can be defined as a two-dimensional mapping relationship between sound event types and semantic constraint rules: R = {r ij |i∈[1,N],j∈[1,M i ]}, where r ij This indicates the type of sound event, c.i The j-th exclusive semantic constraint rule, M i This represents the number of semantic constraint rules specific to the i-th sound event type. For example: c1 = Snoring, corresponding to {r 11 =High loudness, r 12 =The background sound is television sound, r 13 =Nighttime bedroom scene}; c2 = Infantcry, corresponding to {r 21 =Increasing sound intensity, r 22 =No background noise, r 23 =Daytime indoor environment}. This refined rule mechanism enables targeted semantic control, improving the consistency and task fit between structured semantic prompts and audio synthesis. Combining the aforementioned set of sound event types with the set of semantic generation rule constraints, each set of sound event types and its corresponding semantic generation rule can be integrated into a semantic description.

[0080] Example 3:

[0081] To improve the quality of the generated structured semantic prompts, and consequently improve the quality of the audio data subsequently generated based on these prompts, in this application, based on the above embodiments, the step of generating structured semantic prompts corresponding to the at least one sound event type using a pre-trained large-scale language model and the semantic description information includes:

[0082] Using the large-scale language model, based on the semantic description information and pre-configured syntactic structure template prompts, structured semantic prompt instructions with consistent syntactic structures are generated.

[0083] Large-scale language models such as ChatGPT and GPT-4, while possessing powerful language understanding and generation capabilities, often exhibit diverse and uncertain output text formats due to the complexity of their training data and the flexibility of their generation logic, especially without explicit format guidance. For example, given the same semantic description information, the model may generate text descriptions with different sentence structures and varying information order. Furthermore, the synthesis of sound event detection data requires a large number of semantically complete and uniformly formatted prompts to ensure that the subsequent audio generation module can efficiently and accurately read the prompts and perform standardized audio synthesis operations. Therefore, this application pre-configures syntactic structure template prompts. These templates are based on natural language sentence structures and reserve positions for replaceable variables. For example, "c..." i "Natural Language Audio Description", where c iThis represents the i-th sound event type. This template design transforms abstract semantic descriptions into structured expressions. This semantic description, along with a pre-configured syntactic structure template, can then be input into a large-scale language model. Upon receiving the input, the large-scale language model first identifies the template's sentence structure, understanding the semantic meaning and logical relationships of each variable. Based on its knowledge of language patterns and semantic associations learned through training on massive amounts of text, the model reorganizes and optimizes the semantic descriptions filling the variable positions according to the template's defined sentence structure and logic. In generating structured semantic prompts corresponding to at least one sound event type, the model also incorporates its understanding of language expression habits and grammatical rules to ensure that the generated structured semantic prompts are not only syntactically consistent but also fluent, clear, and semantically consistent with the input semantic description.

[0084] For example, semantic description information and pre-configured syntactic structure templates are input into a large-scale language model for processing. Based on the semantic description information, this large-scale language model automatically generates a set of structured semantic prompts that match the semantic description information, which can be represented as: P = {p1, p2, ..., p...} K}, where each structured semantic prompt instruction has the following format: p k =“c i "Natural Language Audio Descriptions". For example, if a unified semantic constraint rule is used, the following example of a structured semantic prompt instruction can be generated: "Please generate 100 natural language audio descriptions for each of the two events, Snoring and InfantCry. The event sounds must be clear and loud, and limited background noise is allowed, consisting of indoor conversations." If a specific semantic constraint rule is used, the sound event types in the semantic description information will be combined one by one. The following example of a structured semantic prompt instruction can be generated: Snoring: "Please generate 100 audio descriptions about 'Snoring,' requiring high event loudness, allowing background television sound, and the scene being a nighttime bedroom." Infantcry: "Please generate 50 audio descriptions about 'baby crying,' requiring increasing sound intensity, no other background noise, and occurring in a daytime indoor environment."

[0085] Because structured semantic prompts are highly standardized and accurate in both format and content, audio generation models can quickly parse the parameters in the prompts and generate audio data that meets the training requirements of the sound event detection model according to the structured semantic prompts. This instruction generation mechanism based on structured semantic prompts fully leverages the generation advantages of large-scale language models while effectively avoiding the problem of inconsistent output formats, greatly improving the efficiency and quality of sound event detection data synthesis. Furthermore, determining the sound event type label based on the structured semantic prompts that correspond one-to-one with the synthesized audio data becomes extremely direct and efficient. Since the structured semantic prompts explicitly contain sound event type identifiers, each audio data point can be quickly and accurately labeled simply by extracting relevant content from the fixed fields of the prompts using preset parsing rules. This automated label generation method eliminates the need for manual annotation, significantly shortening the data annotation cycle and avoiding the subjective biases that may arise from manual annotation, ensuring the accuracy and consistency of the labels. This further improves the overall efficiency and quality of sound event detection data construction, laying a reliable data foundation for subsequent model training, evaluation, and optimization.

[0086] In one example, the step of generating structured semantic prompts corresponding to each of the at least one sound event type based on the semantic description information using a pre-trained large-scale language model includes:

[0087] A multi-round sampling mechanism is adopted to call the large-scale language model multiple times on the semantic description information to generate multiple structured semantic prompts with different expressions that correspond to the semantic description information.

[0088] Large-scale language models (such as ChatGPT and GPT-4) exhibit a degree of randomness in text generation; even with the same input, the output may differ each time. While this randomness can be advantageous in some scenarios, in synthesizing sound event detection data, it's crucial to ensure that the generated instructions not only meet semantic requirements but also possess sufficient diversity to cover different application scenarios. Instructions generated through single-round sampling are often limited to a specific output of the model, making it difficult to simultaneously meet the requirements of accuracy and diversity. Therefore, employing a multi-round sampling mechanism, by adjusting sampling parameters and prompting methods, allows for a systematic exploration of the model's output space, generating multiple candidate instructions with subtle differences. This significantly improves instruction diversity while maintaining semantic consistency. For example, a multi-round sampling mechanism can be used to repeatedly call a large-scale language model on a unified semantic description, generating multiple structured semantic prompts with different expressions that correspond to the stated semantic description.

[0089] In one possible implementation, after obtaining structured semantic prompts for different expressions of the same semantic meaning, before the pre-trained audio generation model generates synthesized audio data consistent with the semantic description information based on each of the structured semantic prompts, the method further includes:

[0090] From the structured semantic prompts that express the same semantic meaning in different ways, select the structured semantic prompts that meet the pre-configured audio generation conditions;

[0091] The audio generation conditions include one or more of the following: semantic filtering and alignment verification, and language model scoring. The semantic filtering and alignment verification verifies whether the structured semantic prompt instruction contains key event keywords specified by the preset semantic constraint rules by calculating the semantic similarity between the structured semantic prompt instruction and the preset semantic constraint rules, and filters structured semantic prompt instructions that meet the requirements based on the similarity and keyword coverage. The language model scoring obtains an evaluation score for the structured semantic prompt instruction using the large-scale language model, and filters structured semantic prompt instructions based on the evaluation score.

[0092] In practical applications, although structured semantic prompts generated from multiple rounds of sampling exhibit diversity in their different expressions of the same semantic meaning, not all structured semantic prompts can effectively drive audio generation models to produce high-quality synthetic audio data. Some structured semantic prompts may suffer from semantic ambiguity, missing key information, or poor compatibility with audio generation models. Therefore, it is necessary to screen these prompts before using a pre-trained audio generation model to generate synthetic audio data based on structured semantic prompts. For example, pre-configured audio generation conditions serve as the basis for screening structured semantic prompts. These audio generation conditions are a set of screening criteria constructed by combining the input specifications of the audio generation model, the characteristics of the target sound event, and practical application requirements. These criteria include, but are not limited to, one or more of the following:

[0093] I. Semantic Filtering and Alignment Verification. Semantic filtering and alignment verification can screen for semantic accuracy. During execution, structured semantic prompts are meticulously compared with a pre-defined semantic constraint rule base. Leveraging advanced semantic technology, the similarity between structured semantic prompts and rules in the semantic space can be accurately calculated. Simultaneously, keywords in the structured semantic prompts are extracted, rigorously verifying whether they contain the key event keywords specified by the semantic constraint rules, and accurately calculating keyword coverage. For example, to generate synthesized audio of the sound "snoring," the relevant instructions should contain the keyword "snoring," and the similarity with the pre-defined semantic constraint rules must reach a certain standard. Only when the semantic similarity reaches the threshold set by the audio generation conditions, and the keyword coverage meets the standard, can the structured semantic prompts proceed to the next stage of processing.

[0094] II. Language Model Scoring. Language model scoring provides a comprehensive evaluation of instruction quality. During execution, the acquired structured semantic prompts can be input back into the model. The large-scale language model will comprehensively evaluate the structured semantic prompts based on pre-defined scoring criteria, considering multiple key dimensions such as semantic accuracy, fluency, and information completeness, and output corresponding scores. For example, if an instruction has issues such as ambiguous semantic expression, logical incoherence, or missing key information, the large-scale language model will give a lower evaluation score. Conversely, instructions that are semantically clear, fluent, and information-complete will receive relatively higher scores. In this way, high-quality structured semantic prompts that better meet the input requirements of the audio generation model can be selected from multiple large-scale language instructions.

[0095] It should be noted that one or more of the above methods can be flexibly selected as audio generation conditions according to application needs, so as to further improve the automation of prompt generation and the ability to control the quality of data generation.

[0096] Example 4:

[0097] To further improve the quality of the generated audio data, based on the above embodiments, the method in this application further includes:

[0098] For each of the synthesized audio data, the verification information of the synthesized audio data is obtained based on the synthesized audio data and the sound event type label corresponding to the synthesized audio data;

[0099] For each of the structured semantic prompt instructions, the feedback information of the structured semantic prompt instruction is determined based on the verification information of each synthesized audio data corresponding to the structured semantic prompt instruction; and the usage of the semantic description information is determined based on the feedback information.

[0100] Based on the usage of each of the semantic description information, the pre-configured semantic constraint rules are adjusted.

[0101] To improve the quality and training value of synthesized audio data, this application allows for the quality review of the acquired synthesized audio data. This review can include one or more of the following methods: analyzing the acoustic quality of the synthesized audio data using acoustic feature algorithms; evaluating the synthesized audio data using an audio quality scoring model, such as performing soft validation using a pre-trained sound event detection model to automatically score and filter the event identifiability in the synthesized audio data; calculating a quality score by comparing the consistency between the soft labels output by the recognition model and the actual labels of the synthesized audio data, thus achieving an automated evaluation and filtering mechanism; and conducting quality reviews of a portion of the generated audio using manual sampling and listening to ensure the clarity of the target event and the consistency of the content with the prompts. For example, in manual sampling and listening, a certain proportion of sample pairs (x...) are randomly selected. i ,y i The audio content was listened to and identified by personnel with experience in audio event recognition. Among them, x in this sample pair... i Represents synthesized audio data, y i Represents synthesized audio data x i The corresponding weak labels are then identified. Next, the content is listened to to verify whether the labeled sound events actually appeared in the synthesized audio data, and whether the sound events were clear, loud, and identifiable. Finally, samples that do not meet the requirements are recorded and removed, and their corresponding semantic structured semantic prompts are analyzed to determine if there are issues such as unclear descriptions, semantic ambiguity, or scene interference.

[0102] It should be noted that a single review method can be used to review synthesized audio data, or a combination of multiple review methods can be used to review synthesized audio data, thereby maximizing the accuracy and reliability of the review.

[0103] The above method establishes a closed-loop synthesis process: "structured semantic prompt instruction generation → synthesized audio data generation → label generation → quality verification." This mechanism helps ensure the credibility of the dataset and the effectiveness of training. Based on the above review method, issues with sound event recognition and semantic consistency in the generated audio can be effectively identified, providing a reliable guarantee for the quality control of synthesized audio data.

[0104] In one possible implementation, the method further includes:

[0105] For each of the synthesized audio data, the verification information of the synthesized audio data is obtained based on the synthesized audio data and the sound event type label corresponding to the synthesized audio data;

[0106] For each of the structured semantic prompt instructions, the feedback information of the structured semantic prompt instruction is determined based on the verification information of each synthesized audio data corresponding to the structured semantic prompt instruction; and the usage of the semantic description information is determined based on the feedback information.

[0107] Based on the usage of each of the semantic description information, the pre-configured semantic constraint rules are adjusted.

[0108] Based on the above embodiments, verification information corresponding to each synthesized audio data can be obtained. On this basis, the verification information of all synthesized audio data generated by the same structured semantic prompt instruction can be aggregated. Taking the structured semantic prompt instruction for "baby crying" as an example, the verification information of all corresponding synthesized audio data can be integrated, including multi-dimensional data such as audio rejection rate, quality score, and acoustic feature matching degree. By analyzing this verification information, such as calculating the average quality score to judge the overall audio quality level, analyzing data reliability based on the rejection rate, and using acoustic feature distribution to mine differences in audio characteristics, feedback information for the structured semantic prompt instruction can be constructed. This aggregation analysis can not only quantitatively evaluate the overall quality level of audio data, but also deeply explore the potential problems behind data anomalies and present them in the form of feedback information. When analyzing the verification information of synthesized audio for "snoring," if the rejection rate is found to be significantly higher than that of audio generated by other instructions, the root cause of the problem can be located by tracing the detailed verification information of the abnormal samples. The final feedback information will clearly show that some audio has problems such as disordered snoring rhythm and blurred sound events, and clearly points out that these problems are related to the lack of limitation on key details such as snoring frequency and sound clarity in the structured semantic prompt instruction.

[0109] Since structured semantic prompts are generated based on semantic description information, this feedback information becomes a crucial basis for evaluating the usage of that information. For example, if a certain type of semantic description information, after being converted into a structured semantic prompt, consistently produces low-quality audio, and the feedback information repeatedly mentions missing key parameters or semantic ambiguity, it indicates that the semantic description information may have defects such as unclear expression or missing key parameters. Based on this, the usage of the semantic description information can be determined according to the feedback information. This allows for the differentiation between efficient, inefficient, and ineffective semantic description information, providing precise data support for subsequent optimization of semantic constraint rules and adjustment of semantic description strategies. This enables end-to-end quality control and iterative upgrades from semantic description to audio generation.

[0110] After obtaining the usage data of each semantic description piece of information, the pre-configured semantic constraint rules can be adjusted based on this data to delete or modify inefficient or invalid semantic descriptions. For example, based on feedback information, the usage of semantic descriptions is evaluated in a tiered manner, accurately classifying them into efficient, inefficient, and invalid categories. Efficient semantic descriptions can be retained as high-quality templates, inefficient information needs targeted modification and improvement, and invalid information is directly eliminated. This optimizes the semantic constraint rule set, and the next round of generation is then performed based on the optimized set of semantic constraint rules. The overall closed-loop process can be summarized as follows: Where I represents the semantic description information set, Prompting represents large-scale language model processing, P represents structured semantic prompting instructions, AudioGen represents audio generation model processing, X represents synthesized audio data, spot-check represents review processing, Feedback represents usage information, and Adjustment represents adjustment processing. * This represents the optimized semantic description information set.

[0111] The data generation closed loop formed by the above evaluation and feedback optimization not only improves the usability of the samples, but also lays the foundation for the introduction of automated verification in the future.

[0112] Example 5:

[0113] The method for synthesizing sound event detection data based on semantic prompts provided in this application will be described below through specific embodiments. Figure 2 A schematic diagram illustrating a specific semantic-cue-based sound event detection data synthesis process provided in this application embodiment is shown, the process including:

[0114] S201: Determine semantic description information based on a pre-configured set of sound event types and a set of semantic constraint rules.

[0115] The semantic description information includes at least one sound event type and semantic constraint rules pre-configured for at least one sound event type.

[0116] For example, the set of sound event types C = {Snoring, InfantCry}, the set of semantic constraint rules R: clear, prominent, generate 100 audio entries, and the semantic description information I is "Please generate 50 audio descriptions for Snoring and InfantCry respectively, requiring the event sounds to be clear and loud".

[0117] S202: Generation of structured semantic prompts: Based on semantic description information, determine structured semantic prompts corresponding to at least one type of sound event by using a pre-trained large-scale language model.

[0118] Taking the above example again, through a pre-trained large-scale language model, based on semantic description information I, the structured semantic prompt instruction p = "Snoring: A person snoring loudly in a quiet room." is determined.

[0119] In one possible implementation, a large-scale language model is used to generate structured semantic prompts with consistent syntactic structures based on semantic description information and pre-configured syntactic structure template prompts.

[0120] S203: Audio Sample Generation: Using a pre-trained audio generation model, synthetic audio data that matches the semantic description information is generated based on each structured semantic prompt instruction.

[0121] For example, let P be the set containing various structured semantic prompt instructions, and let p be any structured semantic prompt instruction. i Input ∈P into a pre-trained audio generation model, and output synthesized audio data x. i The audio data is stored in .wav file format, thus forming a set X containing various synthesized audio data. This audio data set X = {x1, x2, ..., x...} N}, where N represents the total number of sound event types, x N This represents the synthesized audio data corresponding to the Nth sound event type.

[0122] S204: Label Construction and Structure Annotation: For each synthesized audio data, based on the sound event type identifier at a preset position in the structured semantic prompt instruction corresponding to the synthesized audio data, determine the sound event type label corresponding to the synthesized audio data and generate training data pairs.

[0123] For example, construct a weak label set Y = {y} for each synthesized audio data. i},in:

[0124] y i =extract_event(p i )

[0125] The extract_event() function is a tag extraction function based on regular expressions or keyword matching, which can extract sound event types as weak tags from a fixed position in a structured semantic prompt instruction.

[0126] S205: Sample Validation and Closed-Loop Optimization: Evaluate the identifiability of the generated synthetic audio data samples, perform feedback optimization on the semantic description information to obtain semantic description information, and execute S202.

[0127] In one possible implementation, the quality of the acquired synthesized audio data can be reviewed. This review can include one or more of the following methods: analyzing the acoustic quality of the synthesized audio data using acoustic feature algorithms; evaluating the synthesized audio data using an audio quality scoring model, such as performing soft validation using a pre-trained sound event detection model to automatically score and filter the event identifiability in the synthesized audio data; calculating a quality score by comparing the consistency between the soft tags output by the recognition model and the actual tags of the synthesized audio data, thus achieving an automated evaluation and filtering mechanism; and conducting quality reviews of a portion of the generated audio using manual sampling and listening to ensure that the target event is clear and the content is consistent with the prompts. For example, in manual sampling and listening, a certain proportion of sample pairs (x...) are randomly selected. i ,y i The audio content was listened to and identified by personnel with experience in audio event recognition. Among them, x in this sample pair... i Represents synthesized audio data, y i Represents synthesized audio data x i The corresponding weak labels are then identified. Next, the content is listened to to verify whether the labeled sound events actually appeared in the synthesized audio data, and whether the sound events were clear, loud, and identifiable. Finally, sample pairs that do not meet the requirements are recorded and removed, and their corresponding structured semantic prompts are analyzed to determine if there are issues such as unclear descriptions, semantic ambiguity, or scene interference. All synthesized audio verification information generated from the same structured semantic prompt is aggregated to obtain feedback information for the structured semantic prompt. Based on the feedback information, the usage of semantic description information is determined. Based on the usage of each semantic description, the pre-configured semantic constraint rules are adjusted to delete or modify inefficient and invalid semantic description information. For example, the overall closed-loop process can be summarized as follows: Where I represents the semantic description information set, Prompting represents large-scale language model processing, P represents structured semantic prompting instructions, AudioGen represents audio generation model processing, X represents synthesized audio data, spot-check represents review processing, Feedback represents usage information, and Adjustment represents adjustment processing. * This represents the optimized semantic description information set.

[0128] Example, adjust I to add background sounds → I* = "Please generate 50 audio descriptions for each of InfantCry and Snoring, requiring clarity and volume, and also add background sounds from a home environment, such as music and people talking."

[0129] S206: Synthetic training data output and model training: Based on each synthesized audio data and its corresponding soft label, determine the training data for the sound event detection model.

[0130] For example, the structured semantic prompt instruction x i =“Snoring: A person snoring loudly in a quiet room.” and the corresponding soft tag y i ="Snoring", constructing training data pairs (x i ,y i These are used together to support the subsequent model training process.

[0131] The above embodiments can output a large-scale, semantically rich synthetic audio training dataset, which can be directly used in the training or pre-training stage of downstream sound event detection models. This audio training dataset can be used as an independent dataset or jointly trained with real-world labeled data to improve the model's generalization ability and robustness in real-world environments. This method has significant advantages such as high automation, strong data diversity, and low labeling cost. It can achieve high-quality, highly diverse, and highly consistent training sample generation without the need for on-site collection and labeling, providing high-quality training support for sound event detection models and significantly improving the robustness and performance of sound event detection systems in practical application scenarios.

[0132] Example 5:

[0133] This application provides a training method for the sound event detection model based on any of the above embodiments 1-5. Figure 3 This is a schematic diagram illustrating the training process of a sound event detection model provided in an embodiment of this application. The process includes:

[0134] S301: Obtain any audio sample in the sample set and its corresponding real sound event detection label; wherein, the audio sample includes strongly labeled audio samples and weakly labeled audio samples, the weakly labeled audio samples include synthetic audio data obtained based on any of the methods described in embodiments 1-5 above, the real sound event detection label corresponding to the strongly labeled audio sample includes a real sound event type label and a real time detection label, and the real sound event detection label corresponding to the weakly labeled audio sample includes a real sound event type label.

[0135] S302: Using the original sound event detection model, based on the audio sample, obtain the predicted sound event detection result corresponding to the audio sample.

[0136] S303: Based on the predicted sound event detection results and the actual sound event detection labels, the original sound event detection model is trained to obtain a trained sound event detection model.

[0137] The training method for the sound event detection model provided in this application is applied to a computer device, which can be a smart device or a server. The computer device used for training the sound event detection model in this application can be the same as, or different from, the computer device used for synthesizing sound event detection data based on semantic prompts.

[0138] In one possible implementation, sound event detection data synthesis is typically performed offline.

[0139] In this application, a sample repository for training the sound event detection model is pre-built, and audio samples are collected through various channels to construct a sample set. Strongly labeled audio samples can be derived from publicly available standard audio datasets (such as AudioSet, ESC-50, etc.), which have been manually labeled with sound event type labels and precise time detection labels (marking the start and end times of the sound event in the audio). Weakly labeled audio samples are generated according to any of the methods in embodiments 1-5 above, that is, using a large-scale language model combined with semantic description information and semantic constraint rules to generate structured semantic prompts, which then drive the audio generation model to obtain synthesized audio data and automatically obtain its corresponding sound event type labels.

[0140] In one possible implementation, the collected audio samples may be preprocessed, including unifying the audio format (e.g., converting to WAV format), adjusting the sampling rate and quantization bit depth to model adaptation standards (e.g., 44.1 kHz sampling rate, 16-bit quantization).

[0141] When training the original sound event detection model, any audio sample and its corresponding sound event detection label can be obtained from the sample set. This original sound event detection model can be a common convolutional neural network (CNN) structure (such as ResNet, VGG, etc.) or a hybrid structure combining recurrent neural networks (RNNs) to adapt to the processing of audio sequence data. The audio sample is then input into the original sound event detection model. The original sound event detection model processes the input audio sample to obtain the predicted sound event detection result corresponding to the audio sample. This predicted sound event detection result includes the prediction of the sound event type and the prediction of the start and end times of the sound event in the audio. The loss value is determined based on the error between the predicted sound event detection result and the sound event detection label. For example, the cross-entropy loss function is used to measure the difference between the predicted sound event type and the sound event type label, and the mean squared error (MSE) loss function is used to calculate the error between the predicted time detection result and the actual time detection label. Based on the loss value, the parameters of the original sound event detection model are adjusted to obtain the trained sound event detection model.

[0142] In one example, training the original sound event detection model based on the predicted sound event detection results and the actual sound event detection labels to obtain a trained sound event detection model includes:

[0143] For each strongly labeled audio sample, a first loss value is determined based on the predicted sound event detection result of the strongly labeled audio sample and the actual sound event detection label.

[0144] For each of the weakly labeled audio samples, a second loss value is determined based on the predicted sound event detection result of the weakly labeled audio sample and the actual sound event detection label.

[0145] The original sound event detection model is trained based on each of the first loss values ​​and each of the second loss values ​​to obtain a trained sound event detection model.

[0146] In this application, a differential loss calculation and joint optimization strategy can be used to train the sound event detection model based on the different characteristics of strongly labeled and weakly labeled audio samples. For example, for strongly labeled audio samples (which simultaneously contain real sound event type labels and real time detection labels), a first loss value is determined based on the error between the predicted sound event type and the real sound event type label, and the error between the predicted time detection result and the real time detection label. For instance, the cross-entropy loss function is used to measure the type classification error between the predicted sound event type and the sound event type label, and the mean squared error (MSE) loss function is used to calculate the time detection error between the predicted time detection result and the real time detection label. The first loss value is determined based on the type error and the time error. The type classification error and the time detection error can be proportionally fused to form the final first loss value, which can be expressed by the following formula: L1 = δL type +(1-δ)L time Where L1 represents the first loss value, L type The classification error L represents the type. type δ represents the type classification error L type The corresponding ratio, L time This represents the temporal detection error. For weakly labeled audio samples (containing only true sound event type labels), a second loss value is determined based on the error between the predicted sound event type and the true sound event type label for that weakly labeled audio sample. For example, cross-entropy loss or consistency loss can be used to measure the type classification error between the predicted sound event type and the sound event type label. Then, based on the obtained first and second loss values, a comprehensive loss value is determined. Based on this comprehensive loss value, the parameters of the original sound event detection model are adjusted to obtain the trained sound event detection model.

[0147] In one possible implementation, the comprehensive loss value is determined based on the obtained first and second loss values, expressed by the following formula:

[0148] L=βL1+α·L2

[0149] Where L represents the overall loss value, L1 represents the first loss value, L2 represents the second loss value, β represents the weight corresponding to the first loss value, and α represents the weight corresponding to the second loss value. α and β can be pre-configured constant values ​​or hyperparameters that are adjusted during the training process.

[0150] Since the sample set used to train the original sound event detection model contains a large number of audio samples, the above operation is performed on each audio sample. When the preset convergence condition is met, the original sound event detection model is trained and a trained sound event detection model is obtained.

[0151] The preset convergence condition can be that the sum of the loss values ​​corresponding to the audio samples in the current iteration sample set reaches a minimum or tends to stabilize, or the number of iterations for training the original sound event detection model reaches the set maximum number of iterations, etc. These conditions can be flexibly set in practice and are not specifically limited here.

[0152] As one possible implementation, when training the model, the audio samples in the sample set can be divided into a training set, a validation set, and a test set. First, the original sound event detection model is trained based on the training set. Then, the reliability of the trained sound event detection model is verified based on the validation set. Finally, the performance of the trained sound event detection model is tested based on the test set.

[0153] Example 7:

[0154] This embodiment demonstrates the practical application of semantically prompted synthetic audio data in sound event detection tasks, conducting comparative experiments based on the official baseline system of DCASE2023 Task4A. The baseline model adopts the publicly available teacher-student model (CRNN) architecture. The training dataset used in the experiment is a sound event detection dataset previously independently constructed by the applicant, covering two event categories: "infantcry" and "snoring." This dataset contains strongly labeled samples with precise start and end times, as well as weakly labeled samples containing only event category information, and features realistic scene backgrounds and representative noise environments.

[0155] Based on this, the following two experimental schemes were designed:

[0156] Experiment 1 (Replacing an Equal Amount of Weakly Labeled Data): All weakly labeled samples in the original dataset are replaced with an equal amount of synthesized audio samples and their corresponding weak labels generated by the method of this application, while the strong label part remains unchanged;

[0157] Experiment 2 (Replacing the Expanded Weakly Labeled Data): Based on Experiment 1, the synthesized weakly labeled data is expanded to 5 times the size of the original weakly labeled data and completely replaces the original weakly labeled part to construct a large-scale weakly labeled training set.

[0158] Given that the number of weakly labeled audio samples is significantly increased after using the large-scale weakly labeled synthetic samples generated in this application, the traditional strategy of applying the same loss weights to both strong and weak labels during training can easily lead to training imbalance, thus affecting the model's ability to fully learn from strong supervision information. To address this issue, in Experiment 2, a weakly labeled loss weighting mechanism is introduced based on the original loss function, and the following optimized loss function structure is proposed:

[0159] L=βL1+α·L2

[0160] Where L represents the comprehensive loss value, L1 represents the first loss value, L2 represents the second loss value, β represents the weight corresponding to the first loss value, and α represents the weight corresponding to the second loss value. α and β can be pre-configured constant values ​​or hyperparameters adjusted during the training process. This weighting strategy is an adaptive optimization mechanism proposed in this application for the model training part under the introduction of semantically prompted synthetic audio data. It effectively improves the robustness of the overall training process and the model's responsiveness to strongly supervised information, demonstrating good generalizability and practical value.

[0161] This application evaluates model performance based on event-based F1 scores and multi-threshold-based scores—polyphonic sound event detection scores (PSDS1 and PSDS2). The event-based F1 score measures sound event detection performance, using complete events as the calculation unit. It requires that the detected event category and start / end time be sufficiently close to the real event. The calculation formula is as follows:

[0162]

[0163] The accuracy rate is calculated as follows:

[0164]

[0165] The recall rate is calculated as follows:

[0166]

[0167] Wherein, TP (True Positive) is the number of samples correctly predicted as positive; FP (False Positive) is the number of samples incorrectly predicted as positive; and FN (False Negative) is the number of samples incorrectly predicted as negative.

[0168] Event-based F1 scoring uses a complete event as the unit of calculation, focusing on the accuracy of the event level in the detection results. In this evaluation method, an event is considered correct (True Positive, TP) only if the detected event is correctly classified and its start and end times are sufficiently close to the reference event.

[0169] PSDS1 emphasizes the model's detection speed, requiring the system to react quickly during event detection (e.g., triggering alarms, adapting to home automation systems), thus prioritizing recall. PSDS2, on the other hand, prioritizes accurate event classification when reaction time is less critical, emphasizing precision to avoid class confusion. The calculation method for PSDS can be summarized as follows:

[0170]

[0171] Among them, e max denoted as eFPR (effective false positive rate), and r(e) is the ROC curve calculated under multiple operation points.

[0172] The test set audio data is preprocessed and features are extracted. The test data is then fed into a sound event detection model to predict the sound events contained in the test data. Based on the probability values ​​of each event category predicted by the model in each frame, three F1 scores, PSDS1 and PSDS2 are calculated.

[0173] The following are the experimental results of a complex sound event detection data synthesis method based on a large-model semantic prompting and audio generation model:

[0174] system Event-F1 (%) PSDS1 (%) PSDS2 (%) 1 83.09 88.81 95.99 2 84.80 90.77 97.39 3 84.97 91.14 97.61

[0175] System 1 is the official baseline system of DCASE 2023 Task 4A, using the original crying and snoring datasets; System 2 uses the method described in this application, replacing the original weakly labeled data with synthetic data in equal amounts while maintaining the same data volume; System 3 generates and expands the original dataset using the method described in this application by 5 times, while introducing a weighted coefficient for weak label loss to balance the training. Experimental results show that compared with the high-performance official baseline system, the method described in this application achieves stable improvements in several key indicators: System 2 improves Event-F1, PSDS1, and PSDS2 by 1.71%, 1.96%, and 1.40%, respectively; System 3 further improves Event-F1 to 84.97% and PSDS1 / PSDS2 to 91.14% / 97.61% by expanding the weakly labeled data and introducing a weighting mechanism, verifying the effectiveness of the data synthesis and training optimization strategy.

[0176] Experimental conclusions: The semantic prompt-based synthetic audio data method proposed in this application, combined with a weak label loss weighting mechanism, can significantly improve the accuracy and robustness of the sound event detection model, especially in large-scale weak label data scenarios, and has good engineering application value.

[0177] Example 7:

[0178] Based on the same inventive concept, this application also provides a device for synthesizing sound event detection data based on semantic prompts. Figure 4 A schematic diagram of a sound event detection data synthesis device based on semantic prompting provided in this application embodiment, the device comprising:

[0179] The acquisition unit 41 is used to acquire semantic description information of the sound event detection task; wherein, the semantic description information includes at least one sound event type and semantic constraint rules pre-configured for the at least one sound event type;

[0180] The structured semantic generation unit 42 is used to generate structured semantic prompt instructions corresponding to the at least one sound event type based on the semantic description information using a pre-trained large-scale language model; wherein, the structured semantic prompt instruction for any sound event type is used to guide the generation of audio of the corresponding type;

[0181] The audio generation unit 43 is used to generate synthetic audio data that matches the semantic description information based on each of the structured semantic prompt instructions through a pre-trained audio generation model.

[0182] The tag generation unit 44 is used to determine the sound event type tag corresponding to each synthesized audio data based on the structured semantic prompt instruction corresponding to the synthesized audio data.

[0183] In this embodiment, the sound event detection data synthesis device based on semantic prompts is presented in the form of a functional module. Here, a module refers to an application-specific integrated circuit (ASIC), a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0184] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0185] Example 8:

[0186] This application also provides a training device for a sound event detection model. Figure 5 This application provides a schematic diagram of a training device for a sound event detection model, comprising:

[0187] The acquisition module 51 is used to acquire any audio sample in the sample set and its corresponding real sound event detection label; wherein, the audio sample includes strong-labeled audio samples and weak-labeled audio samples, the weak-labeled audio samples include synthesized audio data obtained based on any of the sound event detection data synthesis methods described in embodiments 1-5 above, the real sound event detection label corresponding to the strong-labeled audio sample includes a real sound event type label and a real time detection label, and the real sound event detection label corresponding to the weak-labeled audio sample includes a real sound event type label;

[0188] Processing module 52 is used to obtain the predicted sound event detection result corresponding to the audio sample based on the audio sample using the original sound event detection model;

[0189] Training module 53 is used to train the original sound event detection model based on the predicted sound event detection results and the real sound event detection labels to obtain the trained sound event detection model.

[0190] In this embodiment, the training device for the sound event detection model is presented in the form of functional modules. Here, a module refers to an application-specific integrated circuit (ASIC), a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0191] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0192] Example 9:

[0193] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of this application, such as... Figure 6 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 6 Take a processor 10 as an example.

[0194] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0195] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0196] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device as shown by a landing page for an app. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0197] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0198] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.

[0199] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touchscreen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touchscreen.

[0200] Example 10:

[0201] Based on the above embodiments, this application also provides a computer-readable storage medium storing a computer program executable by a processor. When the program runs on the processor, it causes the processor to perform the following steps:

[0202] Obtain semantic description information for a sound event detection task; wherein the semantic description information includes at least one sound event type and semantic constraint rules pre-configured for the at least one sound event type;

[0203] Using a pre-trained large-scale language model, structured semantic prompts are generated based on the semantic description information, corresponding to each of the at least one sound event type; wherein, the structured semantic prompts for any sound event type are used to guide the generation of the corresponding type of audio.

[0204] Using a pre-trained audio generation model, synthetic audio data that matches the semantic description information is generated based on each of the structured semantic prompts.

[0205] For each of the synthesized audio data, the sound event type label corresponding to the synthesized audio data is determined based on the structured semantic prompt instruction corresponding to the synthesized audio data.

[0206] Since the principle of the computer-readable storage medium in solving the problem is similar to that of the semantic prompt-based sound event detection data synthesis method, the implementation of the computer-readable storage medium can be found in Examples 1-5 of the method, and the repeated parts will not be described again.

[0207] Example 10:

[0208] Based on the above embodiments, this application also provides a computer-readable storage medium storing a computer program executable by a processor. When the program runs on the processor, it causes the processor to perform the following steps:

[0209] Obtain any audio sample in the sample set and its corresponding real sound event detection label; wherein, the audio sample includes strongly labeled audio samples and weakly labeled audio samples, the weakly labeled audio samples include synthesized audio data obtained based on any of the sound event detection data synthesis methods, the real sound event detection label corresponding to the strongly labeled audio sample includes a real sound event type label and a real time detection label, and the real sound event detection label corresponding to the weakly labeled audio sample includes a real sound event type label;

[0210] Based on the audio samples, the predicted sound event detection results corresponding to the audio samples are obtained using the original sound event detection model.

[0211] Based on the predicted sound event detection results and the actual sound event detection labels, the original sound event detection model is trained to obtain a trained sound event detection model.

[0212] Since the principle of the computer-readable storage medium in solving the problem is similar to the training method of the sound event detection model, the implementation of the computer-readable storage medium can be found in Examples 6-7 of the method, and the repeated parts will not be described again.

[0213] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for synthesizing sound event detection data based on semantic prompts, characterized in that, The method includes: Obtain semantic description information for a sound event detection task; wherein the semantic description information includes at least one sound event type and semantic constraint rules pre-configured for the at least one sound event type; Using a pre-trained large-scale language model, structured semantic prompts are generated based on the semantic description information, corresponding to each of the at least one sound event type; wherein, the structured semantic prompts for any sound event type are used to guide the generation of the corresponding type of audio. Using a pre-trained audio generation model, synthetic audio data that matches the semantic description information is generated based on each of the structured semantic prompts. For each of the synthesized audio data, the sound event type label corresponding to the synthesized audio data is determined based on the structured semantic prompt instruction corresponding to the synthesized audio data.

2. The method according to claim 1, characterized in that, The semantic constraint rules include, but are not limited to, one or more of the following: event sound source characteristics, background sound environment type, background sound interference level, scene spatial attributes, temporal structure features, audio data volume, and clarity.

3. The method according to claim 1, characterized in that, The sources of the semantic constraint rules include one or more of the following: a general semantic constraint rule set and a personalized semantic constraint rule set; wherein, the general semantic constraint rule set includes uniform semantic constraint rules applicable to all sound event types; and the personalized semantic constraint rule set includes exclusive semantic constraint rules applicable to specific sound event types.

4. The method as described in claim 1, characterized in that, The process of generating structured semantic prompts corresponding to at least one sound event type based on the semantic description information using a pre-trained large-scale language model includes: Using the large-scale language model, based on the semantic description information and pre-configured syntactic structure template prompts, structured semantic prompt instructions with consistent syntactic structures are generated.

5. The method as described in claim 1 or 4, characterized in that, The process of generating structured semantic prompts corresponding to at least one sound event type based on the semantic description information using a pre-trained large-scale language model includes: A multi-round sampling mechanism is adopted to call the large-scale language model multiple times on the semantic description information to generate multiple structured semantic prompts with different expressions that correspond to the semantic description information.

6. The method as described in claim 5, characterized in that, After obtaining structured semantic prompts for different expressions of the same semantic meaning, before generating synthesized audio data that matches the semantic description information using a pre-trained audio generation model based on each of the structured semantic prompts, the method further includes: From structured semantic prompts expressing the same semantic meaning in different ways, structured semantic prompts that meet pre-configured audio generation conditions are selected. These audio generation conditions include one or more of the following: semantic filtering and alignment verification, and language model scoring. The semantic filtering and alignment verification verifies whether the structured semantic prompts contain key event keywords specified by the pre-configured semantic constraint rules by calculating the semantic similarity between the structured semantic prompts and the preset semantic constraint rules, and selects structured semantic prompts that meet the requirements based on the similarity and keyword coverage. The language model scoring obtains an evaluation score for the structured semantic prompts using the large-scale language model, and selects structured semantic prompts based on the evaluation score.

7. The method as described in claim 5, characterized in that, The method further includes: For each of the synthesized audio data, the verification information of the synthesized audio data is obtained based on the synthesized audio data and the sound event type label corresponding to the synthesized audio data; For each of the structured semantic prompt instructions, the feedback information of the structured semantic prompt instruction is determined based on the verification information of each synthesized audio data corresponding to the structured semantic prompt instruction; and the usage of the semantic description information is determined based on the feedback information. Based on the usage of each of the semantic description information, the pre-configured semantic constraint rules are adjusted.

8. The method as described in claim 1, characterized in that, For each of the synthesized audio data, determining the sound event type label corresponding to the synthesized audio data based on the structured semantic prompt instruction corresponding to the synthesized audio data includes: For each synthesized audio data, the sound event type label corresponding to the synthesized audio data is determined according to the sound event type identifier at a preset position in the structured semantic prompt instruction corresponding to the synthesized audio data.

9. A training method for a sound event detection model, characterized in that, The method includes: Obtain any audio sample in the sample set and its corresponding real sound event detection label; wherein, the audio sample includes strongly labeled audio samples and weakly labeled audio samples, the weakly labeled audio samples include synthetic audio data obtained based on any of the methods described in claims 1-8, the real sound event detection label corresponding to the strongly labeled audio sample includes a real sound event type label and a real time detection label, and the real sound event detection label corresponding to the weakly labeled audio sample includes a real sound event type label; Based on the audio samples, the predicted sound event detection results corresponding to the audio samples are obtained using the original sound event detection model. Based on the predicted sound event detection results and the actual sound event detection labels, the original sound event detection model is trained to obtain a trained sound event detection model.

10. The method as described in claim 9, characterized in that, The step of training the original sound event detection model based on the predicted sound event detection results and the actual sound event detection labels to obtain a trained sound event detection model includes: For each strongly labeled audio sample, a first loss value is determined based on the predicted sound event detection result of the strongly labeled audio sample and the actual sound event detection label. For each of the weakly labeled audio samples, a second loss value is determined based on the predicted sound event detection result of the weakly labeled audio sample and the actual sound event detection label. The original sound event detection model is trained based on each of the first loss values ​​and each of the second loss values ​​to obtain a trained sound event detection model.

Citation Information

Cited By

  • Forklift automatic data synchronous scanning method and system

    CN122285777A