Audio scene analysis method and apparatus, service action execution method

By identifying event information and temporal relationships in audio signals, encapsulating them into structured semantic prompts, and inputting them into a large language model for audio scene analysis, this solves the problem of insufficient accuracy in audio scene analysis in existing technologies, and achieves highly accurate and interpretable audio scene understanding.

CN122224205APending Publication Date: 2026-06-16BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2026-02-06
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing audio scene analysis technologies lack accuracy, especially end-to-end direct input methods which are prone to issues such as confusion in the order of events.

Method used

By identifying the event information carried by the original audio signal, determining the start and end timestamps of the event, encapsulating it into structured semantic prompts according to the temporal relationship, and inputting it into a large language model for audio scene analysis, and combining preset semantic templates and scene categories, constructing instruction prompts to improve the accuracy of analysis.

Benefits of technology

It significantly improves the accuracy of audio scene analysis results generated by large language models, reduces fictitious descriptions, enhances the interpretability and transferability of models, supports concurrent modeling of multiple events, and adapts to different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122224205A_ABST
    Figure CN122224205A_ABST
Patent Text Reader

Abstract

The application provides an audio scene analysis method and device and a service action execution method, relates to the technical field of speech recognition, and comprises the following steps: identifying each event information carried by an original audio signal, the event information comprising an event type and event start and end timestamps; determining the time sequence relationship of each event type according to the event start and end timestamps; encapsulating each event information into a structured semantic prompt according to the time sequence relationship; inputting the structured semantic prompt into a large language model and prompting the large language model to perform audio scene analysis on the structured semantic prompt to obtain an audio scene analysis result output by the large language model. The reasoning basis of the large language model is changed from the original audio signal to an abstract structured semantic prompt, the phenomenon of fictitious description commonly seen in an end-to-end model is significantly reduced, and the accuracy of the audio scene analysis result generated by the large language model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech recognition technology, and in particular to an audio scene analysis method and apparatus, and a business action execution method. Background Technology

[0002] Audio scene analysis technology aims to understand environmental conditions by analyzing audio signals and is widely used in scenarios such as intelligent voice assistants, environmental monitoring, multimedia analysis, and public safety early warning. With the development of AI technology, audio scene understanding has become a core direction of human-computer interaction. Its capabilities have gradually evolved from the early "sound / scene recognition" to "recognizing dynamic events in a scene," requiring the comprehensive judgment of event correlation, temporal logic, and situation.

[0003] Existing technologies typically input audio features directly into large language models to achieve audio scene understanding. However, this end-to-end direct input method often results in audio scene analysis results that deviate from the audio facts and are prone to issues such as confusion regarding the order of events, leading to insufficient accuracy in audio scene analysis results. Summary of the Invention

[0004] This invention provides an audio scene analysis method and apparatus, and a business action execution method, to solve the technical problem of insufficient accuracy of audio scene analysis results in the prior art.

[0005] This invention provides an audio scene analysis method, comprising: Identify the event information carried by the original audio signal, wherein the event information includes the event type and the event start and end timestamps; The temporal relationship of each event type is determined based on the start and end timestamps of the events. Each event information is encapsulated into a structured semantic prompt according to the aforementioned temporal relationship; The structured semantic prompt is input into the large language model, and the large language model is prompted to perform audio scene analysis on the structured semantic prompt, so as to obtain the audio scene analysis results output by the large language model.

[0006] According to an audio scene analysis method provided by the present invention, the identification of various event information carried by the original audio signal includes: The original audio signal is preprocessed to obtain time-spectral characteristics; Extract the time-frequency feature vector of the time-spectrum features; Identify the event information carried by the time-frequency feature vector.

[0007] According to an audio scene analysis method provided by the present invention, the preprocessing includes one or more of sampling rate conversion, noise suppression, and normalization.

[0008] According to an audio scene analysis method provided by the present invention, the large language model performs audio scene analysis on the structured semantic prompt, including: Based on the structured semantic prompts and preset semantic templates, instruction prompts are constructed to prompt the large language model to perform audio scene analysis on the structured semantic prompts; The instruction prompt is input into the large language model.

[0009] According to an audio scene analysis method provided by the present invention, the structured semantic prompts include scene categories; The instruction prompting, constructed based on the structured semantic prompt and a preset semantic template, for prompting the large language model to perform audio scene analysis on the structured semantic prompt, includes: The scene category is written into the semantic template to obtain the instruction prompt, which is used to prompt the large language model to perform audio scene analysis on the sequence of each event information in the structured semantic prompt based on the scene rules of the scene category, and output the audio scene analysis result.

[0010] The present invention also provides a method for executing business actions, comprising: Obtain the audio scene analysis result based on any of the audio scene analysis methods described above; Execute the business actions corresponding to the audio scene analysis results.

[0011] The present invention also provides an audio scene analysis device, comprising: The identification module is used to identify various event information carried by the original audio signal, including event type and event start and end timestamps; The determination module is used to determine the temporal relationship of each event type based on the start and end timestamps of the events; An encapsulation module is used to encapsulate each of the event information into a structured semantic prompt according to the temporal relationship; The prompting module is used to input the structured semantic prompts into the large language model and prompt the large language model to perform audio scene analysis on the structured semantic prompts, so as to obtain the audio scene analysis results output by the large language model.

[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the audio scene analysis method or the business action execution method described above.

[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the audio scene analysis method or the business action execution method described above.

[0014] The audio scene analysis method and apparatus, and business action execution method provided by this invention identify the event information carried by the original audio signal, determine the temporal relationship of each event type according to the start and end timestamps of the events, and encapsulate the event information into structured semantic prompts according to the temporal relationship. This transforms the reasoning basis of the large language model from the original audio signal to the abstract structured semantic prompts, significantly reducing the fictional description phenomenon commonly found in end-to-end models and improving the accuracy of the audio scene analysis results generated by the large language model. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0016] Figure 1 This is a flowchart illustrating the audio scene analysis method provided by the present invention.

[0017] Figure 2 This is a flowchart illustrating the business action execution method provided by the present invention.

[0018] Figure 3 This is a schematic diagram of the audio scene analysis device provided by the present invention.

[0019] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0021] Currently, mainstream audio scene analysis technologies can be categorized into three types: Audio Event Detection (AED): It can locate discrete events but cannot establish event relationships.

[0022] Acoustic Scene Classification (ASC): It can roughly identify scenes but lacks dynamic temporal analysis.

[0023] Audio-Large Language Models: Represented by SALMONN, AudioGPT, and WavLLM, these models directly input audio features into large language models to achieve tasks such as speech recognition and auditory question answering, and are currently a research hotspot.

[0024] The aforementioned mainstream technologies have the following shortcomings: Limited scene reasoning ability: It is difficult to explicitly model the temporal and concurrent logic between events in end-to-end input, and the generated results often deviate from the audio facts, which can easily lead to confusion about the order of events. The modeling of multi-event relationships is lacking: it relies on shallow statistical co-occurrence features, lacks explicit descriptions of the causal chains and logical structures of events, and cannot effectively distinguish the correlation between events; Insufficient interpretability and transferability: Lacking an intermediate event layer, the model struggles to provide reasoning basis, and adaptation to new scenarios requires large-scale retraining, making it difficult to quickly adapt to different application scenarios and meet industrial-grade application requirements.

[0025] In particular, the audio-large language model lacks explicit representation of multiple concurrent sound events, making it difficult to accurately describe sound facts; it cannot express the temporal order, concurrency, and causal relationships between events, resulting in limited scenario reasoning capabilities; and the output results lack interpretability and traceability, making it impossible to directly link with actual business actions.

[0026] The following is combined Figures 1 to 4 This invention describes the audio scene analysis method and apparatus, and the business action execution method.

[0027] Figure 1 This is a flowchart illustrating the audio scene analysis method provided by the present invention, as shown below. Figure 1 As shown, the method includes, but is not limited to, steps S1, S2, S3 and S4.

[0028] Step S1: Identify the event information carried by the original audio signal. The event information includes the event type and the start and end timestamps of the event.

[0029] The raw audio signal can come from a microphone array or an audio acquisition device. Event information refers to the specific sound object data extracted from the raw audio signal, which may include event type, event start and end timestamps (event start time and event end time), and confidence level. Pre-trained audio event detection (AED) models can be used to identify the various event information carried in the raw audio signal.

[0030] Event types include, for example, sounds of heated arguments, physical impacts, and cries for help occurring at school. The start and end timestamps for these events could be: heated arguments: 14:25:10-14:25:18; physical impacts: 14:25:19-14:25:26; cries for help: 14:25:27-14:25:33. The confidence levels could be 92% for heated arguments, 89% for physical impacts, and 95% for cries for help.

[0031] Step S2: Determine the temporal relationship of each event type based on the start and end timestamps of the events.

[0032] Temporal relationships can include sequential relationships and / or concurrent relationships. Sequential relationships represent two event types occurring sequentially, while concurrent relationships represent two event types occurring simultaneously.

[0033] If the end time of the first event type is before the start time of the second event type, then the temporal relationship between the two event types is that the first event type comes first, followed by the second event type. For example, the temporal relationship between the sounds of a heated argument, the sounds of physical impact, and the cries for help is the order of heated argument - physical impact - cries for help, with no concurrent relationship.

[0034] If the start and end timestamps of the first event type overlap with the start and end timestamps of the second event type, then the temporal relationship between the two event types is a concurrent relationship.

[0035] Step S3: Encapsulate the information of each event into a structured semantic prompt according to the temporal relationship.

[0036] Each event information can be encapsulated into machine-readable structured semantic prompts according to temporal relationships. Step S3 can sort discrete event information into a coherent event sequence according to the timeline, and explicitly mark the temporal relationships between events (such as chronological order, synchronous occurrence), encapsulate the event sequence into standardized structured semantic prompts, and organize them in JSON template or table format as input for subsequent large language models.

[0037] Structured semantic hints, for example: {"events": [ {"type": "sound of heated argument", "start_time": "14:25:10", "end_time": "14:25:18", "confidence": 0.92}, {"type": "sound of physical impact", "start_time": "14:25:19", "end_time": "14:25:26", "confidence": 0.89}, {"type": "cries for help", "start_time": "14:25:27", "end_time": "14:25:33", "confidence": 0.95} ], "time_relation": "sequential", "scene_category": "campus_security"}; Among them, events are events, type is the type, start_time is the start time, end_time is the end time, confidence is the confidence level, time_relation is the time sequence relationship, sequential is the order, scene_category is the scene category, and campus_security is the campus security.

[0038] Steps S1-S3 perform fine-grained event-level analysis on the input raw audio signal, extract sound events that characterize scene features, and organize them in a structured manner according to the time dimension. This can transform the complex acoustic signals mixed in the raw audio signal into a clearly structured symbolic event description, clearly representing the key information of "what events happened" and "when the events happened".

[0039] Step S4: Input the structured semantic prompts into the large language model and prompt the large language model to perform audio scene analysis on the structured semantic prompts, and obtain the audio scene analysis results output by the large language model.

[0040] It can prompt large language models to perform deep semantic reasoning on the event sequence in structured semantic prompts. Combined with the preset reasoning rule base, it focuses on analyzing the causal logic, scene evolution and concurrent relationships between events. Under the dual constraints of temporal and causal reasoning, it obtains human-readable scene-level semantic descriptions and classification labels based on the reasoning results. Based on the scene-level semantic descriptions and classification labels, it generates audio scene analysis results.

[0041] Audio scene analysis results can include scene conclusions and explanatory information. That is, the audio scene analysis results can be interpretable conclusions, including scene labels, inferential evidence, and time slices of the inferential evidence. For example, an audio scene analysis result might state, "From 14:25:10 to 14:25:33, continuous sounds of intense arguing, physical impact, and cries for help were detected, consistent with the audio event characteristics of a school bullying incident." Here, the scene label is "school bullying incident," the inferential evidence is the continuous detection of intense arguing, physical impact, and cries for help, and the time slice of the inferential evidence is 14:25:10-14:25:33.

[0042] Step S3 is equivalent to introducing a structured semantic interface between the event detection layer and the semantic reasoning layer, realizing an interpretable mapping from audio events to scene semantics. By transforming the reasoning basis of the large language model from raw audio signals to abstract structured semantic cues, the fictitious description phenomenon common in end-to-end models is significantly reduced, while enhancing the transparency and interpretability of the system's decision-making process.

[0043] As described above, the audio scene analysis method of the present invention identifies the event information carried by the original audio signal, determines the temporal relationship of each event type according to the start and end timestamps of the events, and encapsulates each event information into a structured semantic prompt according to the temporal relationship. This transforms the reasoning basis of the large language model from the original audio signal to the abstract structured semantic prompt, significantly reducing the fictional description phenomenon commonly found in end-to-end models and improving the accuracy of the audio scene analysis results output by the large language model.

[0044] This invention achieves decoupled collaboration between the audio event layer and the semantic reasoning layer, balancing the model's interpretability, transferability, and practical applicability. Compared to existing end-to-end audio-language models, this invention significantly reduces the risk of hallucination-like misrepresentations and provides a generalized audio understanding solution that can be quickly adapted to different scenarios (such as classroom teaching supervision, traffic incident detection, and meeting status recognition) without retraining through templated structured prompts and an extensible knowledge base. Building upon existing event detection results, this invention introduces an audio perception layer and a structured semantic reasoning mechanism, achieving high-level semantic cognition through temporal logic constraints. This overcomes existing technological bottlenecks and achieves interpretable understanding from events to scenarios.

[0045] In one embodiment, step S1 may further include: The original audio signal is preprocessed to obtain its time-spectral characteristics; Extract the time-frequency feature vector of the time-spectral features; Identify the event information carried by the time-frequency feature vector.

[0046] Preprocessing refers to the process of cleaning and standardizing the raw audio signal to obtain standardized time-frequency spectral features. Standardized time-frequency spectral features characterize the energy distribution of sound at different frequencies over time. Extracting the time-frequency feature vector from the time-frequency spectral features is essentially transforming the time-frequency spectral features into a high-dimensional time-frequency feature vector.

[0047] The event information carried by the time-frequency feature vector can be identified through various audio self-supervised pre-trained models (such as WavLM, Hubert, Wav2vec, BEATS, etc.) and / or various audio event classification models (such as CED or PANs).

[0048] This invention converts the original audio signal into time-frequency spectral features and extracts time-frequency feature vectors, effectively filtering out irrelevant interference from the environment while preserving the key fingerprint information of the sound. This hierarchical feature extraction method significantly improves the robustness and accuracy of audio event detection, providing accurate basic data input for subsequent inference in large language models.

[0049] In one embodiment, the preprocessing of the present invention may include one or more of sampling rate conversion, noise suppression, and normalization.

[0050] Sampling rate conversion converts audio collected from different hardware devices into a standard sampling rate required by a large language model to ensure the consistency of input data.

[0051] Noise suppression can utilize Wiener filtering, spectral subtraction, or AI noise reduction algorithms to remove steady-state noise such as white noise and current noise from the background, highlighting the main sound events.

[0052] Normalization is a process of standardizing audio amplitude or spectral energy (such as Z-score normalization), scaling the values ​​to a uniform range (such as [-1, 1]) to prevent volume differences caused by different recording device gains from affecting the judgment of large language models.

[0053] By performing sampling rate conversion, noise suppression, and normalization on the original audio signal, standardized time-spectral characteristics can be obtained, providing high-quality input for subsequent event information recognition.

[0054] In one embodiment, step S4, where the large language model performs audio scene analysis on the structured semantic prompts, may further include: Based on structured semantic prompts and preset semantic templates, instruction prompts are constructed to prompt the large language model to perform audio scene analysis on the structured semantic prompts; Input the command prompt into the large language model.

[0055] Based on structured semantic prompts and according to preset semantic templates, it can automatically generate instruction prompts that conform to the processing specifications of large language models. This can effectively stimulate the domain-specific reasoning ability of large language models, standardize the input and output interfaces of large language models, and ensure that large language models can focus on the analysis of audio scenes according to the expected logic. This is conducive to improving the usability and standardization of audio scene analysis results.

[0056] In one embodiment, the structured semantic prompts of the present invention may include scene categories; Based on structured semantic prompts and preset semantic templates, instruction prompts are constructed to prompt large language models for audio scene analysis of structured semantic prompts. These prompts may further include: The scene category is written into the semantic template to obtain the scene rules based on the scene category for prompting the large language model to perform audio scene analysis on the sequence of event information in the structured semantic prompt, and output the instruction prompt of the audio scene analysis result.

[0057] The semantic template of this invention can be "Based on the xx scenario rules, analyze the causal relationship of the above event sequence, and output the scenario conclusion and explanation information." For the structured semantic prompt in the example above, if the scenario category is campus safety, then writing the scenario category into the semantic template will result in the instruction prompt "Based on the campus safety scenario rules, analyze the causal relationship of the above event sequence, and output the scenario conclusion and explanation information."

[0058] This invention achieves context-aware scene analysis by introducing scene categories and corresponding scene rules. The large language model can dynamically adjust its inference logic based on the specific application context, focusing on key risks or concerns within that scene. This significantly improves the professionalism and accuracy of the analysis results in specific vertical fields and reduces false positives.

[0059] like Figure 2 As shown, the business action execution method provided by the present invention includes: Step S5: Obtain the audio scene analysis results based on any of the above audio scene analysis methods; Step S6: Execute the business actions corresponding to the audio scene analysis results.

[0060] This invention can transmit the scenario conclusions and explanations output by the large language model to the business application layer, and execute corresponding business actions according to predefined strategies, including visualization, business rule triggering, and front-end parameter optimization, to achieve closed-loop control of "detection-understanding-execution".

[0061] Business actions may include: The audio scene analysis results are transmitted to the upper-level application interface for visualization. Pre-defined business rules are automatically triggered based on the identified scene tags, including classroom attendance records, abnormal situation alarms, meeting status annotations, and adaptive adjustment of front-end collected parameters.

[0062] This allows for a complete closed-loop processing flow of event detection, scenario understanding, and business execution, improving the practicality and automation level of the invention in real-world application environments.

[0063] This invention uses the identification and early warning of campus bullying incidents in a campus safety scenario as an example to illustrate the method of executing business actions.

[0064] Application Background: In real time, abnormal events such as physical conflicts and bullying need to be monitored in public areas of the campus (such as corridors, playgrounds, and stairwells). Traditional manual patrols have problems such as delayed response and incomplete coverage. This invention can achieve all-weather automatic early warning through audio event analysis.

[0065] System Deployment: Deploy microphone arrays with noise reduction capabilities in key areas of the campus, set the audio acquisition frequency to 16kHz, enable real-time noise suppression in the preprocessing module (filtering ambient noises such as wind noise and get out of class noise), load the "Campus Event Feature Library" (containing 12 types of events such as arguments, collisions, and cries for help), and pre-set the "Campus Security Reasoning Rule Library" (containing event causal chain judgment logic) in the large language model.

[0066] Specific execution process: 1. The microphone array acquires corridor audio in real time, and the system automatically completes sampling rate conversion and noise suppression, outputting standardized time-frequency spectral characteristics; 2. Extract the Mel spectrum features of the time-frequency spectrum features and identify the event information in the Mel spectrum features: "14:25:10-14:25:18 intense argument (confidence 92%)", "14:25:19-14:25:26 limb impact sound (confidence 89%)", and "14:25:27-14:25:33 cries for help (confidence 95%)". The sequence of events is as follows: sounds of arguing - sounds of physical contact - cries for help (occurring consecutively, without concurrency). Generate structured semantic hints in JSON format: { "events": [ {"type": "sound of heated argument", "start_time": "14:25:10", "end_time": "14:25:18", "confidence": 0.92}, {"type": "sound of physical impact", "start_time":"14:25:19", "end_time": "14:25:26", "confidence": 0.89}, {"type": "cries for help", "start_time": "14:25:27", "end_time": "14:25:33", "confidence": 0.95} ], "time_relation": "sequential", "scene_category": "campus_security"}; 3. Based on the preset semantic template, the system automatically generates the instruction prompt: "Based on the rules of the campus safety scenario, analyze the causal relationship of the above event sequence and output the scenario conclusion and explanation information." The large language model determines that the event sequence matches the "bullying event characteristics" and infers and generates the scene label "school bullying event"; The final audio scene analysis result is: "From 14:25:10 to 14:25:33, the sounds of intense arguing, physical collisions, and cries for help were continuously detected, which are consistent with the audio event characteristics of a school bullying incident." 4. Execution of business operations: Send early warning information to the campus security center, including the location of the incident (corridor zone 3), time, scene tag, and evidence; The system coordinates with the surveillance cameras in the area to focus on and capture images, transmitting the footage back to the security terminal in real time. The campus broadcast system was triggered to play a warning message in the vicinity ("Security personnel have arrived at the scene. Please stop the inappropriate behavior immediately").

[0067] As can be seen from the foregoing, the present invention has the following advantages: 1. The reasoning process is highly interpretable. This invention decomposes the audio scene understanding task into two stages: scene perception and scene reasoning, realizing a structured process from signal processing to semantic generation. This design effectively avoids the illusion problem common in end-to-end black-box models, making the output of each stage traceable and verifiable, greatly improving the transparency and credibility of the reasoning process.

[0068] 2. Supports concurrent modeling of multiple events, adapting to complex audio scenarios. By performing fine-grained identification of audio event information and generating structured event sequences, it is possible to model the temporal and semantic relationships between multiple overlapping sound events, overcoming the limitation of traditional methods that can only identify a single main sound source, and significantly improving the ability to understand scenes in real and complex environments.

[0069] 3. Achieved closed-loop control of detection-understanding-execution. Based on audio event recognition and scene semantic reasoning, this invention further feeds back structured conclusions to the front-end system or external business unit, realizing a complete closed-loop control from sound perception and semantic understanding to business execution, thereby enhancing the system's adaptability and real-time response efficiency.

[0070] 4. The system has a high degree of modularity, making it easy to expand and maintain. This invention adopts a hierarchical and modular design, separating functional layers such as audio event detection, semantic reasoning, and business response. The interfaces between each layer are clear and the functions are independent, supporting flexible replacement or upgrading of specific modules in different application scenarios, and possessing good system scalability and engineering maintainability.

[0071] 5. Possesses good transferability and cross-scenario adaptability. This invention, through a prompt design and knowledge base-driven few-sample instruction optimization mechanism, can be rapidly deployed in various fields such as classroom teaching supervision, meeting status recognition, and traffic incident monitoring, without the need to retrain the underlying acoustic model, significantly reducing application costs and deployment cycle.

[0072] The audio scene analysis device provided by the present invention is described below. The audio scene analysis device described below and the audio scene analysis method described above can be referred to in correspondence.

[0073] like Figure 3 As shown, the audio scene analysis device provided by the present invention includes: The recognition module is used to identify various event information carried by the original audio signal. The event information includes the event type and the start and end timestamps of the event. The determination module is used to determine the timing relationship of each event type based on the start and end timestamps of the events. The encapsulation module is used to encapsulate the information of each event into structured semantic prompts according to the temporal relationship; The prompting module is used to input structured semantic prompts into the large language model and for the large language model to perform audio scene analysis on the structured semantic prompts, thereby obtaining the audio scene analysis results output by the large language model.

[0074] Figure 4The example illustrates the physical structure of an electronic device, which may include a processor, a communications interface, memory, and a communication bus. The processor, communications interface, and memory communicate with each other via the communication bus. The processor can invoke logical instructions from the memory to execute audio scene analysis methods or business action execution methods.

[0075] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0076] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the audio scene analysis method or business action execution method provided by the above methods.

[0077] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the audio scene analysis method or business action execution method provided by the above methods.

[0078] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0079] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An audio scene analysis method, characterized in that, include: Identify the event information carried by the original audio signal, wherein the event information includes the event type and the event start and end timestamps; The temporal relationship of each event type is determined based on the start and end timestamps of the events. Each event information is encapsulated into a structured semantic prompt according to the aforementioned temporal relationship; The structured semantic prompts are input into a large language model, and the large language model is prompted to perform audio scene analysis on the structured semantic prompts, so as to obtain the audio scene analysis results output by the large language model.

2. The audio scene analysis method according to claim 1, characterized in that, The identification of various event information carried by the original audio signal includes: The original audio signal is preprocessed to obtain time-spectral characteristics; Extract the time-frequency feature vector of the time-spectrum features; Identify the event information carried by the time-frequency feature vector.

3. The audio scene analysis method according to claim 2, characterized in that, The preprocessing includes one or more of sampling rate conversion, noise suppression, and normalization.

4. The audio scene analysis method according to claim 1, characterized in that, The prompting, provided by the large language model, involves audio scene analysis of the structured semantic prompting, including: Based on the structured semantic prompts and preset semantic templates, instruction prompts are constructed to prompt the large language model to perform audio scene analysis on the structured semantic prompts; The instruction prompt is input into the large language model.

5. The audio scene analysis method according to claim 4, characterized in that, The structured semantic prompts include scene categories; The instruction prompting, constructed based on the structured semantic prompt and a preset semantic template, for prompting the large language model to perform audio scene analysis on the structured semantic prompt, includes: The scene category is written into the semantic template to obtain the instruction prompt, which is used to prompt the large language model to perform audio scene analysis on the sequence of each event information in the structured semantic prompt based on the scene rules of the scene category, and output the audio scene analysis result.

6. A method for executing a business action, characterized in that, include: Obtain the audio scene analysis result based on the audio scene analysis method according to any one of claims 1-5; Execute the business actions corresponding to the audio scene analysis results.

7. An audio scene analysis device, characterized in that, include: The identification module is used to identify various event information carried by the original audio signal, including event type and event start and end timestamps; The determination module is used to determine the temporal relationship of each event type based on the start and end timestamps of the events; An encapsulation module is used to encapsulate each of the event information into a structured semantic prompt according to the temporal relationship; The prompting module is used to input the structured semantic prompts into the large language model and prompt the large language model to perform audio scene analysis on the structured semantic prompts, so as to obtain the audio scene analysis results output by the large language model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the audio scene analysis method as described in any one of claims 1 to 5 or the business action execution method as described in claim 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the audio scene analysis method as described in any one of claims 1 to 5 or the business action execution method as described in claim 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the audio scene analysis method as described in any one of claims 1 to 5 or the business action execution method as described in claim 6.