Semantized audio processing method and device, computer equipment and storage medium

By analyzing natural language instructions to generate structured processing instructions and dynamically match the tool chain, the problem of high professionalization of existing audio processing technology is solved, and the ease of use and efficiency is improved, and it is suitable for audio processing in the financial and medical fields.

CN120544543APending Publication Date: 2025-08-26平安科技(上海)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510848275.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing audio processing technology has high degree of professionalism and high interaction barriers. Users need to deeply master professional terms, complex operations, and limited creative capabilities, making it difficult to meet the convenient needs of the financial and medical fields.

Method used

By receiving natural language instructions, the structured processing instructions are generated, the processing tool chain is dynamically matched, and the instructions are converted into processing parameters to realize audio processing.

Benefits of technology

It lowers the interactive threshold for audio processing, improves ease of use and efficiency, breaks through creative limitations, and non-professionals can also complete audio processing efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544543A_ABST
    Figure CN120544543A_ABST
Patent Text Reader

Abstract

The invention relates to the fields of artificial intelligence, financial science and technology and digital medical treatment, discloses a semantic audio processing method and device, computer equipment and a storage medium, and can be applied to audio processing in the fields of finance and medical treatment. The method comprises the following steps: receiving a natural language instruction and an original audio input by a user; analyzing the natural language instruction and the original audio, and generating a structured processing instruction; dynamically matching a processing tool chain according to the structured processing instruction, and converting the structured processing instruction into a processing parameter; and processing the original audio according to the processing tool chain and the processing parameters, and generating a target audio. According to the method, the natural language instruction is analyzed to generate the structured processing instruction, traditional terminology input is replaced, the interaction threshold is lowered, and the usability of audio processing is improved; by converting the natural language instruction into the processing parameter, the implementation capability of abstract requirements is improved, and workers in the financial and medical fields can efficiently and conveniently complete audio processing work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence, financial technology and digital medicine, and in particular to semantic audio processing methods, devices, computer equipment and storage media. Background Art

[0002] With the rapid development of digital technology, audio processing has become deeply integrated into people's daily lives and various fields. This is particularly true in the financial and medical fields. For example, in the financial sector, audio processing can be used to organize meeting recordings and optimize customer communication records to ensure the clear transmission and accurate retention of information. It can also be applied to the editing and adjustment of anti-fraud promotional audio to ensure the timely updating of financial knowledge. In the medical industry, audio processing can be used to process audio of medical lectures and case discussions, assisting doctors in learning, communication, and case analysis.

[0003] Currently, there are three main audio processing technologies: First, using traditional digital audio workstations (DAWs), such as Pro Tools and Audacity, to process audio, which can meet diverse audio processing needs; second, using voice assistant integration functions, such as Siri and Google Assistant, which can only achieve basic media control; and third, using automated tools based on preset rules, such as Auto-Tune, which can perform templated batch processing of audio.

[0004] However, these existing audio processing technologies have obvious defects. Specifically, these existing audio processing technologies all have high interaction barriers, requiring users to deeply master and apply professional terms such as "compression ratio" and "Q value", which are difficult for non-professionals to use conveniently. At the same time, the functions of various audio processing tools are scattered, and users often need to frequently switch between multiple software such as RX noise reduction and FL Studio arrangement, which greatly affects work efficiency. In addition, the creative capabilities of existing audio processing technologies and common audio processing tools are very limited, and it is difficult to understand users' abstract needs such as "making a 1980 tape effect". These problems make existing audio processing technologies not convenient for people in the financial and medical fields to use. They need to spend a lot of time learning complex audio processing knowledge and operating procedures, and cannot complete audio processing work efficiently and conveniently. Summary of the Invention

[0005] The present invention provides a semantic audio processing method, apparatus, computer equipment and storage medium to solve the technical problems of the existing audio processing having a high degree of specialization and limited creation.

[0006] First, a semantic audio processing method is provided, including:

[0007] Receive natural language commands and original audio input from the user;

[0008] Parsing the natural language instructions and the original audio to generate structured processing instructions;

[0009] Dynamically matching a processing tool chain according to the structured processing instruction and converting the structured processing instruction into processing parameters;

[0010] The original audio is processed according to the processing tool chain and processing parameters to generate target audio.

[0011] In a second aspect, a semantic audio processing device is provided, comprising:

[0012] A receiving module, configured to receive natural language commands and original audio input by a user;

[0013] A parsing module, configured to parse the natural language instructions and the original audio to generate structured processing instructions;

[0014] a conversion module, configured to dynamically match a processing tool chain according to the structured processing instruction and convert the structured processing instruction into a processing parameter;

[0015] A processing module is used to process the original audio according to the processing tool chain and processing parameters, and generate target audio.

[0016] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned semantic audio processing method when executing the computer program.

[0017] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned semantic audio processing method are implemented.

[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention generates structured processing instructions by parsing natural language instructions, replacing traditional professional terminology input, thereby lowering the interaction threshold, improving the usability of audio processing, and enabling non-professionals to easily operate audio processing; unifying the workflow through dynamic matching processing tool chains, avoiding users from manually switching between multiple software, and improving audio processing efficiency; converting natural language instructions into processing parameters helps to break through creative limitations, further reduces the difficulty of audio processing, and improves the ability to realize abstract or complex creative needs, which is beneficial for financial and medical personnel to complete audio processing work efficiently and conveniently.

[0019] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In addition, in order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the following preferred embodiments are specifically cited and described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is a schematic diagram of an application environment of a semantic audio processing method according to an embodiment of the present invention;

[0021] Figure 2 This is a flowchart of a semantic audio processing method according to an embodiment of the present invention;

[0022] Figure 3 yes Figure 2 A schematic flow chart of a specific implementation of step S20;

[0023] Figure 4 yes Figure 2 A schematic flow chart of a specific implementation of step S30;

[0024] Figure 5 yes Figure 2 Another flow chart of a specific implementation of step S30;

[0025] Figure 6 yes Figure 2 A schematic flow chart of a specific implementation of step S40;

[0026] Figure 7 yes Figure 2 A schematic flow chart of a specific implementation of step S50;

[0027] Figure 8 is a structural diagram of a semantic audio processing device according to an embodiment of the present invention;

[0028] Figure 9 is a structural diagram of a computer device in one embodiment of the present invention;

[0029] Figure 10 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0031] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0032] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0033] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0034] See also Figure 1 and Figure 2 , Figure 1 Schematic diagram of the application scenario of the semantic audio processing method provided by an embodiment of the present invention. Among them, the user can input natural language instructions and upload original audio through the client. The client can communicate with the server through the network; the server can receive the natural language instructions and original audio input by the user through the network; parse the natural language instructions and original audio to generate structured processing instructions; dynamically match the processing tool chain according to the structured processing instructions, and convert the structured processing instructions into processing parameters; process the original audio according to the processing tool chain and processing parameters, and generate target audio. Among them, the client can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.

[0035] See also Figure 2 As shown, Figure 2A schematic flow chart of a semantic audio processing method provided by an embodiment of the present invention. The semantic audio processing method includes the following steps:

[0036] S10: Receive natural language instructions and original audio input by the user.

[0037] Step S10 serves as the starting point of the entire audio processing flow, providing basic data for subsequent audio processing. It is understood that in this embodiment, natural language instructions refer to the audio processing requirements described by the user in spoken language. Expressing audio processing requirements through daily communication can significantly lower the threshold for audio operation. Raw audio refers to the audio input that needs to be processed, which clearly defines the processing object, facilitating targeted processing of the processing object.

[0038] During specific implementation, users only need to input natural language instructions and original audio into the client, and the server can obtain the natural language instructions and original audio from the client and process the original audio. For example, in the financial field, staff can input the natural language instruction "remove the background noise of this customer consultation recording" and upload the original customer consultation recording. The server can then start audio processing for this instruction and recording. In the medical field, doctors can input the instruction "enhance the volume of the expert's explanation in this medical lecture audio" and upload the original audio of the lecture. The server can then optimize the lecture audio. The user operation is convenient, and there is no need to spend a lot of time and energy learning the professional terminology of audio processing, which improves the usability of audio processing and enables non-professionals to easily operate audio processing.

[0039] S20: Parse the natural language instructions and the original audio to generate structured processing instructions.

[0040] Understandably, natural language instructions are often flexible, abstract, and ambiguous. For example, a user might say, "Make this audio sound more engaging," but such a statement lacks clear operational guidance for the system. Raw audio itself contains rich acoustic features, such as volume variations, frequency distribution, and timbre. Without analysis, the system has no way of knowing which parts the user wishes to process. Step S20 converts the user's spoken request into structured processing instructions that the system can understand and execute, clarifying the specific content of the audio processing and laying the foundation for subsequent, precise audio processing.

[0041] In some embodiments of the present invention, Figure 3 As shown, a specific parsing solution is provided, in S20, that is, parsing the natural language instructions and the original audio to generate structured processing instructions, which specifically includes the following steps S21-S22.

[0042] S21: Using a large multimodal language model to synchronously analyze the semantics of the natural language instruction and the audio features of the original audio.

[0043] Among them, the multimodal language large model is an artificial intelligence model that integrates the processing capabilities of multiple data modalities such as text, audio, and images. It achieves a more comprehensive understanding and processing of complex information through joint learning and analysis of data of different modalities. Specifically, the multimodal language large model is based on deep learning architectures, such as Transformer, etc., and through training with a large amount of data, it learns the association and mapping relationship between data of different modalities, thereby having powerful semantic understanding and feature extraction capabilities. In step S21, the processing capabilities of the multimodal language large model for natural language text and audio data are mainly utilized to combine the description of the user's spoken needs with the actual characteristics of the audio, and to unearth the deep needs and processing directions hidden behind the instructions.

[0044] For step S21, the multimodal language model can simultaneously receive and process data in two different modalities, text and audio, such as the natural language instructions and original audio in this embodiment. It can also use powerful deep learning algorithms and massive training data to explore the potential connection between the semantics of the instructions and the audio features. For example, in a financial application scenario, when a user enters the instruction "make this anti-fraud promotion audio more passionate," the multimodal language model not only understands the semantic requirement of "passion," but also simultaneously analyzes the original audio's speech speed, intonation, volume changes, and other features to determine in which aspects the current audio differs from the "passion" effect, thereby formulating a more targeted processing strategy. For another example, in a medical application scenario, for a medical case discussion audio, the user enters "highlight the expert's explanation of the key basis for diagnosis." The multimodal language model, while parsing the instruction, analyzes the voice characteristics of different speakers in the discussion audio, the time distribution of the discussion content, and other aspects, accurately locating the paragraph where the expert explains the key basis for diagnosis and providing accurate information for subsequent processing. This synchronous parsing method can avoid the one-sidedness and inefficiency brought about by single analysis, making the system's understanding of user needs more comprehensive and in-depth, thereby providing strong support for the subsequent generation of accurate structured processing instructions.

[0045] S22: Output a structured processing instruction, and align the structured processing instruction with the audio feature; wherein the structured processing instruction includes an operation type and a target effect.

[0046] Step S22 integrates and standardizes the results of the previous analysis, providing clear, unambiguous, and actionable guidance for subsequent audio processing. Specifically, by organizing the information obtained from parsing natural language instructions and audio features into structured processing instructions that include the operation type and target effect, the system can understand and execute user needs in a standardized manner.

[0047] Among them, the operation type clarifies the specific processing method, such as noise reduction, mixing, editing, adding special effects, etc., and provides a basis for selecting appropriate processing tools; the target effect indicates the state that the audio should reach after processing, such as clearer, more vivid, etc., which is an important reference for setting processing parameters. In addition, the target effect also includes concrete requirements, such as how clear the audio should be and which time period the audio should be cropped.

[0048] Aligning structured processing instructions with audio features further ensures processing accuracy, allowing the system to perform actions based on the actual audio. For example, when processing a conference recording containing multiple speakers, by aligning the target effect with the audio features, the system can accurately identify each speaker's speaking period and apply different processing strategies to different time periods, such as noise reduction during periods with high background noise and volume enhancement for key speech sections. This helps ensure that the entire processing process is orderly, accurate, and efficient.

[0049] Preferably, the structured processing instructions are in JSON format.

[0050] The following is an example implementation of a structured processing instruction:

[0051]

[0052] Among them, "action": "edit_music" means performing an audio processing operation; "steps" is an array containing multiple editing steps, which are executed in sequence. Each step contained in the array is an operation type; "type" is the operation type of this embodiment, which is represented here as one of the steps in "steps".

[0053] "type":"trim" means trimming the audio segment.

[0054] "start":"1:30" indicates the start time of the capture, which is 1 minute and 30 seconds.

[0055] "end":"2:45" indicates the end time of the capture, which is 2 minutes and 45 seconds.

[0056] In this embodiment, "start" and "end" are both target effects. Therefore, this step is to intercept a segment from 1 minute 30 seconds to 2 minutes 45 seconds from the original audio.

[0057] "type":"add_effect" means adding sound effects.

[0058] "effect":"reverb" indicates that the sound effect type to be added is reverb.

[0059] "intensity":"medium" indicates that the reverb intensity is medium.

[0060] In this embodiment, both "effect" and "intensity" are target effects. Therefore, this step adds a medium-intensity reverb effect to the audio.

[0061] S30: Dynamically matching a processing tool chain according to the structured processing instruction, and converting the structured processing instruction into a processing parameter.

[0062] For step S30, by converting the structured processing instructions into a specific executable processing tool chain and processing parameter combination, the audio processing work is ensured to be carried out smoothly; converting natural language instructions into processing parameters helps to break through creative limitations, further reduces the difficulty of audio processing, and improves the ability to realize abstract or complex creative needs, which is beneficial for financial and medical staff to complete audio processing work efficiently and conveniently. It is understandable that in the field of audio processing, there are many processing tools with different functions, such as noise reduction tools, mixing software, audio editors, arrangement tools, etc., each tool is suitable for different processing scenarios and needs. By dynamically matching the processing tool chain, the system can match the appropriate processing tool chain according to the structured processing instructions, so that the system can adopt the processing tool chain and process the original audio according to the processing parameters.

[0063] For example, when a natural language instruction requests noise reduction and mixing of audio, the system will first match the appropriate noise reduction tool, then select the corresponding mixing software, and determine the order in which they are executed, forming a complete processing tool chain. At the same time, converting structured processing instructions into processing parameters is key to ensuring that the processing effect meets user expectations. Different processing tools require specific parameter settings to achieve the desired effect, such as the noise reduction intensity of the noise reduction tool and the volume ratio of the mixing software. By obtaining accurate processing parameters, the system enables each processing tool to operate in an optimal state, thereby ensuring processing quality and results.

[0064] In some embodiments of the present invention, Figure 4 As shown, a specific matching solution is provided, in S30, that is, dynamically matching the processing tool chain according to the structured processing instruction, specifically including the following steps S31-S38.

[0065] S31: Pre-build a mapping relationship table between processing types and processing tool libraries.

[0066] Step S31 is the foundation and prerequisite for implementing a dynamically matched processing toolchain. Its core purpose is to establish a fast indexing mechanism between processing types and processing tools, ensuring that the system can efficiently and accurately match the appropriate tool for different audio processing requirements. Audio processing involves a wide variety of processing types, including noise reduction, mixing, editing, speed change, and pitch shifting. Each processing type corresponds to a variety of processing tools, each with varying functionalities, applicable scenarios, and processing effects. By pre-establishing a mapping table, the system associates processing types with specific tools in the processing tool library, creating a one-to-one mapping relationship. For example, the "noise reduction" processing type is associated with professional noise reduction software such as Audacity and RX, while the "mixing" processing type is associated with mixing tools such as FL Studio and Cubase. When the system receives an operation type in a structured processing instruction, it can quickly locate the corresponding processing tool based on the mapping table, eliminating the need to blindly search through the vast tool library. This significantly reduces tool matching time and improves audio processing efficiency. In addition, the mapping relationship table can be dynamically updated and optimized according to actual needs. As new processing tools emerge or the functions of existing tools are upgraded, the mapping relationship can be adjusted in a timely manner to ensure that the system can always use the most advanced and appropriate tools for audio processing.

[0067] S32: Obtain the operation type of the structured processing instruction.

[0068] S33: Query the mapping relationship table in sequence for processing types that semantically match the operation type.

[0069] For steps S32-S33, after obtaining the operation type of the structured processing instruction, a semantic matching query is performed on the processing type in the pre-built mapping relationship table to further standardize the operation type. Specifically, since the expression of the operation type may be relatively broad or diverse, and the processing types in the mapping relationship table are limited and have been classified and standardized, the operation type in the user instruction can be matched with the standard processing type in the mapping relationship table through semantic matching, ensuring that the system finds the processing method that best meets the user's needs. For example, the operation type input by the user may be "make the sound clearer", which is semantically related to the processing types such as "noise reduction" and "sound quality enhancement" in the mapping relationship table. Through query and matching, the system can determine the specific processing type, thereby more targeted screening of processing tools. In addition, the semantic matching process also takes into account the diversity and flexibility of language expression. Even if there are certain differences in the expression of the operation type, the corresponding processing type can be found through semantic analysis.

[0070] In one embodiment, a fine-tuned multimodal language large model can be used to perform the operation of step S33. During the specific implementation, the mapping relationship table can be checked first to see whether there is a processing type that is the same as the operation type. If so, the acquisition operation of the processing tool can be continued. If not, synonym matching can be performed to calculate the semantic similarity between the operation type and the processing type, and then obtain a processing type that semantically matches the operation type; if the synonym matching operation cannot find the processing type, model reasoning can be generated to infer a processing type with high accuracy in matching the user's expected operation type from multiple groups of candidate types.

[0071] S34: According to the processing type, obtain a corresponding processing tool from the processing tool library.

[0072] S35: Generate the processing tool chain by using the processing tools in the order of the operation types.

[0073] For steps S34-S35, the execution logic sequence of audio processing is established through the processing tool chain to ensure that the processing tasks of multi-target effects can be carried out in an orderly manner. Audio processing usually involves multiple operations, such as noise reduction before mixing, and there are dependencies between different operations. For example, audio editing must be completed before volume normalization can be performed. By analyzing the order of operation types in structured instructions, the system can automatically arrange the tool execution process to avoid processing failures caused by sequence errors. In addition, the generation of the processing tool chain also needs to consider the data compatibility between tools, such as ensuring that the output format of the previous tool can be directly read by the next tool, reducing the time loss and sound quality loss caused by format conversion; at the same time, the setting of the processing tool chain can call the processing tools in sequence, unify the workflow, avoid users manually switching between multiple software, and improve audio processing efficiency.

[0074] In specific implementation, for example, in financial applications, when a user wishes to process financial training audio, if the order of operations is "noise reduction - speech rate adjustment - adding background music," the system will sequentially call Audacity, Time Stretch, and FLStudio processing tools to generate a processing tool chain. For another example, in medical applications, when a user needs to process live surgical audio, the expected order of operations is "audio separation - noise reduction - sound quality enhancement." The system will then call the Spleeter processing tool to separate the doctor's voice from the instrument sound, then call RX for noise reduction, and finally call Adobe Audition for equalization to generate a processing tool chain.

[0075] In one embodiment, a processing tool corresponding to a processing type is retrieved from the processing tool library using a preset priority rule. It is understood that a processing tool may have multiple processing functions, such as an FFmpeg processing tool that supports editing, format conversion, and adding basic effects, while an Audacity processing tool supports editing and effects processing but may be slower. Therefore, a processing tool is matched using a preset priority rule.

[0076] During implementation, priority rules can be set based on the strengths of each processing tool, the specific needs of the user, the frequency of tool invocation, etc. For example, in this embodiment, operations requiring high precision, such as voice extraction, are predefined to prioritize AI processing tools such as Demucs; and batch processing tasks, such as noise reduction, are predefined to prioritize parallelization tools such as RX BatchProcessor.

[0077] In some embodiments of the present invention, Figure 5 As shown, a specific conversion solution is provided, in S30, that is, converting the structured processing instruction into processing parameters, which specifically includes the following steps S36-S39.

[0078] S36: Pre-construct a mapping knowledge base that associates processing effects with processing parameters.

[0079] S37: Obtain the target effect of the structured processing instruction.

[0080] S38: Query the processing effects that semantically match the target effect from the mapping knowledge base in sequence.

[0081] S39: According to the processing effect, corresponding processing parameters are obtained from the mapping knowledge base.

[0082] For steps S36-S39, the parameterized conversion of the target effect, such as the matching of the operation type and the processing tool, is similar. The target effect is semantically parsed, and then the processing effect that is the same or similar to the target effect is searched from the mapping knowledge base to obtain the corresponding processing parameters.

[0083] In this embodiment, the mapping knowledge base collects a large number of audio processing cases and establishes a mapping relationship between processing effects and processing parameters. For example, the "retro tape effect" corresponds to parameters such as "sampling rate reduced to 44.1kHz", "adding tape hiss noise", and "low-frequency attenuation 2dB". This mapping mechanism avoids the technical threshold of users manually adjusting complex parameters, and is especially suitable for non-professional users to quickly achieve the expected results. In addition, the mapping knowledge base supports automatic updates of machine learning algorithms, optimizes parameter combinations by analyzing user historical processing data, and improves the accuracy of processing effects. In the financial and medical fields, the automated conversion of target effects can ensure the consistency of processing results. For example, all compliant recordings use unified noise reduction parameters to avoid inconsistent quality inspection standards due to differences in manual settings.

[0084] The following is an implementation example of converting target effects into processing parameters:

[0085] Target effect: "strong sense of space";

[0086] Processing effect: "enhance the sense of space";

[0087] The processing parameters for the Enhance Spatial mapping include:

[0088] Reverb parameters: decay_time = 2.8s, wet_dry_mix = 45%;

[0089] Delay parameters: feedback=35%, low_pass_cutoff=5kHz.

[0090] In one embodiment, a fine-tuned multimodal language large model can be used to perform the operation of step S38. During the specific implementation, it is possible to first search in the mapping knowledge base whether there is a processing effect that is the same as the target effect. If so, the processing parameter acquisition operation can be continued. If not, synonym matching can be performed to calculate the semantic similarity between the target effect and the processing effect, and then obtain a processing effect that semantically matches the target effect. If the synonym matching operation cannot obtain the processing effect, a model inference can be generated to infer a processing effect with high accuracy that is consistent with the target effect expected by the user from multiple groups of candidate effects.

[0091] S40: Process the original audio according to the processing tool chain and processing parameters, and generate target audio.

[0092] In step S40, the raw audio is processed by invoking the processing tool chain and processing parameters, ensuring the efficiency and stability of the processing flow. In this embodiment, the system adopts a pipeline architecture, treating each processing tool in the processing tool chain as an independent processing node. The raw audio data flows through each node in sequence, and the corresponding processing operation is performed in turn.

[0093] In some embodiments of the present invention, Figure 6 As shown, a specific conversion solution is provided, in S40, that is, processing the original audio according to the processing tool chain and processing parameters, and generating the target audio, which specifically includes the following steps S41-S42.

[0094] S41: Generate a directed acyclic graph according to the processing tool chain.

[0095] A directed acyclic graph (DAG) is a data structure consisting of nodes and directed edges. Specifically, nodes represent tasks or operations, and directed edges represent dependencies between nodes. There are no cycles in the graph. In this embodiment, each node in the DAG represents a processing tool, and directed edges represent the order in which the processing tools are used. The DAG ensures that each processing tool is launched only after its dependent predecessor processing tool has completed, avoiding resource contention and processing errors.

[0096] For example, in financial application scenarios, when processing cross-border financial transaction telephone recordings, the processing tool chain is "voice transcription -> sensitive word filtering -> encrypted storage". The generated directed acyclic graph shows that the sensitive word filtering tool must be started after the voice transcription is completed, and the encrypted storage depends on the output of the first two steps; in financial application scenarios, through directed acyclic graph visualization, the bank's compliance department can also find that the encrypted storage node in the original processing tool chain is started in advance, so that the order can be corrected in time, avoiding the risk of plaintext storage of un-massified data; in medical application scenarios, when processing audio interpretation of genetic testing reports, the processing tool chain is "audio enhancement -> speech recognition -> medical terminology standardization". The directed acyclic graph shows that the medical terminology standardization tool requires the text results of speech recognition as input; in medical application scenarios, when monitoring the audio processing process, the directed acyclic graph finds that the output format of the speech recognition node is incompatible with the standardization tool, which is conducive to adjusting the interface parameters in advance and avoiding errors during batch processing.

[0097] S42: calling processing tools and processing parameters according to the directed acyclic graph to process the original audio and generate target audio.

[0098] For step S42, the calling mechanism of the processing tool based on the directed acyclic graph is the key to ensuring that the audio processing process is strictly executed according to the target, and at the same time, an automated, non-human intervention processing process is realized. In specific implementation, according to the node order of the directed acyclic graph, the processing tools are activated and data is transferred in sequence. For example, the noise reduction tool is first started to process the original audio, generating an intermediate file after noise reduction, and then the file is automatically imported into the mixing tool for the next step of processing. This reduces the time loss of manually triggering tools in traditional batch processing and is particularly suitable for financial call centers or medical service centers with large audio processing volumes. In addition, the directed acyclic graph of this embodiment supports a fault-tolerant mechanism. If the node where a processing tool is located fails to process, such as crashing due to incompatible audio formats, the system automatically rolls back to the previous node and attempts to reprocess using a backup processing tool, such as switching from Audacity to Audition processing tools, to ensure task completion rate. In financial and medical scenarios, this mechanism can ensure that the audio processing of critical audio such as emergency call recordings and major transaction negotiation records is not interrupted, meeting the industry's strict requirements for data integrity.

[0099] The following is an example implementation of directed acyclic graph generation:

[0100] graph TB

[0101] A[original audio]—>B[FFmpeg time cropping]

[0102] B—>C [SoX adds reverb]

[0103] C—>D [AudioLDM generates background sound]

[0104] D—>E[Multi-track mixed output].

[0105] Among them, graph TB is the syntax of the flowchart, which means creating a flowchart that flows vertically from top to bottom; the arrow "—" clearly indicates the execution direction, "A", "B", "C", "D" and "E" represent five nodes, and nodes A->E are executed in sequence. These five nodes executed in sequence represent: original audio, time cropping of original audio through FFmpeg processing tool, adding reverb to the audio processed by node B through SoX processing tool, adding background sound to the audio processed by node C through AudioLDM processing tool, and multi-track mixing output of audio with multiple tracks processed by node D, with the output as the target audio.

[0106] S50: Verify the accuracy of the target audio.

[0107] It is understood that during the audio processing process, due to deviations in the understanding of natural language commands, errors in processing tools, or inaccurate parameter settings, the generated target audio may suffer from problems such as distorted sound quality and content that deviates from user requirements. Step S50 verifies the accuracy of the target audio to ensure that the audio processing results meet user expectations and satisfy actual application needs.

[0108] In some embodiments of the present invention, Figure 7 As shown, a specific verification scheme is provided, in S50, that is, verifying the accuracy of the target audio, which specifically includes the following steps S51-S54.

[0109] S51: Detect the distortion of the target audio through audio fingerprint comparison technology.

[0110] Among them, audio fingerprint matching technology generates digital fingerprints by extracting unique acoustic features (such as spectrograms, Mel-frequency cepstral coefficients, etc.) from audio files, and uses hash algorithms, dynamic time warping and other technologies to compare the similarity of different audio fingerprints. If the fingerprint difference between the target audio and the original audio exceeds a threshold, it is determined that there is distortion. An audio fingerprint is a unique set of digital features of an audio segment, containing information such as spectral distribution, rhythmic patterns, and time series. It is similar to the uniqueness of human fingerprints and can achieve non-invasive detection without manual monitoring of each segment, greatly improving detection efficiency. It is especially suitable for scenarios that process large amounts of audio data. In financial and medical applications, distortion detection can ensure the authenticity and integrity of audio evidence and medical records, and avoid the credibility of key information affected by sound quality problems.

[0111] S52: Input the output audio into the multimodal language model to verify the semantic matching degree between the target audio and the natural language instruction.

[0112] For steps S51-S52, a multi-dimensional evaluation mechanism is used to systematically test the target audio to ensure the accuracy of the audio processing. On the one hand, the physical quality of the audio is tested to ensure that the processing process has not caused irreversible damage to the original audio; on the other hand, the semantic consistency of the audio content and the user's instructions is verified to avoid deviations from the user's original expected processing results. By quantitatively evaluating the distortion and semantic matching and generating an accuracy score, it not only provides an objective measurement standard for audio quality, but also helps the system identify potential defects in the processing flow and provides data support for subsequent optimization and iteration. Especially in fields such as finance and healthcare that have extremely high requirements for the accuracy and reliability of audio content, strict accuracy verification can avoid serious consequences such as compliance risks, misdiagnosis and misjudgment caused by audio quality issues, and ensure the security and professionalism of business processes.

[0113] S53: Generate the accuracy score according to the distortion and the semantic matching degree.

[0114] S54: The target audio whose accuracy score reaches or exceeds the preset threshold is recorded as passed verification, and the target audio whose accuracy score does not reach the preset threshold is recorded as failed verification.

[0115] In steps S53-S54, generating an accuracy score is a key step in quantifying and integrating the two core metrics of distortion and semantic match, aiming to comprehensively assess the quality of the target audio using a unified numerical standard. Distortion reflects the physical integrity of the audio, while semantic match measures the fit between the content and user requirements. These two metrics reflect the reliability of the processing results from technical and semantic perspectives, respectively. In an optional embodiment, the test results of these two metrics can be converted into a 0-100 scoring system using a preset weighting mechanism, such as a 40% weighting for distortion and a 60% weighting for semantic match. For example, if the distortion test result is 90 and the semantic match result is 80, the final accuracy score calculated based on the weighting is 84, indicating that the audio processing requirements for the original audio are generally met. This quantitative scoring not only facilitates the system's rapid screening of qualified audio but also provides users with intuitive quality feedback, enabling them to adjust processing strategies based on the scoring criteria. In the financial and medical fields, standardized accuracy scores can serve as a unified acceptance standard for audio quality, ensuring comparability of results across different processing tasks and reducing manual review costs.

[0116] Setting scoring thresholds and conducting verification assessments is a mechanism for automating audio quality assessment. For example, a pre-set threshold of 80 points serves as the dividing line between acceptable and unacceptable target audio. This threshold can be flexibly adjusted to meet the needs of different application scenarios, such as financial compliance recordings requiring a higher threshold. This automated assessment can quickly distinguish between target audio that meets the standards and audio that requires rework, eliminating the subjectivity and efficiency bottlenecks of manual review.

[0117] S60: Feedback the verification result of the target audio.

[0118] For step S60, feedback of verification results is a link in building a closed loop of interaction between the user and the system, so that the evaluation information of the audio processing can be delivered to the user in a timely manner to support subsequent decision-making and process optimization. In specific implementation, for audio that has passed the verification, the system can feedback the qualified conclusion and scoring details to the client to help users confirm that the processing results meet expectations so that they can be put into use directly; for audio that has not passed the verification, detailed reasons for failure can be fed back, such as "excessive distortion" or "insufficient semantic matching", and the user or system can be guided to automatically adjust the processing parameters and re-execute the processing task. This closed-loop mechanism not only improves the efficiency and consistency of audio processing, but also optimizes threshold settings and processing strategies through data accumulation, such as dynamically adjusting the weights of distortion and semantic matching based on historical verification data.

[0119] As you can imagine, feedback supports diverse outputs, including visual reports, text summaries, and API data, facilitating integration into existing user processing systems. Furthermore, feedback results can serve as a basis for system self-optimization, for example by analyzing the causes of frequent failures and automatically adjusting processing toolchains or parameter configuration strategies. In the financial and medical sectors, transparent feedback can enhance user trust in the system, ensuring that the audio processing process is traceable and auditable, meeting industry compliance requirements.

[0120] It can be seen that in the above scheme, by parsing natural language instructions to generate structured processing instructions, instead of traditional professional terminology input, the interaction threshold is lowered, the usability of audio processing is improved, and non-professionals can easily operate audio processing; the workflow is unified by dynamically matching the processing tool chain, avoiding users from manually switching between multiple software, thereby improving audio processing efficiency; converting natural language instructions into processing parameters helps to break through creative limitations, further reduces the difficulty of audio processing, and improves the ability to realize abstract or complex creative needs, which is beneficial for financial and medical personnel to complete audio processing work efficiently and conveniently.

[0121] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0122] In one embodiment, the present invention provides a semantic audio processing device 100, which corresponds to the semantic audio processing method in the above embodiment. Figure 8 As shown, the semantic audio processing device 100 includes a receiving module 101, a parsing module 102, a conversion module 103, a processing module 104, a verification module 105 and a feedback module 106. The functional modules are described in detail as follows:

[0123] The receiving module 101 is configured to receive natural language instructions and original audio input by a user.

[0124] The parsing module 102 is configured to parse the natural language instructions and the original audio to generate structured processing instructions.

[0125] The conversion module 103 is configured to dynamically match a processing tool chain according to the structured processing instruction and convert the structured processing instruction into a processing parameter.

[0126] The processing module 104 is configured to process the original audio according to the processing tool chain and processing parameters, and generate target audio.

[0127] The verification module 105 is configured to verify the accuracy of the target audio.

[0128] The feedback module 106 is configured to provide feedback on the verification result of the target audio.

[0129] In one embodiment, the parsing module 102 is specifically configured to:

[0130] Using a large multimodal language model to synchronously analyze the semantics of the natural language instructions and the audio features of the original audio;

[0131] Outputting a structured processing instruction and aligning the structured processing instruction with the audio feature; wherein the structured processing instruction includes an operation type and a target effect.

[0132] In one embodiment, the conversion module 103 is specifically configured to:

[0133] Pre-build a mapping table between processing types and processing tool libraries;

[0134] Obtaining an operation type of the structured processing instruction;

[0135] Querying the processing types that semantically match the operation type from the mapping relationship table in sequence;

[0136] According to the processing type, obtaining a corresponding processing tool from the processing tool library;

[0137] The processing tools are used to generate the processing tool chain in the order of the operation types.

[0138] In one embodiment, the conversion module 103 is further configured to:

[0139] Pre-build a mapping knowledge base that associates treatment effects with treatment parameters;

[0140] Obtaining a target effect of the structured processing instruction;

[0141] sequentially querying the processing effects that semantically match the target effect from the mapping knowledge base;

[0142] According to the processing effect, corresponding processing parameters are obtained from the mapping knowledge base.

[0143] In one embodiment, the processing module 104 is specifically configured to:

[0144] Generate a directed acyclic graph based on the processing tool chain;

[0145] The processing tools and processing parameters are called according to the directed acyclic graph to process the original audio and generate the target audio.

[0146] In one embodiment, the verification module 105 is specifically configured to:

[0147] Detecting the distortion of the target audio using audio fingerprint comparison technology;

[0148] Input the output audio into a multimodal language model to verify the semantic match between the target audio and the natural language instruction;

[0149] generating the accuracy score according to the distortion and the semantic matching;

[0150] The target audio whose accuracy score reaches or exceeds the preset threshold is recorded as passed verification, and the target audio whose accuracy score does not reach the preset threshold is recorded as failed verification.

[0151] For the specific limitations of the semantic audio processing device 100, please refer to the limitations of the semantic audio processing method above, which will not be repeated here. The various modules in the above-mentioned semantic audio processing device 100 can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0152] In one embodiment, a computer device 200 is provided. The computer device 200 may be a server, and its internal structure diagram may be as shown in FIG. Figure 9 As shown. The computer device 200 includes a processor 220, a memory and a network interface 250 connected via a system bus 210. Among them, the processor 220 of the computer device is used to provide computing and control capabilities. The memory of the computer device 200 includes a non-volatile and / or volatile storage medium, an internal memory 240. The non-volatile storage medium 230 stores an operating system 231, a computer program 232 and a database 233. The internal memory 240 provides an environment for the operation of the operating system and computer program in the non-volatile storage medium 230. The network interface 250 of the computer device 200 is used to communicate with an external client through a network connection. When the computer program is executed by the processor 220, it realizes the functions or steps of a semantic audio processing method server. That is, when the processor 220 executes the computer program, the following steps are implemented:

[0153] Receive natural language commands and original audio input from the user;

[0154] Parsing the natural language instructions and the original audio to generate structured processing instructions;

[0155] Dynamically matching a processing tool chain according to the structured processing instruction and converting the structured processing instruction into processing parameters;

[0156] The original audio is processed according to the processing tool chain and processing parameters to generate target audio.

[0157] In one embodiment, a computer device 300 is provided. The computer device 300 may be a client, and its internal structure diagram may be as follows: Figure 10 As shown. The computer device includes a processor 320, a memory, a network interface 350, a display screen 370 and an input device 360 ​​connected via a system bus 310. Among them, the processor 320 of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium 330 and an internal memory 340. The non-volatile storage medium 330 stores an operating system 331 and a computer program 332. The internal memory provides an environment for the operation of the operating system 331 and the computer program 332 in the non-volatile storage medium 330. The network interface 350 of the computer device 300 is used to communicate with an external server through a network connection. When the computer program is executed by the processor 320, it implements the functions or steps of the client side of a semantic audio processing method. That is, when the processor 320 executes the computer program 332, the following steps are implemented:

[0158] Receive natural language commands and original audio input from the user;

[0159] Parsing the natural language instructions and the original audio to generate structured processing instructions;

[0160] Dynamically matching a processing tool chain according to the structured processing instruction and converting the structured processing instruction into processing parameters;

[0161] The original audio is processed according to the processing tool chain and processing parameters to generate target audio.

[0162] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0163] Receive natural language commands and original audio input from the user;

[0164] Parsing the natural language instructions and the original audio to generate structured processing instructions;

[0165] Dynamically matching a processing tool chain according to the structured processing instruction and converting the structured processing instruction into processing parameters;

[0166] The original audio is processed according to the processing tool chain and processing parameters to generate target audio.

[0167] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant description in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0168] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0169] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0170] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A semantic audio processing method, characterized in that: include: Receive natural language commands and original audio input from the user; Parsing the natural language instructions and the original audio to generate structured processing instructions; Dynamically matching a processing tool chain according to the structured processing instruction and converting the structured processing instruction into processing parameters; The original audio is processed according to the processing tool chain and processing parameters to generate target audio.

2. The semantic audio processing method according to claim 1, characterized in that The step of parsing the natural language instruction and the original audio to generate a structured processing instruction includes: Using a large multimodal language model to synchronously analyze the semantics of the natural language instructions and the audio features of the original audio; Outputting a structured processing instruction and aligning the structured processing instruction with the audio feature; wherein the structured processing instruction includes an operation type and a target effect.

3. The semantic audio processing method according to claim 2, characterized in that The dynamically matching a processing tool chain according to the structured processing instruction and converting the structured processing instruction into a processing parameter includes: Pre-build a mapping table between processing types and processing tool libraries; Obtaining an operation type of the structured processing instruction; Querying the processing types that semantically match the operation type from the mapping relationship table in sequence; According to the processing type, obtaining a corresponding processing tool from the processing tool library; The processing tools are used to generate the processing tool chain in the order of the operation types.

4. The semantic audio processing method according to claim 3, characterized in that The dynamically matching a processing tool chain according to the structured processing instruction and converting the structured processing instruction into a processing parameter further includes: Pre-build a mapping knowledge base that associates treatment effects with treatment parameters; Obtaining a target effect of the structured processing instruction; sequentially querying the processing effects that semantically match the target effect from the mapping knowledge base; According to the processing effect, corresponding processing parameters are obtained from the mapping knowledge base.

5. The semantic audio processing method according to claim 4, characterized in that: The processing of the original audio according to the processing tool chain and the processing parameters to generate the target audio includes: Generate a directed acyclic graph based on the processing tool chain; The processing tools and processing parameters are called according to the directed acyclic graph to process the original audio and generate the target audio.

6. The semantic audio processing method according to claim 1, characterized in that: After processing the original audio according to the processing tool chain and the processing parameters and generating the target audio, the method further includes: Verifying the accuracy of the target audio; Feedback the verification result of the target audio.

7. The semantic audio processing method according to claim 6, characterized in that: Verifying the accuracy of the target audio includes: Detecting the distortion of the target audio using audio fingerprint comparison technology; Input the output audio into a multimodal language model to verify the semantic match between the target audio and the natural language instruction; generating the accuracy score according to the distortion and the semantic matching; The target audio whose accuracy score reaches or exceeds the preset threshold is recorded as passed verification, and the target audio whose accuracy score does not reach the preset threshold is recorded as failed verification.

8. A semantic audio processing device, characterized in that: include: A receiving module, configured to receive natural language commands and original audio input by a user; A parsing module, configured to parse the natural language instructions and the original audio to generate structured processing instructions; a conversion module, configured to dynamically match a processing tool chain according to the structured processing instruction and convert the structured processing instruction into a processing parameter; A processing module is used to process the original audio according to the processing tool chain and processing parameters, and generate target audio.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the semantic audio processing method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the semantic audio processing method according to any one of claims 1 to 7 are implemented.