Data processing method and device, equipment, medium and product

By using multimodal fusion processing of audio and visual data, the problem of accuracy in separating multi-source audio signals was solved, achieving higher separation accuracy and reliability.

CN121708951APending Publication Date: 2026-03-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411311422.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing technologies have low accuracy and poor separation effect in multi-source audio signal recognition.

Method used

By combining audio and visual data in a multimodal fusion processing approach, sound source separation is performed on multi-source audio signals. Multimodal fusion processing and optimization techniques, including speech alignment and splicing, and semantic analysis, are used to improve the accuracy of audio signal separation.

Benefits of technology

It improves the accuracy and reliability of separating multi-source audio signals, reduces semantic or syntactic errors, and ensures semantic accuracy and coherence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708951A_ABST
    Figure CN121708951A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, equipment, a medium and a product, and the method comprises the steps: obtaining multi-modal data collected in a business scene, and the multi-modal data comprises audio data and visual data; the audio data comprises multi-sound-source audio signals obtained by carrying out audio signal collection on N objects in the service scene, the visual data comprises visual signals of the N objects synchronously collected in the process of collecting the multi-sound-source audio signals, and N is an integer greater than 1; based on the audio data and the visual data, sound source separation is carried out on the multi-sound-source audio signals in a multi-modal fusion processing mode, an audio separation result is obtained, and the audio separation result comprises the audio signal of each object; and performing optimization processing on the audio signal of each object in the audio separation result to obtain an optimized audio signal of each object. The audio signals and the visual signals are fused, sound source separation is carried out on the multi-sound-source audio signals, and the accuracy of audio separation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a data processing method, apparatus, device, medium and product, specifically to a data processing method, a data processing apparatus, a computer device, a computer-readable storage medium and a computer program product. Background Technology

[0002] In intelligent voice recognition scenarios such as autonomous driving and smart homes, voice interaction between users and the system is often involved. The system needs to perform voice recognition to process the user's instructions. During the voice recognition process, the audio signals collected by the system may involve mixed audio signals generated by multiple objects (i.e., multi-source audio signals). Therefore, how to separate the audio signals from multiple sources to identify the audio signal of each object is a hot topic in voice recognition.

[0003] Currently, the common method is to directly identify and analyze the audio signals from multiple sound sources to separate the audio signals of each object. However, this method of sound source separation has low accuracy and poor separation effect. Summary of the Invention

[0004] This application provides a data processing method, apparatus, device, medium, and product that can perform source separation of multi-source audio signals through multimodal fusion processing of audio and visual signals, thereby improving the accuracy of audio separation.

[0005] On one hand, embodiments of this application provide a data processing method, the method comprising:

[0006] Acquire multimodal data collected in the business scenario. The multimodal data includes audio data and visual data. The audio data includes multi-source audio signals obtained by collecting audio signals from N objects in the business scenario. The visual data includes visual signals of N objects that are simultaneously collected during the process of collecting the multi-source audio signals. N is an integer greater than 1.

[0007] Based on audio and visual data, a multimodal fusion processing method is used to separate the audio signals from multiple sources, resulting in audio separation results; the audio separation results include the audio signal of each object.

[0008] The audio signal of each object in the audio separation result is optimized separately to obtain the optimized audio signal of each object.

[0009] On one hand, embodiments of this application provide a data processing apparatus, the apparatus comprising:

[0010] The acquisition unit is used to acquire multimodal data collected in the business scenario. The multimodal data includes audio data and visual data. The audio data includes multi-source audio signals obtained by acquiring audio signals from N objects in the business scenario. The visual data includes visual signals of N objects that are simultaneously acquired during the acquisition of the multi-source audio signals. N is an integer greater than 1.

[0011] The processing unit is used to perform source separation on multi-source audio signals based on audio data and visual data using a multimodal fusion processing method, and obtain audio separation results; the audio separation results include the audio signal of each object;

[0012] The processing unit is also used to optimize the audio signal of each object in the audio separation result to obtain the optimized audio signal of each object.

[0013] In one possible implementation, the processing unit uses a multimodal fusion processing method to separate the multi-source audio signals based on audio and visual data, obtaining audio separation results, which are then used to perform the following operations:

[0014] The audio data is processed by signal separation to obtain audio processed data, which includes audio signals of N objects separated from multi-source audio signals;

[0015] Visual data is processed visually to obtain visually processed data.

[0016] A multimodal fusion processing method is adopted to perform multimodal fusion on audio processing data and visual processing data in order to separate the audio signals from multiple sources and obtain the audio separation result.

[0017] In one possible implementation, the processing unit performs signal separation processing on the audio data to obtain processed audio data, which is then used to perform the following operations:

[0018] A sound source separation model is used to separate audio signals from multiple sound sources in audio data to obtain separated audio data; the sound source separation model includes any one of the following: a dual-channel speech separation model, an audio conversion model, and a multi-channel convolutional neural network model;

[0019] The separated audio data is subjected to audio optimization processing to obtain processed audio data.

[0020] In one possible implementation, the separated audio data includes first audio signals corresponding to N objects, where each first audio signal is represented as the i-th first audio signal, i being a positive integer and 1 ≤ i ≤ N; the audio processing data includes N second audio signals; the processing unit performs audio optimization processing on the separated audio data to obtain audio processing data, which is used to perform the following operations:

[0021] Beamforming technology is used to perform audio optimization processing on the i-th first audio signal in the separated audio data to obtain the second audio signal after optimization of the i-th first audio signal;

[0022] The optimized N second audio signals are combined into audio processing data;

[0023] The audio optimization processing includes one or more of the following: speech enhancement and noise suppression.

[0024] In one possible implementation, the visual processing data includes N visual feature signals, each visual feature signal being generated by an object; the processing unit performs visual processing on the visual data to obtain visual processing data, which is used to perform the following operations:

[0025] Visual data is processed for visual recognition to obtain N candidate visual signals; a candidate visual signal refers to an object that is synchronously generated by the object in the process of generating the corresponding audio signal.

[0026] N second audio signals are acquired, and time alignment processing is performed on the N candidate visual signals and the N second audio signals to correct the N candidate second audio signals, thereby obtaining the corrected N visual feature signals;

[0027] Among them, a second audio signal generated by an object is used to correct the candidate visual signal generated synchronously by the object.

[0028] In one possible implementation, the processing unit employs a multimodal fusion processing method to perform multimodal fusion on the audio processing data and the visual processing data to separate the audio signals from multiple sources, obtaining an audio separation result for performing the following operations:

[0029] Obtain environmental parameters for the business scenario;

[0030] The weighted fusion strategy corresponding to the business scenario is determined based on environmental parameters. The weighted fusion strategy includes the weight coefficients to be fused, which include the first weight of each second audio signal in the audio processing data and the second weight of each visual feature signal in the visual processing data.

[0031] Based on weighting coefficients, the second audio signal of any object in the audio processing data and the visual feature signal of the object are weighted and fused to obtain the audio separation result.

[0032] In one possible implementation, the processing unit uses a multimodal fusion processing method to separate the multi-source audio signals based on audio and visual data, obtaining audio separation results, which are then used to perform the following operations:

[0033] Acquire location data collected in the business scenario. Location data refers to the data obtained by using LiDAR technology to collect the spatial location of N objects in the business scenario.

[0034] Based on location data, frequency data, and visual data, a multimodal fusion processing method is used to separate the sound sources of multi-source audio signals, and the audio separation results are obtained.

[0035] In one possible implementation, the processing unit is also used to perform the following operations:

[0036] Multimodal fusion processing is performed on audio and visual data to separate the audio signals from multiple sources, resulting in the first separation result of the audio signals from multiple sources.

[0037] Multimodal fusion processing is performed on audio data and location data to separate the audio signals from multiple sources, resulting in a second separation result for the audio signals from multiple sources.

[0038] The first separation result and the second separation result are fused to obtain the audio separation result.

[0039] In one possible implementation, any one of the N objects is represented as object j, where j is a positive integer and 1 ≤ j ≤ N; the processing unit optimizes the audio signal of each object in the audio separation result to obtain the optimized audio signal of each object, which is used to perform the following operations:

[0040] Using speech alignment and splicing technology, the audio signal of object j in the audio separation result is optimized in the first stage to obtain the corrected audio signal of object j;

[0041] Semantic analysis techniques are used to perform secondary optimization processing on the modified audio signal of object j, resulting in the optimized audio signal of object j.

[0042] In one possible implementation, the processing unit employs speech alignment and splicing techniques to perform a first-level optimization on the audio signal of object j in the audio separation result, obtaining a corrected audio signal of object j, which is used to perform the following operations:

[0043] The audio signal of object j is aligned using voice alignment technology to obtain the aligned audio signal;

[0044] The aligned audio signals are spliced ​​using voice splicing technology to obtain the corrected audio signal for object j.

[0045] In one possible implementation, the processing unit employs semantic analysis techniques to perform secondary optimization processing on the optimized audio signal of object j, obtaining the optimized audio signal of object j, which is then used to perform the following operations:

[0046] The natural language processing model is invoked to perform semantic analysis on the modified audio signal of object j, resulting in the analyzed text.

[0047] If the analyzed text contains the text content of other objects besides object j among N objects, then the analyzed text is semantically corrected to obtain the optimized audio signal of object j.

[0048] In one possible implementation, optimizing the audio signal includes task instructions indicated by any object in the business scenario; the processing unit is also configured to perform the following operations:

[0049] The business processing system outputs the task instructions indicated by the optimized audio signals of each object in the business scenario; where the business scenario includes intelligent driving scenario or smart home scenario;

[0050] Obtain feedback data from at least one object in response to task instructions, and obtain environmental data surrounding the business scenario;

[0051] Based on feedback data and environmental data, the business processing system undergoes adaptive learning; the adaptively learned business processing system is then used to perform audio recognition on various objects within the business scenario.

[0052] On one hand, embodiments of this application provide a computer device, which includes a processor and a memory; the memory stores a computer program; when the computer program is executed by the processor, it performs the above-described data processing method.

[0053] On one hand, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the aforementioned data processing method.

[0054] On the one hand, embodiments of this application provide a computer program product, which includes a computer program. When the computer program is executed by a processor, it performs the above-described data processing method.

[0055] In this embodiment, multimodal data collected in a business scenario is acquired. This multimodal data includes audio data and visual data. The audio data includes multi-source audio signals obtained by collecting audio signals from N objects in the business scenario. The visual data includes visual signals of N objects simultaneously collected during the acquisition of multi-source audio signals, where N is an integer greater than 1. Based on the audio and visual data, a multimodal fusion processing method is used to separate the multi-source audio signals to obtain an audio separation result. This audio separation result includes the audio signal of each object. The audio signal of each object in the audio separation result is then optimized to obtain an optimized audio signal for each object. It can be seen that this application can perform multimodal fusion of video and audio data, thereby enabling the separation of multi-source audio signals by combining visual signals. Because the audio and visual signals are considered from multiple dimensions for source separation, the accuracy of source separation is improved. Furthermore, the optimization processing of the audio signals of each object in the audio separation result further improves the accuracy of the separated audio signals. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0057] Figure 1 This is a schematic diagram illustrating the principle of sound source separation processing provided in an embodiment of this application;

[0058] Figure 2 This is a schematic diagram of the architecture of a data processing system provided in an embodiment of this application;

[0059] Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application;

[0060] Figure 4a This is a schematic diagram of a process for acquiring multi-source audio signals provided in an embodiment of this application;

[0061] Figure 4b This is a schematic diagram of a process for acquiring visual data provided in an embodiment of this application;

[0062] Figure 5 This is a flowchart illustrating another data processing method provided in an embodiment of this application;

[0063] Figure 6 This is a schematic diagram of the structure and processing flow of a business processing system provided in an embodiment of this application;

[0064] Figure 7 This is a flowchart illustrating a data processing scenario provided in an embodiment of this application;

[0065] Figure 8 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;

[0066] Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0067] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0068] This application provides a data processing solution applicable to intelligent business scenarios such as intelligent driving and smart home scenarios. This solution can collect multimodal data (such as visual and audio data) in these scenarios and achieve multi-source audio signal separation based on multimodal data fusion processing. Specifically, it combines visual signals to separate multi-source audio signals, thereby improving the accuracy of audio separation. The principle of the data processing solution provided in this application is briefly described below:

[0069] (1) Acquire multimodal data collected in business scenarios (such as intelligent driving scenarios and smart home scenarios). Here, multimodal data includes business data in at least one modality. For example, multimodal data includes audio data and visual data. Here, audio data includes multi-source audio signals obtained by collecting audio signals from N objects in the business scenario. Here, visual data includes visual signals of N objects simultaneously collected during the process of collecting multi-source audio signals; N is an integer greater than 1.

[0070] (2) Based on audio and visual data, a multimodal fusion processing method is used to separate the audio signals from multiple sources, resulting in an audio separation result containing the audio signal of each object. The multimodal fusion processing method here refers to the process of separating the audio signals from multiple sources by combining visual signals. Since data from multiple modalities are combined, the effect of source separation can be improved.

[0071] (3) Optimize the audio signal of each object in the audio separation result to obtain the optimized audio signal of each object. Here, optimization processing such as speech recognition, semantic analysis, and semantic correction is performed on the audio signals of each object separated in the above steps. This can automatically detect and correct the parts of the audio signal mixed with other objects, ensuring the accuracy and coherence of the semantics.

[0072] As can be seen from the above, on the one hand, in the process of audio separation, this application can combine the visual signals of the object to perform multimodal fusion of audio signals from multiple sound sources. Since the data from both visual and audio modes are fused, the accuracy of audio separation can be improved. On the other hand, the audio signals of each object in the audio separation result can be further optimized to avoid semantic or syntactic errors, thereby further improving the accuracy and reliability of the separated audio signals.

[0073] The following is an introduction to the key technical terms involved in this application.

[0074] I. Multimodal data.

[0075] Multimodal data, as the name suggests, refers to business data collected in multiple modalities within a business scenario. These modalities include any one or more of audio, visual, and location modalities. Audio data is audio data, visual data is visual data, and location data is location data. Audio data includes multi-source audio signals obtained by collecting audio signals from N objects in the business scenario; visual data includes visual signals of N objects simultaneously collected during the multi-source audio signal acquisition process; and location data includes spatial location signals of N objects simultaneously collected during the multi-source audio signal acquisition process. Therefore, business data in any modality is used to characterize the signal generated by an object in the corresponding modality. Since multimodal data contains business data from at least one modality, this application can achieve multimodal data fusion, thereby enabling source separation of multi-source audio signals based on business data from multiple modalities.

[0076] II. Multi-source audio signals and visual signals.

[0077] Multi-source audio signals refer to audio signals generated by multiple sound sources. Here, one sound source is used to indicate the different spatial positions or directions of an object in a business scenario. Therefore, the audio signal of one sound source refers to the audio signal collected by a sensor (such as a microphone) targeting an object in the business scenario. In this application, multi-source audio signals refer to the data obtained after collecting audio signals from N objects (N is an integer greater than 1, such as 2, 3, 4, etc.) in a business scenario. Assuming the business scenario includes N objects (such as object 1, object 2, object 3), N microphones can be deployed in the business scenario to form a microphone array. One microphone in this array is used to collect audio signals from different directions in the business scenario to obtain multi-source audio signals. Since any microphone may simultaneously collect audio signals from different objects, such as a microphone collecting audio signal 1 from object 1, while simultaneously collecting audio signal 2 from object 2 and audio signal 3 from object 3, there may be overlapping or confused audio parts in the multi-source audio signals. Therefore, this application mainly involves source separation of multi-source audio signals generated by N objects to obtain the audio signal of each object, so that each object can perform corresponding business processing in the business scenario based on its own audio signal.

[0078] Visual signals refer to signals captured when an object is in a visual state (such as facial expressions, lip movements, etc.). Here, the visual signal is generated synchronously with the audio signal produced by the object; for example, if an object (such as a user) says, "You look beautiful today," the user's facial expressions and mouth movements (such as lip movements) can be captured simultaneously, and these captured facial expressions and mouth movements are used as the visual signal for that user. It can be seen that since visual signals and audio signals are generated synchronously, and the user's facial expressions and mouth movements are usually different depending on what they say, meaning different audio signals correspond to different visual signals, this application can use visual signals to provide auxiliary explanations of multi-source audio signals during the audio separation process, thereby improving the audio separation effect.

[0079] III. Sound source separation.

[0080] Sound source separation refers to the process of separating the audio signals of different objects from a multi-source audio signal. Since a multi-source audio signal is obtained by collecting audio signals from N objects, sound source separation means separating the audio signals of each of the N objects from the multi-source audio signal. Please see [link to relevant documentation]. Figure 1 , Figure 1 This is a schematic diagram illustrating the principle of sound source separation processing provided in an embodiment of this application. For example... Figure 1As shown, the principle of this sound source separation processing flow mainly involves separating multi-source audio signals. Here, the multi-source audio signals are audio signals collected from three objects: object 1, object 2, and object 3. That is, the multi-source audio signals are a mixed audio signal from the above three objects. The process of separating the multi-source audio signals mainly refers to separating and processing the audio signals generated by each object individually from the multi-source audio signals to obtain the individual audio signals of each object, such as audio signal 1 of object 1, audio signal 2 of object 2, and audio signal 3 of object 3.

[0081] Based on data modality, sound source separation methods include: single-modal sound source separation and multi-modal sound source separation. Single-modal sound source separation refers to the processing method that directly separates the sound source from the single-modal audio data. Multi-modal sound source separation refers to the processing method that separates the sound source from multiple audio sources based on multi-modal data (such as visual data and audio data). Optionally, multi-modal sound source separation can also include processing methods that separate the sound source from multiple audio sources based on visual data, location data, and audio data. Since multi-modal data can provide more dimensional (or modal) business data, the multi-modal sound source separation method (i.e., multi-modal fusion processing method) adopted in this application can improve the accuracy of sound source separation.

[0082] IV. Optimization of audio signal processing.

[0083] In this application, the optimization processing of audio signals mainly refers to the semantic analysis optimization process of the audio signals of each object after sound source separation. Since the audio signal here refers to the audio signal after sound source separation, there may be cases of incomplete or inaccurate separation during the sound source separation process. For example, the audio signal 1 obtained after separating object 1 may contain parts of the audio signal of other objects 2; or the audio signal 1 obtained after separating object 1 may contain truncated sentences or other incomplete signals. Therefore, this application requires further optimization processing of the audio signals after separating each object. This optimization processing includes semantic analysis processing based on NLP (Natural Language Processing) technology. NLP technology refers to the technology of computer processing and understanding user language, mainly involving speech recognition, semantic analysis, text understanding, and other processing technologies. Based on the above semantic analysis technology, semantic correction, content completion, and other optimization processing can be performed on the separated audio signals, thereby improving the accuracy and reliability of the audio signals of each separated object.

[0084] It should be specifically noted that the data involved in the data processing of this application (such as multimodal data, audio data, visual data, multi-source audio signals, visual signals, etc.) requires the permission or consent of the target user when the above embodiments of this application are applied to specific products or technologies. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the region, adhering to the principles of legality, legitimacy, and necessity, and not involving the acquisition of data types prohibited or restricted by laws and regulations. In some optional embodiments, the related data involved in the embodiments of this application is obtained after separate authorization from the target user. Additionally, when obtaining separate authorization from the target user, the intended use of the related data is explained to the target user.

[0085] The data processing system provided in this application will be described in detail below.

[0086] Please see Figure 2 , Figure 2 This is a schematic diagram of the architecture of a data processing system provided in an embodiment of this application. Figure 2 As shown in the diagram, the architecture of this data processing system may include at least: a server 204 and a terminal device cluster; wherein, the terminal device cluster includes: a first terminal device 201, a second terminal device 202, a third terminal device 203, and other multiple terminal devices; it should be understood that the number of terminal devices included in the terminal device cluster is only for illustrative purposes, and the embodiments of this application do not limit the number and type of terminal devices. Any terminal device in the terminal device cluster can be directly or indirectly connected to the server 204 via a network; the aforementioned network may include, but is not limited to: wired networks and wireless networks, wherein the wired network includes: local area networks, metropolitan area networks, and wide area networks, and the wireless network includes: Bluetooth, WIFI (Wireless Fidelity, a standard wireless local area network), and other networks that implement wireless communication.

[0087] The terminal devices in the terminal device cluster can be: mobile phones, tablets, laptops, desktop computers, gaming devices, in-vehicle devices, aircraft, wearable devices (such as smartwatches, smart bracelets, pedometers, etc.), virtual reality devices (such as VR (Virtual Reality) devices, AR (Augmented Reality) devices), etc. It is understood that the types of terminal devices in the terminal device cluster can be the same or different. For example, the first terminal device 201 can be a laptop, the second terminal device 202 can be a desktop computer, and the third terminal device 203 can be a tablet. This application does not limit the number or type of terminal devices in the terminal device cluster.

[0088] Server 204 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0089] The following describes the data interaction process between the first terminal device 201 and the server 204 during data processing, using any terminal device (e.g., the first terminal device 201) as an example:

[0090] ① The first terminal device 201 can deploy audio acquisition devices and video acquisition devices in any voice interaction business scenario (such as intelligent driving scenario, smart home scenario, etc.); wherein, the audio acquisition device (such as a microphone array) is used to collect audio data of N objects in the business scenario, and the video acquisition device (such as a camera array) is used to collect visual data of N objects in the business scenario. Furthermore, the first terminal device 201 can acquire multimodal data in the business scenario through the audio acquisition device (such as a microphone array) and the video acquisition device (such as a camera array), where N is an integer greater than 1.

[0091] ② The first terminal device 201 sends the collected multimodal data to the server 204; wherein, the multimodal data includes audio data and visual data; the audio data includes multi-source audio signals obtained by collecting audio signals from N objects in the business scenario, and the visual data includes visual signals of N objects simultaneously collected during the process of collecting multi-source audio signals.

[0092] ③ Server 204 uses a multimodal fusion processing method to separate the audio signals from multiple sources based on audio and visual data, and obtains an audio separation result that includes the audio signal of each object.

[0093] ④ Server 204 optimizes the audio signal of each object in the audio separation result to obtain the optimized audio signal of each object. For example, server 204 can use NLP technology to perform text recognition and semantic analysis on the audio signal, and then perform semantic correction and content completion optimization processes according to the text recognition and semantic analysis results.

[0094] ⑤ Server 204 returns the optimized audio signals of each object to the first terminal device 201. Subsequently, the first terminal device 201 can control the business processing system to perform business processing according to the optimized audio signals in the business scenario. For example, if the business scenario is an intelligent driving scenario, then the first terminal device 201 can be an intelligent vehicle. The intelligent vehicle can then control the vehicle system to perform corresponding operations according to the task instructions indicated by the optimized audio signals. For example, if the task instruction is "check in with the air conditioning", then the operation of turning on the air conditioning in the car can be performed.

[0095] It should be noted that the above data processing procedure is for illustrative purposes only and does not limit the specific execution process of the first terminal device and the server. Optionally, the first terminal device can perform the process of "separating the audio signals from multiple sources based on audio data and visual data using a multimodal fusion processing method to obtain audio separation results," and then send the audio separation results to the server, which will then perform the audio signal optimization processing; or, the complete data processing flow described above can be executed independently by the first terminal device or the server in the data processing system.

[0096] In one possible implementation, the data processing system provided in this application embodiment can be deployed on a blockchain node. For example, the first terminal device 201, the second terminal device 202, the third terminal device 203, and the server 204 can all be treated as blockchain node devices, jointly forming a blockchain network. Therefore, the data processing flow executed in this application can be executed on the blockchain, which can ensure the fairness and impartiality of the data processing flow, make the above flow traceable, and ensure data security during the data processing process, thereby improving the security and reliability of the entire data processing flow.

[0097] Based on the data processing system provided in this application, on the one hand, during the audio separation process, this application can combine the visual signals of the object to perform multimodal fusion of multi-source audio signals. Since the data from both visual and audio modes are fused, the accuracy of audio separation can be improved. On the other hand, the audio signals of each object in the audio separation result can be further optimized to avoid semantic or syntactic errors, thereby further improving the accuracy and reliability of the separated audio signals.

[0098] It is understood that the data processing system described in the embodiments of this application is for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0099] The specific embodiments of the data processing scheme of this application will be described in detail below with reference to the accompanying drawings.

[0100] Please see Figure 3 , Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application. This data processing method can be performed by a computer device (such as...). Figure 2 Execute on any of the terminal devices or servers shown.

[0101] like Figure 3 As shown, the data processing method includes, but is not limited to, the following steps S301-S303:

[0102] S301: Acquire multimodal data collected in the business scenario; multimodal data may include audio data and visual data.

[0103] Specifically, multimodal data can include audio data and visual data. Audio data includes multi-source audio signals obtained by acquiring audio signals from N objects in the business scenario, while visual data includes visual signals of N objects simultaneously acquired during the acquisition of multi-source audio signals; N is an integer greater than 1. The acquisition process for audio and video data is described in detail below.

[0104] (1) The process of acquiring audio data.

[0105] In one possible implementation, an audio acquisition device is deployed in the business scenario. This device is used to acquire audio signals from N objects within the scenario. For example, the audio acquisition device could be a microphone array, which consists of multiple microphones located at different spatial positions within the business scenario (e.g., different coordinates and directions), thus forming an array structure. Please see [link to relevant documentation]. Figure 4a , Figure 4a This is a schematic diagram illustrating a process for acquiring multi-source audio signals, provided in an embodiment of this application. Figure 4a As shown, a microphone array can be deployed in a business scenario. This array includes multiple microphones, such as microphone a, microphone b, microphone c, and microphone d. Each microphone can be deployed at a different spatial location within the business scenario and collect audio signals from that location. Figure 4aAs shown, each microphone in the microphone array is responsible for capturing audio signals from a specific direction (e.g., microphone a is used to capture audio signals from above in the business scenario, microphone b is used to capture audio signals from the front in the business scenario, and microphone c is used to capture audio signals from the right in the business scenario), so as to ensure that the audio signals generated by N objects located in different directions in the business scenario can be collected. Based on this, the microphone array can be used to collect and process audio signals from N objects at different spatial locations in the business scenario to obtain audio data.

[0106] Furthermore, based on the audio data, it is possible to determine the multi-source audio signals of N objects. Determining the multi-source audio signals includes the following two cases:

[0107] Scenario 1: Using audio data as a multi-source audio signal. Since a microphone array can capture audio signals generated by N objects in a business scenario, the collected audio data includes multi-source audio signals generated by these N objects. Therefore, the audio data collected by the microphone from the N objects can be directly used as the multi-source audio signal generated by the N objects.

[0108] Scenario 2: Audio preprocessing is performed on the audio data to obtain multi-source audio signals. This audio preprocessing includes one or more of the following: echo cancellation, noise suppression, normalization, and voice activity detection. Specifically: ① Echo Cancellation: A technique used to eliminate echoes. An echo occurs when sound played by a loudspeaker is picked up again by a microphone and transmitted back to the other party; ② Noise Suppression: A technique used to reduce background or environmental noise, often used to improve voice communication quality; ③ Voice Activity Detection (VAD): A technique used to detect the presence of active audio signals in multi-channel media audio, thereby improving the processing efficiency of audio data. Through these various audio preprocessing methods, the audio data acquired by the microphone array can be optimized to obtain multi-source audio signals, thus improving the audio quality of the multi-source audio signals to be separated, thereby enhancing the audio separation effect.

[0109] (2) The process of acquiring visual data.

[0110] In one possible implementation, a video capture device is deployed in the business scenario to capture visual signals from N objects within the scenario. For example, the video capture device could be a camera array, consisting of multiple cameras located at different spatial positions within the business scenario (e.g., different coordinates and orientations), thus forming an array structure. See also... Figure 4b , Figure 4b This is a schematic diagram of a process for acquiring visual data provided in an embodiment of this application. For example... Figure 4b As shown, the camera array deployed in the business scenario includes: Camera 1, Camera 2, Camera 3, Camera 4, and Camera 5. Different cameras are located in different spatial positions (e.g., directions) within the business scenario. For example, Camera 1 is located to the left, Camera 2 is located above, Camera 3 is located in front, Camera 4 is located to the right, and Camera 5 is located behind. Cameras in different directions are responsible for acquiring visual signals from objects in their respective directions. Since visual signals are generated alongside audio signals—that is, visual signals are generated only when an object generates audio signals—this application utilizes... Figure 4b The camera array shown can simultaneously acquire and process visual signals from N objects in a business scenario to obtain visual data during the generation of multi-source audio.

[0111] Optionally, after the camera array acquires visual signals from N objects in the business scenario, candidate visual data can be obtained. Then, image preprocessing is required on the candidate visual data to obtain the final visual data. This image preprocessing includes one or more of the following: normalization, image scaling, and image denoising. For example: ① Multiple images containing objects can be extracted from the candidate visual data, and each image can be normalized by mean normalization or variance normalization. Normalization unifies the image format, providing robust input images for data processing. ② Image scaling refers to reducing or enlarging the image size (e.g., length, width, height) to unify the size of images acquired by different cameras, improving the efficiency of visual data processing. ③ Image denoising can include extracting multiple images containing objects from the candidate visual data and then performing denoising operations on each image to improve image clarity. By using the various image preprocessing methods described above, the candidate visual data acquired by the camera array can be optimized to obtain visual data, thereby improving the quality of the visual signals of each object and facilitating multimodal fusion.

[0112] Optionally, multimodal data may also include location data, which is obtained by collecting data on the spatial locations of N objects in the business scenario.

[0113] (3) The process of collecting location data.

[0114] In one possible implementation, if the business scenario belongs to a predefined scenario (such as a high-precision or complex environment, for example, an extreme noise environment or a nighttime environment), a positioning device can be deployed in the business scenario to collect the spatial position signals of N objects in the business scenario. For example, the positioning device can be a LiDAR (Light Detection and Ranging) system. LiDAR can collect the three-dimensional spatial position (such as three-dimensional coordinates or orientation) signals of each object in the business scenario, thereby obtaining position data. In this implementation, since LiDAR has higher accuracy than cameras, it is generally suitable for low-light environments or scenarios with severe visual interference. Therefore, in the above scenarios, the position data collected by LiDAR can significantly improve the system's ability to locate sound sources in each object in the business scenario and improve the subsequent accuracy of separating multi-source audio signals.

[0115] S302: Based on audio data and visual data, a multimodal fusion processing method is used to separate the audio signals from multiple sources and obtain the audio separation result.

[0116] The audio separation result includes the audio signal of each of the N objects. Since the multimodal data includes business data acquired in the visual modality (i.e., visual data) and business data acquired in the audio modality (i.e., audio data), this application can achieve multimodal fusion processing based on visual data and audio data, thereby separating the audio signals from multiple sources to obtain the audio signal of each object.

[0117] In one possible implementation, the methods for source separation of multi-source audio signals mainly include: single-modal separation and multi-modal separation. Single-modal separation primarily refers to directly separating the sources of the multi-source audio signals based on audio data to obtain the audio separation result. Multi-modal separation primarily refers to separating the sources of the multi-source audio signals based on multi-modal data to obtain the audio separation result. This application mainly involves multi-modal separation methods, and based on different types of multi-modal data, multi-modal separation methods can be specifically divided into the following types:

[0118] Method 1: Audio + Vision Multimodal Separation. Specifically, Method 1 mainly involves source separation of multi-source audio signals based on audio and visual data to obtain audio separation results.

[0119] Method 2: Multimodal separation using audio, vision, and location. Specifically, Method 2 mainly involves source separation of multi-source audio signals based on audio data, visual data, and location data to obtain audio separation results.

[0120] Method 3: Multimodal separation using audio and location data. Specifically, Method 3 mainly involves source separation of multi-source audio signals based on audio data and location data to obtain audio separation results.

[0121] This application mainly involves the fusion processing of data based on two modalities, audio and vision, to separate the sources of multi-source audio signals. In one possible implementation, a computer device uses a multimodal fusion processing method based on audio data and visual data to separate the sources of multi-source audio signals and obtain the audio separation result, specifically including the following processes (1)-(3):

[0122] (1) Perform signal separation processing on the audio data to obtain processed audio data, which includes audio signals of N objects separated from multi-source audio signals. Specifically, the process of performing signal separation processing on the audio data includes the following steps ①-②:

[0123] ① A sound source separation model is used to separate audio signals from multiple sound sources in audio data, resulting in separated audio data. The sound source separation model includes any one of the following: a dual-channel speech separation model (such as a dual-path recurrent neural network (DPRNN) model), an audio conversion model (such as a Transformer model or other models based on the Transformer structure), or a multi-channel convolutional neural network model. It should be understood that the sound source separation model here can be any neural network model with sound source separation capabilities; this application does not specifically limit the structure of the sound source separation model. Through this sound source separation model, preliminary separation processing can be performed on the multi-source audio signals acquired by the microphone array, thereby separating the audio signals from different sources.

[0124] ② The separated audio data undergoes audio optimization processing to obtain processed audio data. The separated audio data includes first audio signals corresponding to N objects, where each first audio signal is represented as the i-th first audio signal, i being a positive integer and 1 ≤ i ≤ N; the processed audio data includes N second audio signals. Based on this, the process of the computer device performing audio optimization processing on the separated audio data to obtain processed audio data includes the following steps: using beamforming technology, the i-th first audio signal in the separated audio data is optimized to obtain an optimized second audio signal; then, the N optimized second audio signals are combined to form the processed audio data. The audio optimization processing here includes one or more of speech enhancement and noise suppression. Beamforming technology can further enhance audio signals acquired in a specific direction and suppress noise and interference in other directions, ensuring that the separated audio signals are clear and clean, thus improving the audio separation effect.

[0125] (2) Visual data is processed to obtain visually processed data. This visually processed data includes N visual feature signals, each generated by an object. The main process of visual processing is aligning the visual signals with the audio signals to enhance the accuracy of speech separation. Specifically, the visual processing includes the following steps ①-②:

[0126] ① Perform visual recognition processing on the visual data to obtain N candidate visual signals; a candidate visual signal refers to the signal generated synchronously by an object during the process of generating a corresponding audio signal. Here, candidate visual signals can include signals generated by the object's facial expressions, mouth movements (such as lip movements), and other visual states.

[0127] ② Obtain the N second audio signals obtained after the above audio separation, and perform time alignment processing on the N candidate visual signals and the N second audio signals to correct the N candidate second audio signals, resulting in N corrected visual feature signals. Specifically, a second audio signal generated by an object is used to correct the candidate visual signals synchronously generated by that object. For example, the second audio signal generated by any object i among the N objects is represented as the i-th second audio signal, and the candidate visual signals synchronously generated by object i are from the original video frames captured by the camera (including lip movements and facial expressions). Then, after performing time synchronization and alignment processing on the original video frames using the i-th second audio signal, a visual feature signal representing the synchronization of these lip movements, facial expressions, and audio signals can be output. This corrected visual feature signal is then used by the system to further combine with the audio signal for speech separation, alignment, and correction. Based on this, by performing time synchronization and alignment processing on the visual and audio signals, it is possible to ensure that multimodal data is fused within the same time window.

[0128] (3) A multimodal fusion processing method is adopted to fuse audio processing data and visual processing data in a multimodal manner to separate the audio signals from multiple sources and obtain the audio separation result. Specifically, the multimodal fusion processing includes the following steps ①-③:

[0129] ① It can obtain environmental parameters of the business scenario, such as the number of objects N in the business scenario and the noise level in the business scenario.

[0130] ② A weighted fusion strategy corresponding to the business scenario is determined based on environmental parameters. This weighted fusion strategy defines the weight coefficients to be fused, which include the first weight of each second audio signal in the audio processing data and the second weight of each visual feature signal in the visual processing data. These weight parameters are dynamically adjusted based on the environmental parameters of the business scenario, thus adapting to environmental changes and improving the accuracy of weighted fusion of visual and audio data.

[0131] ③ Based on weighting coefficients, the second audio signal of each object in the audio processing data and the corresponding visual feature signal of the object are weighted and fused to obtain the audio separation result. The audio processing data includes N second audio signals, and the visual processing data includes N visual feature signals. Each object is associated with one second audio signal and one visual feature signal. The weighting coefficients defined by the weighted fusion strategy can fuse the second audio signal and visual feature signal of each object to obtain the signal fusion result (i.e., the audio signal) for each object. Taking any object as an example, the second audio signal of the object is represented as Y, and the visual feature signal of the object is represented as S. The first weight of the second audio signal Y is a, and the second weight of the visual feature signal S is b. Then, the signal fusion result (i.e., the audio signal) for this object is represented as: a*Y + b*S; where a + b = 1. The first weight a and the second weight b can be dynamically adjusted based on the environment. If the business scenario is a noisy environment, the first weight a is relatively lower; if the visual feature signal is used to indicate that the user's face is not clearly visible, the second weight b is relatively lower. By employing a weighted fusion strategy, the features of audio signals are combined with those of visual signals; and the weighting coefficients are dynamically adjusted according to the current environment (such as noise level and number of speakers), which can maximize the separation of audio signals and improve the accuracy of audio signal recognition.

[0132] Based on this, in the sound source separation process shown in steps (1)-(3) above, the DPRNN or Transformer model can perform preliminary sound source separation on the audio data. By performing lip-reading recognition and time synchronization and alignment processing on the visual data, the visual signal of each object can be corrected to obtain a visual feature signal that is visually synchronized with the audio signal of the corresponding object. Then, according to the weight coefficients defined by the weighted fusion strategy, the audio signal and visual signal of each object are weighted and fused. The features of an object from audio and video are weighted and fused, and the visual features of lip reading recognition can be introduced on the basis of the defective audio signal, thereby correcting the errors or missing parts in the audio signal and improving the accuracy of audio separation.

[0133] S303: Optimize the audio signal of each object in the audio separation result to obtain the optimized audio signal of each object.

[0134] Here, the audio signal for each object refers to the audio signal after sound source separation in step S302. Since the sound source separation process may result in incomplete or inaccurate separation—for example, the audio signal 1 obtained after separating object 1 may contain parts of the audio signal from other objects 2; or the audio signal 1 obtained after separating object 1 may contain truncated sentences or other incomplete signals—this application requires further optimization processing of the audio signals after separating each object. This optimization processing can include: optimization processing using NLP technology (e.g., semantic analysis, semantic recognition and correction techniques) and optimization processing using speech technology (e.g., automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition techniques). This embodiment mainly involves optimization processing of the audio signals of each object based on NLP technology. The optimization process performed using NLP technology is described below.

[0135] In one possible implementation, any one of the N objects is represented as object j, where j is a positive integer and 1 ≤ j ≤ N. The computer device optimizes the audio signal of each object in the audio separation result to obtain the optimized audio signal of each object, specifically including the following steps (1)-(2):

[0136] (1) Using speech alignment and splicing techniques, the audio signal of object j in the audio separation result is optimized at the first stage to obtain the corrected audio signal of object j. Among them, speech alignment and splicing techniques are used to rearrange and merge the truncated or mixed speech segments to restore complete and coherent sentences. For example, speech alignment and splicing techniques specifically include: speech alignment technology and speech splicing technology; the so-called speech alignment technology can align the audio signal of each separated object (such as object j) to correct the speech overlap or truncation problem caused by multiple objects; the so-called speech splicing technology is used to restore the complete sentence based on the identified speech segments and context information using splicing algorithms.

[0137] Specifically, the execution process of the above-mentioned first-level optimization includes the following steps ①-②:

[0138] ① The audio signal of object j is aligned using speech alignment technology to obtain the aligned audio signal. Here, the speech alignment technology can be Dynamic Time Warping (DTW). As the name suggests, DTW aligns two sequences of different lengths representing the same type of thing in time. Specifically, the DTW algorithm is often used in speech recognition scenarios. Because the same letter, pronounced by different objects, may have different signal lengths, but after recording the sound, its audio signals are very similar, only slightly misaligned in time. In this case, a function (such as the DTW algorithm) is needed to lengthen or shorten the lengths of each audio signal to reduce the error between them, thereby achieving signal alignment.

[0139] ② The aligned audio signal is spliced ​​using speech splicing technology to obtain the corrected audio signal for object j. Specifically, for truncated sentences, the system uses a splicing algorithm to reconstruct the complete sentence based on the identified speech fragments and contextual information.

[0140] (2) Semantic analysis technology is used to perform secondary optimization processing on the corrected audio signal of the current object j to obtain the optimized audio signal of object j. Semantic analysis technology analyzes the text content and context of the audio signal of each object (such as object j) to understand the actual meaning of the statement and infer its rationality. In this embodiment, semantic analysis technology is used to detect and correct confusing and incomplete parts in the speech recognition result (i.e., the separated audio signal of each object) to achieve the purpose of speech optimization.

[0141] Specifically, the execution process of the above-mentioned secondary optimization process includes the following steps ①-②:

[0142] ① A Natural Language Processing (NLP) model is invoked to perform semantic analysis on the corrected audio signal of object j, resulting in analyzed text. The NLP model can be any network structure, such as LSTM (Long Short-Term Memory), CNN (Convolutional Neural Network), or RNN (Recurrent Neural Network). This application does not impose specific limitations on the NLP model structure. Based on this, a deep learning-based NLP model can perform text recognition on audio signals, and then perform contextual analysis on the recognized text to understand the semantic structure and intent of the sentences.

[0143] ② If the analyzed text contains text content of objects other than object j among N objects, then semantic correction processing is performed on the analyzed text to obtain the optimized audio signal of object j. Specifically, when the system detects that content from other speakers is mixed into a sentence, it automatically separates and corrects semantic errors in the sentence through contextual reasoning, thereby improving the accuracy of the audio signal of each separated object.

[0144] In the optimization process shown in steps (1)-(2) above, on the one hand, speech alignment and splicing technology can be used to further correct the audio signals of the separated objects, thereby effectively processing the truncated or mixed audio signals and recovering the complete sentences; on the other hand, NLP and other technologies can be used to perform contextual semantic analysis on the corrected audio signals, so as to correctly separate and correct the sentences according to the context, thereby further performing speech optimization processing and obtaining more accurate optimized audio signals.

[0145] In this embodiment, multimodal data collected in a business scenario is acquired. This multimodal data includes audio data and visual data. The audio data includes multi-source audio signals obtained by collecting audio signals from N objects in the business scenario. The visual data includes visual signals of N objects simultaneously collected during the acquisition of multi-source audio signals, where N is an integer greater than 1. Based on the audio and visual data, a multimodal fusion processing method is used to separate the multi-source audio signals to obtain an audio separation result. This audio separation result includes the audio signal of each object. The audio signal of each object in the audio separation result is then optimized to obtain an optimized audio signal for each object. It can be seen that this application can perform multimodal fusion of video and audio data, thereby enabling the separation of multi-source audio signals by combining visual signals. Because the audio and visual signals are considered from multiple dimensions for source separation, the accuracy of source separation is improved. Furthermore, the optimization processing of the audio signals of each object in the audio separation result further improves the accuracy of the separated audio signals.

[0146] Please see Figure 5 , Figure 5 This is a flowchart illustrating another data processing method provided in an embodiment of this application. This data processing method can be performed by a computer device (such as...). Figure 2 Execute on any of the terminal devices or servers shown.

[0147] like Figure 5 As shown, the data processing method includes, but is not limited to, the following steps S501-S505:

[0148] S501: Obtain multimodal data in business scenarios.

[0149] Multimodal data includes business data collected in at least one modality. Modalities here include, for example, audio modality, visual modality, location modality, etc. Business data in the audio modality is represented as audio data, business data in the visual modality is represented as visual data, and business data in the location modality is represented as location data. The specific data included in this multimodal data can be categorized as follows:

[0150] Case 1: Multimodal data includes audio data and visual data. Audio data comprises multi-source audio signals obtained by acquiring audio signals from N objects in the business scenario; visual data comprises visual signals of N objects simultaneously acquired during the acquisition of the multi-source audio signals; N is an integer greater than 1.

[0151] Scenario 2: Multimodal data includes audio data and location data. Audio data comprises multi-source audio signals obtained by acquiring audio signals from N objects in the business scenario; location data comprises the spatial location signals of the N objects simultaneously acquired during the acquisition of multi-source audio signals.

[0152] Scenario 3: Multimodal data includes audio data, visual data, and location data. Specifically, audio data includes multi-source audio signals obtained by collecting audio signals from N objects in the business scenario; visual data includes visual signals of the N objects simultaneously acquired during the acquisition of the multi-source audio signals; and location data includes spatial location signals of the N objects simultaneously acquired during the acquisition of the multi-source audio signals.

[0153] It should be noted that detailed procedures for collecting multimodal data (such as audio and visual data) can be found in [reference needed]. Figure 3 The detailed process in step S301 of the embodiments will not be repeated here.

[0154] Therefore, depending on the different data included in the multimodal data, the method for source separation of multi-source audio signals based on multimodal data in this application may also vary. Specifically, the source separation methods for multi-source audio signals in the embodiments of this application include the following: Method 1, source separation of multi-source audio signals based on audio data and visual data; Method 2, source separation of multi-source audio signals based on audio data and location data; Method 3, source separation of multi-source audio signals based on audio data, visual data, and location data. The source separation process (multimodal fusion process) involved in each of the above three different methods will be described in detail below.

[0155] S5021: Based on audio data and visual data, a multimodal fusion method is used to separate the audio signals from multiple sources to obtain the audio separation result.

[0156] In one possible implementation, the computer device uses a multimodal fusion processing method to separate the audio signals from multiple sources based on audio and visual data, and the specific process of obtaining the audio separation result includes the following steps (1)-(3):

[0157] (1) Perform signal separation processing on the audio data to obtain processed audio data, which includes audio signals of N objects separated from multi-source audio signals. Here, a source separation model can be used to perform signal separation processing on the audio data. For example, the source separation model can include: a dual-channel speech separation model (such as a dual-path recurrent neural network (DPRNN) model) or an audio conversion model (such as a Transformer model or other models based on the Transformer structure). In this embodiment, taking the DPRNN model and a Transformer-based model as examples, in the source separation task, the input and output of the DPRNN model and the Transformer-based model are respectively: a mixed audio signal (i.e., multi-source audio signals) and separate audio signals of each speaker. The principle process of the above two models will be explained below in conjunction with the input and output.

[0158] I. DPRNN (Dual-Path Recurrent Neural Network) Model:

[0159] 1. Input (multi-source audio signal): refers to the mixed audio signal captured by a microphone array, which is a mixture of audio signals from multiple speakers (objects) (such as in a scenario where multiple people are speaking at the same time).

[0160] 2. Processing steps:

[0161] ① Block processing: DPRNN first divides the long audio signal into multiple short blocks, which helps to handle long-term dependencies. Each short block represents a small segment of mixed audio signal.

[0162] ② Dual-path network: includes time path and spatial path:

[0163] Time path: Used to handle local dependencies in time, that is, to capture short-term audio changes within a short block.

[0164] Spatial path: Used to handle long-term dependencies across blocks, ensuring that global information can be processed.

[0165] ③ Recurrent Network Processing: The RNN processes the information in the short blocks on each path, and understands the context of the time series through the recurrent network in order to better separate speech.

[0166] 3. Output (separated speech signal): This refers to the independent speech segments of each speaker. Through the dual-path processing of DPRNN, the system can output clean audio signals from different speakers. In other words, the system effectively separates multiple sound sources in mixed speech into individual audio streams.

[0167] II. Transformer-based models:

[0168] 1. Input (multi-source audio signal): This is also a mixture of audio signals from multiple speakers (objects) (such as a scenario where multiple people are speaking at the same time). The mixed audio signal captured by the microphone array is usually preprocessed (such as short-time Fourier transform, STFT) to convert it into spectral data.

[0169] 2. Processing steps:

[0170] ① Self-Attention Mechanism: Through the self-attention mechanism, the Transformer model, while processing the audio signal at each time point, simultaneously pays attention to all other time points in the entire audio signal. In this way, the model can understand the long-term dependencies in the speech signal, thereby distinguishing the voices of different speakers.

[0171] ② Parallel processing: Unlike traditional recurrent neural networks (RNNs), Transformer can process the entire audio signal in parallel, rather than processing each time segment step by step, thus enabling more efficient processing of long-term sequences.

[0172] ③ Global modeling: Transformer captures the speech features of each speaker in multi-source audio signals by modeling global information, which helps to perform speech separation more accurately.

[0173] 3. Output (separated speech signals): The model ultimately outputs independent speech streams from multiple speakers. Because the Transformer can globally focus on the entire audio signal, it can effectively separate individual sound sources in complex speech overlap scenarios and output clear, individual speech signals.

[0174] In summary, both DPRNN and Transformer-based source separation methods take mixed speech signals (i.e., multi-source audio signals) captured by microphone arrays as input. These mixed speech signals contain the voices of multiple speakers and may have overlap and noise in real-world scenarios. Therefore, the two models process these signals using different mechanisms (e.g., DPRNN uses a block processing mechanism for signal processing, while the Transformer model uses a self-attention mechanism). Ultimately, both methods output multiple separated audio signals, effectively breaking down the original mixed speech signal into independent, clear audio signals for each speaker. These separated speech signals can be used for subsequent speech recognition, semantic correction, and other processing steps to ensure that the speech recognition system can accurately handle speech input in multi-source environments.

[0175] (2) Visual data is processed to obtain visual processing data. The visual processing data includes N visual feature signals, each of which is generated by an object. The main process of visual processing is to align the visual signals with the audio signals to enhance the accuracy of speech separation.

[0176] (3) A multimodal fusion processing method is adopted to perform multimodal fusion on the audio processing data and the visual processing data to separate the multi-source audio signals and obtain the audio separation result, which includes the audio signal of each object. Specifically, the above audio processing data includes N second audio signals, the visual processing data includes N visual feature signals, and each object is associated with one second audio signal and one visual feature signal, wherein:

[0177] The second audio signal serves as the foundational signal. The system captures audio signals through a microphone array; these signals are the system's basic input and contain the object's vocal data. However, audio signals are often affected by noise, speech overlap, echoes, etc., in complex environments, which may lead to recognition errors or incompleteness. Therefore, it is necessary to combine the audio signal for correction.

[0178] Visual feature signals are used for lip reading. The system captures the speaker's (i.e., the subject's) lip movements and facial expressions using a camera array to form visual feature signals. These signals are then processed by deep learning algorithms such as convolutional neural networks (CNNs) to extract lip-reading features related to the spoken content. These lip-reading features represent information such as the speaker's lip opening and closing, and mouth shape in the video, which can assist in speech recognition, especially when the speech signal quality is poor.

[0179] Therefore, the multimodal fusion process here specifically refers to the weighted fusion of the second audio signal and visual feature signal for each object. The system uses a multimodal fusion strategy to weightedly fuse features from audio and vision. The purpose of this fusion is to introduce visual feature signals from lip reading into defective audio signals, thereby correcting errors or missing parts in the speech and improving the effectiveness of sound source separation.

[0180] S5022: Based on audio data and location data, a multimodal fusion method is used to separate the audio signals from multiple sources and obtain the audio separation result.

[0181] Specifically, if the business scenario belongs to a preset scenario (such as a high-precision or complex environment, for example, an extreme noise environment or a nighttime environment), a positioning device (such as a LiDAR) can be deployed in the business scenario. This positioning device is used to collect the spatial position signals of N objects in the business scenario. Based on this, the three-dimensional spatial position (such as three-dimensional coordinates or direction) signals of each object in the business scenario can be collected by the LiDAR, thereby obtaining position data. In one possible implementation, the computer device uses a multi-modal fusion method to separate the audio signals from multiple sources based on the audio data and the position data, and the specific process of obtaining the audio separation result includes the following steps (1)-(3):

[0182] (1) Perform signal separation processing on the audio data to obtain audio processing data, which includes audio signals of N objects separated from multi-source audio signals.

[0183] (2) The location data is processed to obtain location processing data, which includes the spatial location signal of each object (reflecting the three-dimensional spatial location of the corresponding object in the business scenario). Specifically, LiDAR is used to acquire the three-dimensional spatial location data of each object in the business scenario, and this data is combined with audio signals for sound source localization and separation. In this implementation, the accuracy of LiDAR is higher than that of ordinary cameras, making it more suitable for low-light environments or scenarios with severe visual interference. Therefore, using location data for sound source separation can significantly improve the system's accuracy in locating and separating multi-source audio signals.

[0184] (3) A multimodal fusion processing method is adopted to perform multimodal fusion on audio processing data and location processing data to separate the audio signals from multiple sources and obtain an audio separation result. The audio separation result includes the audio signal of each object. Specifically, the audio processing data contains N second audio signals, and the visual processing data contains N spatial location signals. Each object is associated with one second audio signal and one spatial location signal. Therefore, the multimodal fusion of audio processing data and location processing data is essentially a weighted fusion process of the second audio signal and spatial location signal of each object. The weighted fusion here is mainly a process of assisting in the identification of the second audio signal of the corresponding object based on the spatial location signal of each object.

[0185] Optionally, the computer device performs multimodal fusion of audio processing data and location processing data as follows: ① It can obtain the weighting coefficients to be fused, which include the first weight of each second audio signal and the third weight of each spatial location signal. ② Based on the weighting coefficients, the second audio signal of any object in the audio processing data and the corresponding spatial location signal of the object are weighted and fused to obtain the audio separation result. In this implementation, the location data of each object (such as three-dimensional spatial coordinates and orientation data) collected by LiDAR and the audio data are time-synchronized and multimodal fused. The system can maintain efficient audio signal separation and recognition capabilities in complex environments. Especially under extreme conditions (such as tunnels, nighttime, etc.), the spatial positioning information provided by LiDAR can significantly improve the overall performance of the system.

[0186] Based on this, in high-precision and complex environments (such as extreme noise environments or nighttime environments), due to the limited or inaccurate acquisition of visual data, the embodiments of this application can use LiDAR to spatially locate objects in the business scenario to collect location data. Then, based on the multimodal fusion processing of location data and audio data, the purpose of sound source separation of multi-source audio signals can also be achieved. Since the spatial positioning information provided by LiDAR is more accurate, it can also significantly improve the sound source separation effect of multi-source audio signals.

[0187] S5023: Based on audio data, visual data, and location data, a multimodal fusion method is used to separate audio signals from multiple sources to obtain audio separation results.

[0188] In one possible implementation, the computer device uses a multimodal fusion method to separate the audio signals from multiple sources based on audio data, visual data, and location data, and the specific process of obtaining the audio separation result includes the following steps (1)-(4):

[0189] (1) Perform signal separation processing on the audio data to obtain audio processing data, which includes audio signals of N objects separated from multi-source audio signals.

[0190] (2) Visual data is processed to obtain visual processing data. The visual processing data includes N visual feature signals, each of which is generated by an object. The main process of visual processing is to align the visual signals with the audio signals to enhance the accuracy of speech separation.

[0191] (3) The location data is identified and processed to obtain location processing data, which includes the spatial location signal of each object (used to reflect the three-dimensional spatial location of the corresponding object in the business scenario).

[0192] (4) Multimodal fusion processing is performed on the audio processing data, visual processing data, and location processing data to separate the audio signals from multiple sources and obtain the audio separation result. Specifically, the multimodal fusion processing process is as follows: ① The weighting coefficients to be fused can be obtained, which include the first weight of each second audio signal, the second weight of each visual feature signal, and the third weight of each spatial location signal. ② Based on the weighting coefficients, the second audio signal of any object in the audio processing data, the corresponding visual feature signal of the object, and the corresponding spatial location signal of the object are weighted and fused to obtain the audio signal of the object.

[0193] In this implementation, business data (such as audio data, visual data, and location data) in each modality can be processed separately to obtain the audio signal, visual feature signal, and spatial location signal of each object in each modality. Then, the signals of each object in each modality are weighted and fused to obtain the final separated audio signal of the corresponding object. Because the signals from the audio modality, visual modality, and spatial modality are combined, the accuracy of sound source separation can be improved.

[0194] In another possible implementation, the computer device uses a multimodal fusion method to separate the audio signals from multiple sources based on audio data, visual data, and location data, and the specific process of obtaining the audio separation result includes the following steps (1)-(3):

[0195] (1) Perform multimodal fusion processing on audio data and visual data to separate the audio signals from multiple sources and obtain the first separation result of the audio signals from multiple sources, which includes N first separation signals.

[0196] (2) Perform multimodal fusion processing on audio data and location data to separate the audio signals from multiple sources and obtain a second separation result of the audio signals from multiple sources. The second separation result includes N second separation signals.

[0197] (3) The first separation result and the second separation result are fused to obtain the audio separation result. The fusion process here includes weighted fusion of the first separation signal and the second separation signal of each object to obtain the audio signal of the corresponding object.

[0198] In this implementation, weighted fusion of service data from any two modalities (such as multimodal fusion of audio and visual data, and multimodal fusion of audio and location data) yields the corresponding separation results. These different separation results are then weighted and fused again to obtain the final audio separation result. Similarly, by combining signals from the audio, visual, and spatial modalities, the accuracy of sound source separation can be improved.

[0199] In the process shown in steps S5021-S5023 above, this application embodiment provides three different multimodal fusion methods. In actual application scenarios, the appropriate multimodal fusion method can be flexibly selected according to the specific scenario requirements to perform multimodal fusion processing, thereby separating and processing the audio signals of multiple sound sources to obtain the audio signals of each object.

[0200] S503: Using speech alignment and splicing technology, the audio signal of object j in the audio separation result is optimized in the first stage to obtain the corrected audio signal of object j.

[0201] In one possible implementation, any one of the N objects is represented as object j, where j is a positive integer and 1 ≤ j ≤ N. Specifically, the execution process of the above-mentioned first-level optimization includes the following steps ①-②:

[0202] ① The audio signal of object j is aligned using a speech alignment technique to obtain the aligned audio signal. The speech alignment technique here can be the DTW algorithm, which can perform time alignment on audio signals of different lengths to ensure that the audio signals are synchronized in time.

[0203] ② The aligned audio signal is spliced ​​using speech splicing technology to obtain the corrected audio signal for object j. Specifically, for truncated sentences, the system uses a splicing algorithm to reconstruct the complete sentence based on the recognized audio signal and context information.

[0204] S504: Using semantic analysis technology, perform secondary optimization processing on the modified audio signal of object j to obtain the optimized audio signal of object j.

[0205] In one possible implementation, semantic analysis technology analyzes the text content and context of the audio signal of any object (such as object j) to understand the actual meaning of the statement and infer its rationality. In this embodiment, semantic analysis technology is used to detect and correct confusing and incomplete parts in the speech recognition result (i.e., the audio signal of each separated object) to achieve the purpose of speech optimization. Specifically, the specific execution process of the above-mentioned secondary optimization process includes the following steps ①-②:

[0206] ① A Natural Language Processing (NLP) model is invoked to perform semantic analysis on the corrected audio signal of object j, resulting in analyzed text. The NLP model can be any network structure, such as LSTM (Long Short-Term Memory), CNN (Convolutional Neural Network), or RNN (Recurrent Neural Network). This application does not impose specific limitations on the NLP model structure. Based on this, a deep learning-based NLP model can perform text recognition on audio signals, and then perform contextual analysis on the recognized text to understand the semantic structure and intent of the sentences.

[0207] ② If the analyzed text contains text content of objects other than object j among N objects, then semantic correction processing is performed on the analyzed text to obtain the optimized audio signal of object j. Specifically, when the system detects that content from other speakers is mixed into a sentence, it automatically separates and corrects semantic errors in the sentence through contextual reasoning, thereby improving the accuracy of the audio signal of each separated object.

[0208] S505: Outputs optimized audio signals for each object.

[0209] In one possible implementation, the optimized audio signal includes the task instructions indicated by any object in the business scenario. Outputting the optimized audio signal for each object means outputting the identified task instructions for each object, thereby triggering the business processing system to execute the task instructions indicated by each object within the business scenario. These business scenarios include intelligent driving scenarios or smart home scenarios; for example, in an intelligent driving scenario, the in-vehicle system can execute task instruction 1 (such as playing music) for one object, and in a smart home scenario, the home system can execute task instruction 2 (such as opening curtains, turning on the air conditioner, etc.) for another object.

[0210] Optionally, the computer device can also perform the following operations: ① Output task instructions indicated by optimized audio signals of various objects in the business processing system of the business scenario (such as the in-vehicle system in the intelligent driving scenario, or the home system in the smart home scenario); ② Obtain feedback data of at least one object in response to the task instructions, and obtain environmental data around the business scenario. The feedback data can be data generated by any object performing a feedback operation (such as a confirmation operation, modification operation, or deletion operation) in response to the output task instructions, and the environmental data can be data characterizing the in-vehicle environment, such as noise levels inside the vehicle and the clarity of camera footage; ③ Perform adaptive learning on the business processing system based on the feedback data and environmental data. The adaptively learned business processing system is used to perform audio recognition on various objects in the business scenario. Optionally, the system can also be personalized according to the voice characteristics and usage habits of different users to improve the voice recognition accuracy for specific users. In this implementation, the business processing system continuously collects feedback data and environmental data from objects and continuously optimizes system parameters through adaptive learning algorithms, thereby improving the overall performance of the system.

[0211] In summary, this application provides an intelligent speech alignment and correction system based on multimodal fusion and semantic analysis (i.e., the aforementioned business processing system), mainly applicable to speech recognition and processing in business scenarios with multiple speakers (such as N objects), such as smart cars and smart homes. This business processing system executes the process shown in steps S501-S505 of the embodiments of this application to achieve accurate speech separation, alignment, splicing, and semantic correction of multiple objects in the business scenario. The module structure and complete processing flow of the business processing system are described below with reference to the accompanying drawings:

[0212] Please see Figure 6 , Figure 6 This is a schematic diagram illustrating the structure and processing flow of a business processing system provided in an embodiment of this application. For example... Figure 6 As shown, the business processing system mainly includes the following modules: multimodal data acquisition module, multimodal fusion and signal processing module, speech recognition and semantic analysis module, and user interaction and feedback module. Each module undertakes a specific function and works collaboratively to achieve the overall system's goals. The complete processing flow of the collaborative work of all modules includes the following steps S1-S13:

[0213] (1) Multimodal data acquisition module: responsible for acquiring multimodal data (such as audio data and visual data).

[0214] S1. The microphone array acquires audio signals.

[0215] S2, The camera array acquires visual signals.

[0216] (2) Multimodal fusion and signal processing module: responsible for multimodal fusion based on multimodal data to separate audio signals from multiple sound sources and obtain audio signals of each object.

[0217] S3, sound source separation algorithms (such as DPRNN or Transformer model) process audio signals.

[0218] S4. Beamforming technology enhances signals in a specific direction (i.e., speech enhancement of audio signals).

[0219] S5, the lip-reading recognition algorithm processes visual signals (using CNN to recognize lip movements and facial expressions).

[0220] S6. Time synchronization and alignment processing (aligning the time of audio signals and visual signals).

[0221] S7. Multimodal fusion output of separated speech signals (combining audio signals with visual signals to output clear speech of each object separately).

[0222] (3) Speech recognition and semantic analysis module: responsible for optimizing the audio signals of each separated object, for example, rearranging and merging the truncated or mixed audio signals to restore complete and coherent sentences.

[0223] S8. Audio signal correction algorithm (e.g., using DTW algorithm to align audio signal).

[0224] S9. Statement concatenation restores complete statements (concatenates truncated statement fragments).

[0225] S10, NLP model analyzes text semantics (understanding sentence structure and intent based on context).

[0226] S11, Semantic correction separation and correction of obfuscated content (correcting statement content mixed with other objects).

[0227] (4) User interaction and feedback module: responsible for data interaction with users.

[0228] S12. Output the recognition results (such as the optimized audio signal for each object).

[0229] S13. User feedback collection and system optimization (adaptive learning and optimization of the business processing system based on the collected feedback data to improve overall system performance).

[0230] Based on this, through Figure 6The collaborative work between the modules in the provided business processing system enables high-precision speech recognition and semantic correction in multi-speaker scenarios (such as intelligent driving scenarios, smart home scenarios, etc.), thereby improving the user experience in the above-mentioned business scenarios.

[0231] As described above, the solution proposed in this application is applicable to any intelligent interaction business scenario, such as intelligent driving scenarios and smart home scenarios. The following example, using an intelligent driving scenario as an illustration, demonstrates the processing flow of the solution proposed in this application.

[0232] Please see Figure 7 , Figure 7 This is a flowchart illustrating a data processing scenario provided in an embodiment of this application. For example... Figure 7As shown, the data processing scenario is an intelligent driving scenario, which mainly involves a vehicle 701 and a server 702. The vehicle 701 is used to collect multimodal data of passengers inside the vehicle, and the server 702 is used to process the multimodal data (such as sound source separation, audio recognition, and audio correction). Based on the vehicle 701 and server 702, multi-speaker sound source separation and speech correction effects can be achieved in the intelligent driving scenario. Specifically, the data processing flow executed in this intelligent driving scenario is as follows: ① The vehicle 701 is equipped with a microphone array and a camera array, used to collect audio and visual signals from passengers (objects) inside the vehicle, thereby obtaining multimodal data. This multimodal data includes: audio data collected by the microphone array from N objects (such as object 1 and object 2), and visual data collected by the camera array from N objects (such as object 1 and object 2). ② Vehicle 701 sends the collected multimodal data to server 702. Server 702 is responsible for data processing operations such as sound source separation, audio recognition, and audio correction of the multimodal data. Specifically, the multimodal data includes multi-source audio signals generated by object 1 and object 2. Server 702 can perform sound source separation on the multi-source audio signals based on audio data and visual data to obtain audio signal 1 of object 1 and audio signal 2 of object 2. ③ Server 702 can also use semantic analysis technology to optimize the separated results (audio signal 1 of object 1 and audio signal 2 of object 2). The optimization process here includes, but is not limited to, semantic correction, content completion, and typo correction, to obtain the recognition result of the speech recognition and correction system on the multimodal data. The recognition result includes: optimized audio signal 1 of object 1 and optimized audio signal 2 of object 2. ④ Server 702 returns the recognition results (i.e., optimized audio signal 1 for object 1 and optimized audio signal 2 for object 2) to vehicle 701. Vehicle 701 can output the above recognition results through the in-vehicle display interface (such as S700). For example, the optimized audio signal 1 for object 1 can be output as "set the air conditioning temperature to 25 degrees", and the optimized audio signal 2 for object 2 can be output as "play a popular song". Optionally, the in-vehicle display interface S700 provides the user with a confirmation option or an editing option, allowing the user to edit the output optimized audio signal (such as correction, addition, etc.) in the in-vehicle display interface S700, thereby providing a feedback mechanism for the user. Subsequent user feedback data will be recorded and used to update the system's recognition model, continuously improving the accuracy of the system's speech recognition.

[0233] This application provides an intelligent speech alignment and correction system based on multimodal fusion and semantic analysis, which significantly improves the accuracy and robustness of speech recognition in multi-speaker scenarios. The specific technical effects include the following three aspects: First, by combining audio and visual signals, speech signals from different directions and speakers can be separated more accurately. This fusion method performs particularly well in noisy environments (such as moving cars), enabling the speech recognition system to maintain high-precision separation even under multiple interference sources. Second, the introduction of speech alignment and splicing techniques solves the problem of speech signals being truncated or confused when multiple speakers are speaking simultaneously. Furthermore, through dynamic time warping algorithms and contextual semantic analysis, the system can identify and correct incomplete or confused sentences, ensuring the semantic integrity and accuracy of the output text and reducing the false recognition rate. Third, a user interaction and feedback mechanism is designed. By continuously collecting user feedback data and environmental data, the system can adaptively optimize the recognition model, gradually improving the accuracy of speech recognition for specific users and its personalized processing capabilities.

[0234] The following describes the relevant apparatus of the data processing scheme provided in the embodiments of this application.

[0235] It should be noted that, in the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program with a predetermined function, which works together with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0236] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a data processing apparatus provided in an embodiment of this application. The data processing apparatus 800 can be used to execute corresponding steps in the data processing method provided in the embodiment of this application. Specifically, the data processing apparatus 800 may include:

[0237] The acquisition unit 801 is used to acquire multimodal data collected in the business scenario. The multimodal data includes audio data and visual data. The audio data includes multi-source audio signals obtained by acquiring audio signals from N objects in the business scenario. The visual data includes visual signals of N objects that are simultaneously acquired during the acquisition of the multi-source audio signals. N is an integer greater than 1.

[0238] The processing unit 802 is used to perform source separation on multi-source audio signals based on audio data and visual data using a multimodal fusion processing method, and obtain audio separation results; the audio separation results include the audio signal of each object;

[0239] The processing unit 802 is also used to optimize the audio signal of each object in the audio separation result to obtain the optimized audio signal of each object.

[0240] In one possible implementation, the processing unit 802 performs source separation on the multi-source audio signals based on audio data and visual data using a multimodal fusion processing method to obtain audio separation results, which are then used to perform the following operations:

[0241] The audio data is processed by signal separation to obtain audio processed data, which includes audio signals of N objects separated from multi-source audio signals;

[0242] Visual data is processed visually to obtain visually processed data.

[0243] A multimodal fusion processing method is adopted to perform multimodal fusion on audio processing data and visual processing data in order to separate the audio signals from multiple sources and obtain the audio separation result.

[0244] In one possible implementation, the processing unit 802 performs signal separation processing on the audio data to obtain audio processed data, which is used to perform the following operations:

[0245] A sound source separation model is used to separate audio signals from multiple sound sources in audio data to obtain separated audio data; the sound source separation model includes any one of the following: a dual-channel speech separation model, an audio conversion model, and a multi-channel convolutional neural network model;

[0246] The separated audio data is subjected to audio optimization processing to obtain processed audio data.

[0247] In one possible implementation, the separated audio data includes first audio signals corresponding to N objects, where each first audio signal is represented as the i-th first audio signal, i is a positive integer and 1≤i≤N; the audio processing data includes N second audio signals; the processing unit 802 performs audio optimization processing on the separated audio data to obtain audio processing data, which is used to perform the following operations:

[0248] Beamforming technology is used to perform audio optimization processing on the i-th first audio signal in the separated audio data to obtain the second audio signal after optimization of the i-th first audio signal;

[0249] The optimized N second audio signals are combined into audio processing data;

[0250] The audio optimization processing includes one or more of the following: speech enhancement and noise suppression.

[0251] In one possible implementation, the visual processing data includes N visual feature signals, each visual feature signal being generated by an object; the processing unit 802 performs visual processing on the visual data to obtain visual processing data, which is used to perform the following operations:

[0252] Visual data is processed for visual recognition to obtain N candidate visual signals; a candidate visual signal refers to an object that is synchronously generated by the object in the process of generating the corresponding audio signal.

[0253] N second audio signals are acquired, and time alignment processing is performed on the N candidate visual signals and the N second audio signals to correct the N candidate second audio signals, thereby obtaining the corrected N visual feature signals;

[0254] Among them, a second audio signal generated by an object is used to correct the candidate visual signal generated synchronously by the object.

[0255] In one possible implementation, the processing unit 802 employs a multimodal fusion processing method to perform multimodal fusion on the audio processing data and the visual processing data to separate the audio signals from multiple sources, obtaining an audio separation result for performing the following operations:

[0256] Obtain environmental parameters for the business scenario;

[0257] The weighted fusion strategy corresponding to the business scenario is determined based on environmental parameters. The weighted fusion strategy includes the weight coefficients to be fused, which include the first weight of each second audio signal in the audio processing data and the second weight of each visual feature signal in the visual processing data.

[0258] Based on weighting coefficients, the second audio signal of any object in the audio processing data and the visual feature signal of the object are weighted and fused to obtain the audio separation result.

[0259] In one possible implementation, the processing unit 802 performs source separation on the multi-source audio signals based on audio data and visual data using a multimodal fusion processing method to obtain audio separation results, which are then used to perform the following operations:

[0260] Acquire location data collected in the business scenario. Location data refers to the data obtained by using LiDAR technology to collect the spatial location of N objects in the business scenario.

[0261] Based on location data, frequency data, and visual data, a multimodal fusion processing method is used to separate the sound sources of multi-source audio signals, and the audio separation results are obtained.

[0262] In one possible implementation, the processing unit 802 is further configured to perform the following operations:

[0263] Multimodal fusion processing is performed on audio and visual data to separate the audio signals from multiple sources, resulting in the first separation result of the audio signals from multiple sources.

[0264] Multimodal fusion processing is performed on audio data and location data to separate the audio signals from multiple sources, resulting in a second separation result for the audio signals from multiple sources.

[0265] The first separation result and the second separation result are fused to obtain the audio separation result.

[0266] In one possible implementation, any one of the N objects is represented as object j, where j is a positive integer and 1 ≤ j ≤ N; the processing unit 802 optimizes the audio signal of each object in the audio separation result to obtain the optimized audio signal of each object, which is used to perform the following operations:

[0267] Using speech alignment and splicing technology, the audio signal of object j in the audio separation result is optimized in the first stage to obtain the corrected audio signal of object j;

[0268] Semantic analysis techniques are used to perform secondary optimization processing on the modified audio signal of object j, resulting in the optimized audio signal of object j.

[0269] In one possible implementation, the processing unit 802 employs speech alignment and splicing technology to perform a first-level optimization process on the audio signal of object j in the audio separation result, obtaining a corrected audio signal of object j, which is used to perform the following operations:

[0270] The audio signal of object j is aligned using voice alignment technology to obtain the aligned audio signal;

[0271] The aligned audio signals are spliced ​​using voice splicing technology to obtain the corrected audio signal for object j.

[0272] In one possible implementation, processing unit 802 employs semantic analysis technology to perform secondary optimization processing on the optimized audio signal of object j, obtaining the optimized audio signal of object j, which is then used to perform the following operations:

[0273] The natural language processing model is invoked to perform semantic analysis on the modified audio signal of object j, resulting in the analyzed text.

[0274] If the analyzed text contains the text content of other objects besides object j among N objects, then the analyzed text is semantically corrected to obtain the optimized audio signal of object j.

[0275] In one possible implementation, optimizing the audio signal includes task instructions indicated by any object in the business scenario; the processing unit 802 is also configured to perform the following operations:

[0276] The business processing system outputs the task instructions indicated by the optimized audio signals of each object in the business scenario; where the business scenario includes intelligent driving scenario or smart home scenario;

[0277] Obtain feedback data from at least one object in response to task instructions, and obtain environmental data surrounding the business scenario;

[0278] Based on feedback data and environmental data, the business processing system undergoes adaptive learning; the adaptively learned business processing system is then used to perform audio recognition on various objects within the business scenario.

[0279] In this embodiment, multimodal fusion of video and audio data is possible, thereby enabling sound source separation of multi-source audio signals by combining visual signals. Since sound source separation considers both audio and visual signals from multiple dimensions, the accuracy of sound source separation is improved. Furthermore, the audio signals of each object in the audio separation result are optimized, further improving the accuracy of the separated audio signals.

[0280] Please see Figure 9 , Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device 900 is used to execute the steps performed by the computer device in the aforementioned method embodiments. The computer device 900 may include one or more devices (such as servers, nodes, terminals, etc.) or internal components (such as chips, software modules, or hardware modules). The computer device may include at least one processor 901 and a communication interface 902. Further optionally, the computer device may also include at least one memory 903 and a bus 904. Additionally, the processor 901, communication interface 902, and memory 903 are connected via the bus 904.

[0281] (1) The processor 901 is a module that performs arithmetic and / or logical operations. Specifically, it may be one or a combination of processing modules such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor unit (MPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), a coprocessor (to assist the central processing unit in completing corresponding processing and applications), and a micro controller unit (MCU).

[0282] (2) The communication interface 902 can be used to provide information input or output to at least one processor 901. And / or, the communication interface 902 can be used to receive data sent externally and / or send data externally, and can be a wired link interface including such as an Ethernet cable, or a wireless link interface (Wi-Fi, Bluetooth, general wireless transmission, vehicle short-range communication technology, and other short-range wireless communication technologies, etc.). The communication interface 902 can serve as a network interface.

[0283] (3) The memory 903 is used to provide storage space, in which data such as the operating system and computer programs (including program instructions) can be stored. The memory 903 can be one or a combination of random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), etc.

[0284] In practice, the processor 901 executes the program instructions stored in the memory 903 to cause the computer device to perform the following operations:

[0285] Acquire multimodal data collected in the business scenario. The multimodal data includes audio data and visual data. The audio data includes multi-source audio signals obtained by collecting audio signals from N objects in the business scenario. The visual data includes visual signals of N objects that are simultaneously collected during the process of collecting the multi-source audio signals. N is an integer greater than 1.

[0286] Based on audio and visual data, a multimodal fusion processing method is used to separate the audio signals from multiple sources, resulting in audio separation results; the audio separation results include the audio signal of each object.

[0287] The audio signal of each object in the audio separation result is optimized separately to obtain the optimized audio signal of each object.

[0288] In one possible implementation, processor 901 uses a multimodal fusion processing method to separate the multi-source audio signals based on audio and visual data, obtaining audio separation results, which are then used to perform the following operations:

[0289] The audio data is processed by signal separation to obtain audio processed data, which includes audio signals of N objects separated from multi-source audio signals;

[0290] Visual data is processed visually to obtain visually processed data.

[0291] A multimodal fusion processing method is adopted to perform multimodal fusion on audio processing data and visual processing data in order to separate the audio signals from multiple sources and obtain the audio separation result.

[0292] In one possible implementation, processor 901 performs signal separation processing on the audio data to obtain audio processed data, which is used to perform the following operations:

[0293] A sound source separation model is used to separate audio signals from multiple sound sources in audio data to obtain separated audio data; the sound source separation model includes any one of the following: a dual-channel speech separation model, an audio conversion model, and a multi-channel convolutional neural network model;

[0294] The separated audio data is subjected to audio optimization processing to obtain processed audio data.

[0295] In one possible implementation, the separated audio data includes first audio signals corresponding to N objects, where any one of the first audio signals is represented as the i-th first audio signal, i being a positive integer and 1 ≤ i ≤ N; the audio processing data includes N second audio signals; the processor 901 performs audio optimization processing on the separated audio data to obtain audio processing data, which is used to perform the following operations:

[0296] Beamforming technology is used to perform audio optimization processing on the i-th first audio signal in the separated audio data to obtain the second audio signal after optimization of the i-th first audio signal;

[0297] The optimized N second audio signals are combined into audio processing data;

[0298] The audio optimization processing includes one or more of the following: speech enhancement and noise suppression.

[0299] In one possible implementation, the visual processing data includes N visual feature signals, each visual feature signal being generated by an object; the processor 901 performs visual processing on the visual data to obtain visual processing data, which is used to perform the following operations:

[0300] Visual data is processed for visual recognition to obtain N candidate visual signals; a candidate visual signal refers to an object that is synchronously generated by the object in the process of generating the corresponding audio signal.

[0301] N second audio signals are acquired, and time alignment processing is performed on the N candidate visual signals and the N second audio signals to correct the N candidate second audio signals, thereby obtaining the corrected N visual feature signals;

[0302] Among them, a second audio signal generated by an object is used to correct the candidate visual signal generated synchronously by the object.

[0303] In one possible implementation, processor 901 employs a multimodal fusion processing method to perform multimodal fusion on audio processing data and visual processing data to separate the audio signals from multiple sources, obtaining audio separation results for performing the following operations:

[0304] Obtain environmental parameters for the business scenario;

[0305] The weighted fusion strategy corresponding to the business scenario is determined based on environmental parameters. The weighted fusion strategy includes the weight coefficients to be fused, which include the first weight of each second audio signal in the audio processing data and the second weight of each visual feature signal in the visual processing data.

[0306] Based on weighting coefficients, the second audio signal of any object in the audio processing data and the visual feature signal of the object are weighted and fused to obtain the audio separation result.

[0307] In one possible implementation, processor 901 uses a multimodal fusion processing method to separate the multi-source audio signals based on audio and visual data, obtaining audio separation results, which are then used to perform the following operations:

[0308] Acquire location data collected in the business scenario. Location data refers to the data obtained by using LiDAR technology to collect the spatial location of N objects in the business scenario.

[0309] Based on location data, frequency data, and visual data, a multimodal fusion processing method is used to separate the sound sources of multi-source audio signals, and the audio separation results are obtained.

[0310] In one possible implementation, processor 901 is also used to perform the following operations:

[0311] Multimodal fusion processing is performed on audio and visual data to separate the audio signals from multiple sources, resulting in the first separation result of the audio signals from multiple sources.

[0312] Multimodal fusion processing is performed on audio data and location data to separate the audio signals from multiple sources, resulting in a second separation result for the audio signals from multiple sources.

[0313] The first separation result and the second separation result are fused to obtain the audio separation result.

[0314] In one possible implementation, any one of the N objects is represented as object j, where j is a positive integer and 1 ≤ j ≤ N; the processor 901 optimizes the audio signal of each object in the audio separation result to obtain the optimized audio signal of each object, which is used to perform the following operations:

[0315] Using speech alignment and splicing technology, the audio signal of object j in the audio separation result is optimized in the first stage to obtain the corrected audio signal of object j;

[0316] Semantic analysis techniques are used to perform secondary optimization processing on the modified audio signal of object j, resulting in the optimized audio signal of object j.

[0317] In one possible implementation, processor 901 employs speech alignment and splicing technology to perform a first-level optimization process on the audio signal of object j in the audio separation result, obtaining a corrected audio signal of object j, which is used to perform the following operations:

[0318] The audio signal of object j is aligned using voice alignment technology to obtain the aligned audio signal;

[0319] The aligned audio signals are spliced ​​using voice splicing technology to obtain the corrected audio signal for object j.

[0320] In one possible implementation, processor 901 employs semantic analysis techniques to perform secondary optimization processing on the optimized audio signal of object j, obtaining the optimized audio signal of object j, which is then used to perform the following operations:

[0321] The natural language processing model is invoked to perform semantic analysis on the modified audio signal of object j, resulting in the analyzed text.

[0322] If the analyzed text contains the text content of other objects besides object j among N objects, then the analyzed text is semantically corrected to obtain the optimized audio signal of object j.

[0323] In one possible implementation, optimizing the audio signal includes task instructions indicated by any object in the business scenario; the processor 901 is also used to perform the following operations:

[0324] The business processing system outputs the task instructions indicated by the optimized audio signals of each object in the business scenario; where the business scenario includes intelligent driving scenario or smart home scenario;

[0325] Obtain feedback data from at least one object in response to task instructions, and obtain environmental data surrounding the business scenario;

[0326] Based on feedback data and environmental data, the business processing system undergoes adaptive learning; the adaptively learned business processing system is then used to perform audio recognition on various objects within the business scenario.

[0327] In this embodiment, multimodal fusion of video and audio data is possible, thereby enabling sound source separation of multi-source audio signals by combining visual signals. Since sound source separation considers both audio and visual signals from multiple dimensions, the accuracy of sound source separation is improved. Furthermore, the audio signals of each object in the audio separation result are optimized, further improving the accuracy of the separated audio signals.

[0328] According to one aspect of this application, embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions. When a processor executes the program instructions, it can perform the methods described in the corresponding embodiments above; therefore, further details will not be repeated here. For technical details not disclosed in the embodiments of the computer-readable storage medium involved in this application, please refer to the description of the method embodiments of this application. As an example, the program instructions can be deployed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network.

[0329] According to one aspect of this application, embodiments of this application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods described in the foregoing embodiments. For technical details not disclosed in the embodiments of the computer program product involved in this application, please refer to the description of the method embodiments of this application, which will not be repeated here.

[0330] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product, which includes one or more computer programs. When the computer program is loaded and executed on a computer device, it generates, in whole or in part, the processes or functions described in the embodiments of this application; the computer device can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer program can be stored in or transmitted through a computer-readable storage medium; the computer program can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible to the computer device or a data processing device such as a server or data center that integrates one or more available media; wherein, the available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0331] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. A data processing method, characterized in that, include: Acquire multimodal data collected in the business scenario, the multimodal data including audio data and visual data; the audio data includes multi-source audio signals obtained by collecting audio signals from N objects in the business scenario, and the visual data includes visual signals of the N objects simultaneously collected during the process of collecting the multi-source audio signals; N is an integer greater than 1; Based on the audio data and the visual data, a multimodal fusion processing method is used to separate the multi-source audio signals to obtain an audio separation result; the audio separation result includes the audio signal of each of the objects. The audio signal of each object in the audio separation result is optimized to obtain the optimized audio signal of each object.

2. The method as described in claim 1, characterized in that, The method of using multimodal fusion processing to separate the multi-source audio signals based on the audio data and the visual data to obtain audio separation results includes: The audio data is subjected to signal separation processing to obtain audio processing data, which includes audio signals of N objects separated from the multi-source audio signals; The visual data is subjected to visual processing to obtain visually processed data; A multimodal fusion processing method is adopted to perform multimodal fusion on the audio processing data and the visual processing data in order to separate the multi-source audio signals and obtain the audio separation result.

3. The method as described in claim 2, characterized in that, The step of performing signal separation processing on the audio data to obtain processed audio data includes: A sound source separation model is used to perform signal separation processing on the multi-source audio signals in the audio data to obtain separated audio data; wherein, the sound source separation model includes any one of the following: a dual-channel speech separation model, an audio conversion model, and a multi-channel convolutional neural network model; The separated audio data is subjected to audio optimization processing to obtain processed audio data.

4. The method as described in claim 3, characterized in that, The separated audio data includes first audio signals corresponding to N objects, where any one of the first audio signals is represented as the i-th first audio signal, i is a positive integer and 1≤i≤N; the audio processing data includes N second audio signals. The audio optimization processing of the separated audio data to obtain processed audio data includes: Beamforming technology is used to perform audio optimization processing on the i-th first audio signal in the separated audio data to obtain the optimized second audio signal of the i-th first audio signal; The optimized N second audio signals are combined into audio processing data; The audio optimization processing includes one or more of the following: speech enhancement and noise suppression.

5. The method as described in claim 2, characterized in that, The visual processing data includes N visual feature signals, where each visual feature signal is generated by an object; the visual processing of the visual data to obtain visual processing data includes: The visual data is subjected to visual recognition processing to obtain N candidate visual signals; each candidate visual signal refers to an object synchronously generated by the object during the process of generating a corresponding audio signal. N second audio signals are acquired, and time alignment processing is performed on the N candidate visual signals and the N second audio signals to correct the N candidate second audio signals, thereby obtaining N corrected visual feature signals. A second audio signal generated by an object is used to correct the candidate visual signal generated synchronously by the object.

6. The method as described in claim 5, characterized in that, The method employs multimodal fusion processing to fuse the audio processing data and the visual processing data in a multimodal manner, thereby separating the multi-source audio signals and obtaining audio separation results, including: Obtain the environmental parameters of the business scenario; Based on the environmental parameters, a weighted fusion strategy corresponding to the business scenario is determined. The weighted fusion strategy includes weight coefficients to be fused. The weight coefficients include a first weight for each second audio signal in the audio processing data and a second weight for each visual feature signal in the visual processing data. Based on the weighting coefficients, the second audio signal of any object in the audio processing data and the visual feature signal of the object are weighted and fused to obtain the audio separation result.

7. The method according to any one of claims 1-6, characterized in that, The method of using multimodal fusion processing to separate the multi-source audio signals based on the audio data and the visual data to obtain audio separation results includes: Acquire location data collected in the business scenario, wherein the location data refers to the data obtained after using LiDAR technology to collect the spatial location of the N objects in the business scenario; Based on the location data, the audio data, and the visual data, a multimodal fusion processing method is used to separate the multi-source audio signals to obtain the audio separation result.

8. The method as described in claim 7, characterized in that, The method further includes: The audio data and the visual data are subjected to multimodal fusion processing to separate the multi-source audio signals and obtain a first separation result of the multi-source audio signals; The audio data and the location data are subjected to multimodal fusion processing to separate the multi-source audio signals and obtain a second separation result of the multi-source audio signals; The first separation result and the second separation result are fused together to obtain the audio separation result.

9. The method as described in claim 1, characterized in that, Any one of the N objects is represented as object j, where j is a positive integer and 1 ≤ j ≤ N; the optimization processing of the audio signal of each object in the audio separation result to obtain the optimized audio signal of each object includes: Using speech alignment and splicing technology, the audio signal of object j in the audio separation result is subjected to first-level optimization processing to obtain the corrected audio signal of object j; Semantic analysis technology is used to perform secondary optimization processing on the modified audio signal of object j to obtain the optimized audio signal of object j.

10. The method as described in claim 9, characterized in that, The process employs speech alignment and splicing technology to perform a first-level optimization on the audio signal of object j in the audio separation result, obtaining the corrected audio signal of object j, including: The audio signal of object j is aligned using voice alignment technology to obtain an aligned audio signal; The aligned audio signal is spliced ​​using voice splicing technology to obtain the corrected audio signal for object j.

11. The method as described in claim 9, characterized in that, The step of employing semantic analysis technology to perform secondary optimization processing on the optimized audio signal of object j to obtain the optimized audio signal of object j includes: The modified audio signal of object j is semantically analyzed by calling a natural language processing model to obtain the analyzed text; If the analyzed text contains text content of other objects among the N objects besides object j, then the analyzed text is semantically corrected to obtain the optimized audio signal of object j.

12. The method as described in claim 1, characterized in that, The optimized audio signal includes a task instruction indicated by any object in the business scenario; the method further includes: The business processing system of the aforementioned business scenario outputs task instructions indicated by optimized audio signals of various objects; wherein, the business scenario includes intelligent driving scenario or smart home scenario; Obtain feedback data from at least one object in response to the task instruction, and obtain environmental data surrounding the business scenario; Based on the feedback data and the environmental data, the business processing system undergoes adaptive learning; wherein, the adaptively learned business processing system is used to perform audio recognition on various objects in the business scenario.

13. A data processing apparatus, characterized in that, include: The acquisition unit is used to acquire multimodal data collected in the business scenario. The multimodal data includes audio data and visual data. The audio data includes multi-source audio signals obtained by acquiring audio signals from N objects in the business scenario. The visual data includes visual signals of the N objects that are synchronously acquired during the acquisition of the multi-source audio signals. N is an integer greater than 1. The processing unit is configured to perform source separation on the multi-source audio signals based on the audio data and the visual data using a multimodal fusion processing method to obtain an audio separation result; the audio separation result includes the audio signal of each of the objects; The processing unit is further configured to optimize the audio signal of each object in the audio separation result to obtain an optimized audio signal for each object.

14. A computer device, characterized in that, include: Memory and processor; The memory stores one or more computer programs; A processor for loading one or more computer programs to implement the data processing method as described in any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, performs the data processing method as described in any one of claims 1-12.

16. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, performs the data processing method as described in any one of claims 1-12.