Audio signal processing method and electronic equipment
By establishing the mapping relationship between the target object's voice signal and the audio channel, combining position information and voiceprint characteristics, dynamically updating the mapping relationship, the problem of lowering audio signal recognition accuracy in multi-person speech scenes is solved, and stable audio signal processing is achieved.
Patent Information
- Application Number
- CN202510572936.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-18
AI Technical Summary
In multi-person speaking scenarios, the prior art cannot maintain stable tracking of each sound source, resulting in a decrease in the audio signal recognition accuracy.
By establishing the mapping relationship between the target object's voice signal and the audio channel, using position information and voiceprint characteristics, the mapping relationship is dynamically updated to adapt to the location changes and identity recognition of the target object, and combining the historical version mapping relationship backtracking mechanism and overlapping audio signal processing, we ensure the accurate allocation and identification of the audio signal.
It improves the accuracy and stability of the audio signal recognition process, adapts to complex and changeable environments, and ensures the continuity and efficiency of speech recognition.
Smart Images

Figure CN120340518A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of signal processing, and in particular, to an audio signal processing method and an electronic device. Background Art
[0002] In the field of audio signal processing, the scenario of multiple people speaking has become a common situation in practical applications. During long-term recording, the prior art cannot maintain stable tracking of each sound source, resulting in a reduction in the accuracy of audio signal recognition. Summary of the Invention
[0003] In view of this, the present disclosure provides an audio signal processing method and an electronic device.
[0004] According to a first aspect of the present disclosure, there is provided an audio signal processing method, including:
[0005] Obtain a first audio signal, where the first audio signal carries at least voice signals of two different target objects;
[0006] Based on the voice signals of each of the target objects in the first audio signal, determine a mapping relationship, where the mapping relationship represents the correspondence between the voice signals of each of the target objects and audio channels;
[0007] Based on the mapping relationship, identify the second audio signal to obtain voice signals corresponding to at least one target object, where the second audio signal is determined after the first audio signal is obtained.
[0008] According to an embodiment of the present disclosure, before determining the mapping relationship based on the voice signals of each of the target objects in the first audio signal, the method further includes:
[0009] Obtain the position information of each of the target objects;
[0010] Determine the audio channels corresponding to the position information matching each of the target objects.
[0011] According to an embodiment of the present disclosure, the identifying the second audio signal based on the mapping relationship includes:
[0012] Obtain the second audio signal, and process the second audio signal to obtain at least one target voice signal;
[0013] For any one of the target voice signals, based on the target voice signal, determine whether the position information of each of the target objects changes;
[0014] If it is determined that the position information of the target object changes, update the mapping relationship, and based on the updated mapping relationship, identify the second audio signal;
[0015] If it is determined that the location information of the target object has not changed, the second audio signal is identified based on the mapping relationship.
[0016] According to an embodiment of the present disclosure, the method further includes:
[0017] If it is determined that the location information of the target object has changed, the registered voiceprint feature of the target object is obtained, and the mapping relationship is updated based on the voiceprint feature of the target object, so that the mapping relationship includes the correspondence between the voiceprint feature of the target object and the audio channel.
[0018] According to an embodiment of the present disclosure, determining whether the location information of each target pair has changed based on the target voice signal includes:
[0019] Obtain the target feature in the target voice signal;
[0020] If it is determined that the target feature matches the feature indicated by the mapping relationship, it is determined that the location information has not changed;
[0021] If it is determined that the target feature does not match the feature indicated by the mapping relationship, it is determined that the location information has changed;
[0022] Wherein, the target feature includes at least one of a voice feature and a semantic feature.
[0023] According to an embodiment of the present disclosure, the target object is the second object, and the voiceprint feature of the second object is not registered. After obtaining the target feature in the target voice signal, it further includes:
[0024] In the case where the difference degree between the target feature of the second object and the target feature recorded in the mapping relationship is greater than a preset threshold, the target feature recorded for the second object in the mapping relationship is updated.
[0025] According to an embodiment of the present disclosure, the target object is the first object, and the voiceprint feature of the first object is registered. The method further includes:
[0026] In the case where it is determined that there is an overlapping audio signal in the second audio signal, the time stamp of the overlapping audio signal is obtained, and the overlapping audio signal is an audio signal in which the voice signals of at least two objects overlap;
[0027] Extract the voice signal of the first object in the overlapping audio signal based on the voiceprint feature of the first object;
[0028] Based on the time stamp of the overlapping audio signal, replace the overlapping audio signal with the voice signal of the first object.
[0029] According to an embodiment of the present disclosure, the situations of the change in the position information of each of the target objects include at least one of the following:
[0030] The relative positions between the target objects change;
[0031] One or more other target objects are added;
[0032] One or more of the target objects are replaced by other target objects.
[0033] According to an embodiment of the present disclosure, the method further includes:
[0034] In the case where the recognition of the second audio signal fails, obtain a mapping relationship of a historical version, where the mapping relationship of the historical version is the mapping relationship of a preset number of rounds before the current mapping relationship;
[0035] Based on the mapping relationship of the historical version, recognize the second audio signal.
[0036] A second aspect of the present disclosure provides an electronic device, including:
[0037] An audio device for collecting a first audio signal, where the first audio signal carries at least voice signals of two different target objects;
[0038] A processor for determining a mapping relationship based on the voice signals of the target objects in the first audio signal, where the mapping relationship represents the corresponding relationship between the voice signals of the target objects and audio channels; and recognizing a second audio signal based on the mapping relationship to obtain voice signals corresponding to at least one target object, where the second audio signal is determined after the first audio signal.
[0039] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0041] Figure 1 Schematically shows a flowchart of an audio signal processing method provided by an embodiment of the present disclosure;
[0042] Figures 2A to 2D Schematically shows schematic diagrams of various scenarios of changes in position information provided by an embodiment of the present disclosure;
[0043] Figure 3A block diagram of an electronic device provided by an embodiment of the present disclosure is schematically shown. Detailed implementation manners
[0044] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.
[0045] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0046] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0047] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0048] In the embodiments of the present disclosure, in terms of the collection, update, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the involved data (for example, including but not limited to user personal information), they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. In particular, necessary measures are taken for user personal information to prevent illegal access to user personal information data and to safeguard user personal information security, network security, and national security.
[0049] Embodiments of the present disclosure provide an audio signal processing method and an electronic device. Among them, the audio signal processing method includes:
[0050] Obtain a first audio signal, where the first audio signal carries at least two different target objects' speech signals;
[0051] Determine a mapping relationship based on the voice signals of the target objects in the first audio signal, where the mapping relationship represents the correspondence between the voice signals of the target objects and the audio channels;
[0052] Based on the mapping relationship, identify the voice signals corresponding to at least one target object from the second audio signal, where the second audio signal is obtained after the first audio signal.
[0053] By adopting the embodiments of the present disclosure, a mapping relationship between the voiceprint features of the voice signals of the target objects and the audio channels is established through the first audio signal, and the voice signals of the target objects in the subsequent second audio signal are accurately assigned to the matching audio channels by using this mapping relationship, so as to achieve continuous and stable audio signal recognition, thereby improving the accuracy of the audio signals during the recognition process.
[0054] The electronic device in the embodiments of the present disclosure refers to a multi-channel device with the ability to collect, process, and analyze audio signals. Such electronic devices include multi-microphone array devices such as intelligent conference systems, remote collaboration platforms, and smart speakers; voice interaction terminals such as voice translators, real-time transcription devices, and intelligent assistants; and scenario-based voice solutions such as far-field pickups with spatial audio processing capabilities, voiceprint recognition access control systems, and multi-person online education platforms. In addition, professional audio processing components such as the sound enhancement unit of medical stethoscope digitization devices, multi-dialect audio recognition systems for judicial court trial records, and industrial environmental noise monitoring and analysis instruments; vertical field audio processing units such as remote bank voice verification modules in the financial field, acoustic anomaly detectors in the security industry, and multi-track audio reconstruction engines in media production also belong to the category of electronic devices described in the embodiments of the present disclosure. At the same time, special scenario audio devices such as smart home multi-person voice control centers in the Internet of Things environment, in-vehicle multi-passenger voice interaction systems, and multi-party dialogue recognition devices in aviation communication, as well as hierarchical architecture audio systems such as collaborative processing nodes of distributed microphone networks, cloud-based voiceprint recognition servers, and edge computing audio processing units, are all within the technical adaptation scope of the embodiments of the present disclosure.
[0055] It should be noted that the embodiments of the present disclosure do not limit the specific type of the electronic device. The above examples are only illustrative descriptions, and the technical solutions of the present disclosure can be applied to any device with the ability to collect, process, and analyze audio signals.
[0056] Next, an audio signal processing method provided in the embodiments of the present disclosure will be introduced.
[0057] Figure 1 The flowchart of the audio signal processing method is shown.
[0058] As Figure 1As shown, the audio signal processing method may further include operations S101 to S103.
[0059] Operation S101: Obtain a first audio signal, where the first audio signal carries at least the voice signals of two different target objects;
[0060] Operation S102: Based on the voice signals of the respective target objects in the first audio signal, determine a mapping relationship, where the mapping relationship represents the correspondence between the voice signals of the respective target objects and the audio channels;
[0061] Operation S103: Based on the mapping relationship, identify the second audio signal to obtain the voice signals corresponding to at least one target object, where the second audio signal is obtained after the first audio signal.
[0062] In operation S101, an audio signal is an electrical signal representing the variation of sound in the time domain, and usually, an audio device such as a microphone converts sound waves into an electrical digital signal form that can be processed and analyzed. In the embodiments of the present disclosure, the first audio signal specifically refers to the original audio data containing the voice content of multiple target objects collected by the electronic device at the initial stage.
[0063] Specifically, the first audio signal usually carries the voice signals of two different target objects. For example, in a meeting scenario, it may contain the speech content of multiple participants; in a home environment, it may contain the conversations of multiple family members; or in a teaching scenario, it may contain the interactive voices of teachers and students. The first audio signal may be a situation where the target objects speak alternately, or a situation where multiple people speak simultaneously to form overlapping voices.
[0064] Among them, a target object refers to the sound source body that generates the voice signal, and in the embodiments of the present disclosure, it can be understood as an individual speaker whose voice signal needs to be recognized. A target object usually refers to meeting participants, teachers and students in the classroom, family members, or any individual who needs to be separately recognized in a multi-person voice environment.
[0065] Similarly, the voice signal of a target object refers to the electrical signal corresponding to the sound wave emitted by a specific target object and captured by the audio device. That is, the voice content of a specific object that needs to be separately recognized and extracted from the mixed audio.
[0066] In practical applications, the method for obtaining the first audio signal is to collect audio data in the acoustic environment through an audio device, thereby forming an original signal containing the voice content of multiple target objects.
[0067] In some embodiments, the audio device includes but is not limited to microphones, microphone arrays, omnidirectional microphones, directional microphones, pickups, and other acoustic sensors capable of converting sound waves into electrical signals. In addition, the number of audio devices can be one or more. When multiple audio devices are used, they can form an array structure to improve the accuracy of sound collection and the ability of spatial perception.
[0068] Regarding the installation position of the audio device, in a feasible implementation, the audio device can be integrated into an electronic device, such as a microphone or microphone array built into an electronic device such as a smartphone, tablet computer, smart speaker, or conference system terminal. The above integration method is convenient for users to carry and use without additional equipment.
[0069] In another feasible implementation, the audio device can also be independent of the electronic device and connected to the electronic device as an external device, such as an external microphone, USB pickup, or professional audio capture card, so as to provide higher-quality audio capture capabilities;
[0070] In yet another feasible implementation, the built-in and external audio devices can be used simultaneously to form a larger-scale microphone array to cover a wider spatial range and obtain richer spatial acoustic information, which is particularly suitable for large conference rooms or professional scenarios that require high-precision audio capture.
[0071] In operation S102, the mapping relationship refers to the corresponding connection established between the voice signals of each target object and the audio channels. In the embodiments of the present disclosure, it can be understood as a data structure that associates the relevant features of the voice signals of a specific target object with the audio channels.
[0072] Among them, the audio channel refers to an independent data path for transmitting and processing the audio signals of a specific sound source. In the embodiments of the present disclosure, it can be understood as an independent audio track or processing path corresponding to the voice signal of a specific target object, which is used for separately storing, processing, and analyzing the voice signal of this target object.
[0073] It should be noted that the audio channels in the embodiments of the present disclosure are different from the physical channels in the traditional stereo channels or multi-channel systems. Instead, they specifically refer to the logical audio tracks separately allocated for each identified target object during the audio recognition process. Each audio channel can be regarded as a virtual voice signal container, which not only stores the voice content of a specific target object but may also be associated with the location information, voiceprint characteristics, and other audio characteristic parameters of this target object. The above voice channel division method enables the system to independently operate on the voice signals of each target object, thereby achieving more accurate voice recognition in an overlapping voice scenario.
[0074] In order to establish an effective mapping relationship, in a feasible implementation, a mapping relationship between a voice signal and an audio channel can be established based on the characteristic parameters of the voice signal; wherein, the characteristic parameters include but are not limited to voiceprint characteristics, voice characteristics (such as pitch, timbre, speech rate), voice energy distribution characteristics, and time-frequency domain characteristics, etc.
[0075] In another feasible implementation, a clustering algorithm can be used to analyze and classify the voice segments in the first audio signal, gather the voice segments with similar acoustic characteristics into the same category, with each category corresponding to a target object, and then assign an independent audio channel to each category, thereby establishing a mapping relationship.
[0076] In yet another feasible implementation, time continuity and semantic coherence can be combined to assist in establishing the mapping relationship, that is, analyze the context association and logical continuity of the voice content. When a certain voice segment has an obvious semantic association with the historical voice in a specific audio channel in terms of content, it can be attributed to that channel.
[0077] In an actual application scenario, although voiceprint recognition has high accuracy, it also requires a large amount of computing resources and is difficult to operate efficiently when processing a large amount of audio data in real time. By establishing a mapping relationship, the system only needs to quickly determine the correspondence between the target object and the audio channel using simple feature judgments, and only trigger a complete voiceprint recognition when necessary; it can allocate the mixed audio signal to the matching audio channel, thereby realizing the efficient recognition and processing of the audio signal.
[0078] In operation S103, the second audio signal refers to the audio data continuously obtained after the mapping relationship is established. In the embodiments of the present disclosure, it can be understood as the subsequent audio data containing the voice content of multiple target objects continuously collected by the electronic device after the initial mapping relationship is established, and is used for real-time recognition processing through the established mapping relationship.
[0079] Specifically, after establishing the mapping relationship between the target object and the audio channel based on the first audio signal, the electronic device continuously receives the second audio signal through the audio acquisition device. The system extracts the voice characteristics in the second audio signal and compares them with the characteristic parameters recorded in the mapping relationship to quickly determine the target object corresponding to each voice segment and the audio channel to which it should be assigned.
[0080] Subsequently, the system performs channel-level speech recognition on the second audio signal based on the allocation relationship to obtain independent speech signals corresponding to each target object. By performing real-time speech recognition on the second audio signal using the established mapping relationship, it is possible to ensure the stability and accuracy of speech recognition when the system processes continuously input mixed audio data, thereby providing a clear and non-overlapping single sound source input for subsequent operations such as speech recognition and semantic analysis, and improving the performance and efficiency of the overall speech processing system.
[0081] By adopting the embodiment of the present disclosure, a mapping relationship between the voiceprint features of the target object speech signal and the audio channel is established through the first audio signal, and the speech signals of each target object in the subsequent second audio signal are accurately allocated to the matching audio channels by using this mapping relationship, realizing continuous and stable audio signal recognition, thereby improving the accuracy of the audio signal during the recognition process.
[0082] In actual application scenarios, the audio signal processing system usually faces a complex multi-person interaction environment. For example, in a meeting room, classroom or home space, multiple target objects may be distributed at different positions and there may be changes in their positions. When the position information changes, it will cause the audio channels in the mapping relationship to change, resulting in the system being unable to correctly track and recognize the speech signals of the target objects. Especially during a long-term audio processing process, the target objects may change seats, move around or new target objects may join, and these changes will cause the original mapping relationship to fail, and the system cannot accurately allocate the speech signals to the corresponding audio channels, ultimately affecting the accuracy and stability of speech recognition.
[0083] To solve the above problems, on the basis of the above embodiment, as an alternative embodiment, the following operations may also be performed before performing operation S102:
[0084] Obtain the position information of each of the target objects and determine the audio channels corresponding to the position information matching each of the target objects.
[0085] Among them, the position information refers to the coordinate data or relative position description of the target object in space, and in the embodiment of the present disclosure, it can be understood as a data set representing the spatial position of the target object in a specific environment, which is used to assist the system in identifying and tracking the target object and provide reference information in the spatial dimension for audio signal recognition.
[0086] In the embodiments of the present disclosure, each audio channel can be assigned a corresponding relationship with a specific spatial position or area. Thus, the system can allocate the voice signal of a target object to a matching audio channel based on the position information of the target object, thereby realizing an effective association between the spatial position and the audio processing path. For example, when it is detected that the target object is located at a specific position in a meeting room, the system can automatically route the voice signal at this position to a pre-specified audio channel to ensure the correct classification and processing of the voice signal.
[0087] In a feasible implementation manner, the position information can be obtained through an audio device. The system uses the acoustic signals collected by a microphone array to estimate the azimuth angle and distance of the target object by analyzing parameters such as the time difference of arrival of sound waves and the direction of arrival of sound waves. For a system equipped with multiple distributed microphones, through triangulation or multi-point positioning technology, the signal characteristics collected by each microphone can be comprehensively analyzed to achieve precise spatial positioning of the target object. The above acoustic-based position acquisition method does not require additional hardware and is applicable to scenarios with only audio devices, such as a telephone conference system or a smart speaker, etc.
[0088] In another feasible implementation manner, the position information can be obtained through a camera device. The system captures video images through a camera and uses computer vision technologies such as face detection, pose estimation, or target tracking algorithms to identify the target object in the environment and determine its spatial position. For a system equipped with multiple cameras, more accurate three-dimensional position data can be obtained through multi-view fusion technology. The above vision-assisted position acquisition method provides more intuitive and accurate spatial information and is particularly suitable for application scenarios that require both audio and video processing, such as video conferencing and distance learning.
[0089] In yet another feasible implementation manner, the position information can be obtained by combining radio frequency signal positioning technology. The system can detect the signal strength or signal transmission time of a mobile device carried by the target object through wireless communication technologies such as Bluetooth and WiFi to infer its spatial position. Such technology is particularly suitable for scenarios that require high-precision indoor positioning, such as intelligent meeting rooms and collaborative spaces, etc., and can provide higher positioning accuracy to assist the system in accurately distinguishing different target objects at adjacent positions.
[0090] In still another feasible implementation manner, the position information can be obtained through a multi-source data fusion method. The system comprehensively utilizes various sensor data such as audio, video, and radio frequency, integrates various types of position information through data fusion, overcomes the limitations of a single sensor, and improves the accuracy of position estimation. The above fusion-based positioning method is applicable to high-precision positioning requirements in complex environments, such as large conferences and multi-person collaborative spaces, etc., and can maintain stable position tracking capabilities in situations where the target object moves, is blocked, or there is signal interference.
[0091] It should be noted that the location information acquisition methods in the embodiments of the present disclosure are not limited to the above several types. The system can select the most suitable location acquisition technology or a combination of multiple technologies according to the actual application environment, device configuration, and accuracy requirements. In addition, the representation form of location information can also be flexible and diverse, which can be accurate three-dimensional coordinates, relative position descriptions (such as "front left", "rear right corner"), or predefined location area numbers.
[0092] In addition, location information mainly plays an auxiliary role in quickly and preliminarily identifying audio signals during the audio signal recognition process. Since voiceprint recognition, although highly accurate, consumes a large amount of computing resources, it may cause system latency when processing a large amount of audio data in real time. By first using location information to quickly pre-identify the voice signal of the target object to the corresponding audio channel, the real-time performance and processing efficiency can be significantly improved. For example, when the system detects a voice signal from a specific location, it will first allocate it to the audio channel that may correspond to the target object, and then perform voiceprint matching to confirm the identity, significantly reducing the computing load of the system.
[0093] By adopting the embodiments of the present disclosure, through the association between location information and audio channels, it is possible to maintain the accuracy and stability of voice recognition even when the location information of the target object changes, thereby further enhancing the adaptability of audio signal processing in the actual application environment.
[0094] The following introduces the audio signal processing method in the embodiments of the present disclosure when the location information changes.
[0095] Based on the above embodiments, as an optional embodiment, operation S102 may specifically further include the following operations:
[0096] Operation S201, obtaining a second audio signal, and processing the second audio signal to obtain at least one target voice signal;
[0097] Operation S202, for any target voice signal, determining whether the location information of each target object changes based on the target voice signal;
[0098] Operation S203, if it is determined that the location information of the target object changes, updating the mapping relationship, and performing recognition on the second audio signal based on the updated mapping relationship;
[0099] Operation S204, if it is determined that the location information of the target object does not change, performing recognition on the second audio signal based on the mapping relationship.
[0100] In operation S202, the target voice signal refers to the audio signal corresponding to the voice content generated by a single target object identified from the second audio signal. In the embodiments of the present disclosure, it can be understood as the voice signal obtained after preliminary identification processing, which is used to determine whether the identity and position of the target object have changed, so as to decide whether to update the mapping relationship.
[0101] Specifically, the system analyzes each identified target voice signal to determine whether the position information of the target object has changed. Exemplarily, the sound features in the target voice signal can be extracted and compared with the features indicated by the mapping relationship. If it is detected that the feature parameters do not match the features indicated in the mapping relationship, it may indicate that the position of the target object has changed and the mapping relationship needs to be updated. Based on the above embodiments, as a feasible embodiment, the present disclosure provides corresponding mapping relationship update strategies for different types of position information change situations. The change in position information mainly includes, but is not limited to, one or more of the following situations:
[0102] Situation 1: The relative positions between the target objects have changed;
[0103] Situation 2: One or more other target objects are added;
[0104] Situation 3: One or more target objects are replaced by other target objects.
[0105] Figures 2A to 2D FIG. shows a schematic diagram of scenarios with various changes in position information provided by the embodiments of the present disclosure. The following will be described in conjunction with Figures 2A to 2D the above process
[0106] As Figure 2A shown, in the initial state, the meeting scene is divided into multiple regions, including the left region 211, the right region 212, and the upper region 213. The target object A201 is located in the left region 211, the target object B202 is located in the right region 212, and the target object C203 is located in the upper region 213. The system has established a complete mapping relationship, specifically: region 211 - target object A201 - audio channel 1, region 212 - target object B202 - audio channel 2, region 213 - target object C203 - audio channel 3.
[0107] For the above situation 1, that is, when the relative positions between the target objects have changed, as Figure 2BAs shown, the target object A201 moves from area 211 to area 212, while the target object B202 moves from area 212 to area 211, and the two exchange positions. The target object C203 remains unchanged in area 213. In operation S204, the system first detects that the target voice signals of audio channel 1 and audio channel 2 do not match the features recorded in the original mapping relationship. Through further analysis, the system finds that the feature of the voice signal currently assigned to audio channel 1 (corresponding to area 211) is more matched with the reference feature of audio channel 2 in the mapping relationship, while the feature of the voice signal assigned to audio channel 2 (corresponding to area 212) is more matched with the reference feature of audio channel 1 in the mapping relationship. Based on this cross-matching relationship, the system determines that the target objects A201 and B202 have exchanged positions.
[0108] For the above situation, the system updates the mapping relationship as: area 211 - target object B202 - audio channel 1, area 212 - target object A201 - audio channel 2, while the mapping relationship of area 213 - target object C203 - audio channel 3 remains unchanged. The system exchanges the target object feature parameters of audio channel 1 and audio channel 2, associates the feature parameters of target object B202 with audio channel 1, and associates the feature parameters of target object A201 with audio channel 2. The above update strategy ensures that the system can correctly track each target object. Even if they are rearranged in space, their voice signals can be assigned to the correct audio channels, thus maintaining the accuracy of speech recognition.
[0109] For the above situation 2, that is, when adding one or more other target objects, such as Figure 2C As shown, the system identifies that a new lower area 214 has been added to the meeting scenario. Based on the original three target objects, a new target object D204 is added, located in the lower area 214. In operation S204, when the system processes the second audio signal, it finds that there is a voice part that cannot be matched to the existing audio channels, and the feature of this part of the voice does not match any of the target object features recorded in the mapping relationship, indicating the emergence of a new target object.
[0110] For the above situation, the system extension mapping relationship is as follows: Region 211 - Target Object A201 - Audio Channel 1, Region 212 - Target Object B202 - Audio Channel 2, Region 213 - Target Object C203 - Audio Channel 3, Region 214 - Target Object D204 - Audio Channel 4. The system creates a new Audio Channel 4 for the newly added Target Object D204. The system first keeps the mapping relationship between the original target objects and audio channels unchanged, then extracts the voice features of the new target object, initializes the reference feature parameters of Audio Channel 4, and adds them to the mapping relationship. The updated mapping relationship will include four audio channels, corresponding to the voice signals of four target objects respectively. The above incremental update strategy can not only adapt to the situation of newly added target objects, but also ensure the continuity and stability of the system operation, and ensure the integrity of audio processing.
[0111] For the above Situation 3, that is, when one or more target objects are replaced by other target objects, such as Figure 2D As shown, Target Object C203 leaves the upper Region 213, and its position is replaced by the new Target Object E205. In operation S204, when the system processes the second audio signal, it is found that the features of the voice signal assigned to Audio Channel 3 are significantly different from the reference features of Target Object C203 recorded in the mapping relationship, but the location information (Region 213) is close to the original Target Object C203, indicating that the target object at this position has been replaced.
[0112] For the above situation, the system updates the mapping relationship as follows: Region 211 - Target Object A201 - Audio Channel 1, Region 212 - Target Object B202 - Audio Channel 2, Region 213 - Target Object E205 - Audio Channel 3. The system first keeps the channel number and spatial association of Audio Channel 3 (the corresponding relationship with Region 213) unchanged, then extracts the voice features of the new Target Object E205, and updates the reference feature parameters of Audio Channel 3. Through the above replacement update, the system can correctly assign the voice signal of Target Object E205 to Audio Channel 3 while keeping the mapping relationships of other audio channels unchanged. The above update strategy is applicable to scenarios where the seats are fixed but the participants may change, such as shift workstations, fixed seats in classrooms or designated positions in meeting rooms, etc. The system can quickly adapt to the replacement of target objects according to the continuity of spatial positions.
[0113] In practical applications, the system may also face more complex changes in location information, such as combinations or mixed forms of the above three basic situations. For such complex scenarios, the system adopts a step-by-step processing strategy in operation S204, first processes the clear replacement relationships, then identifies the newly added target objects, and finally processes the changes in relative positions. This hierarchical processing method can reduce the difficulty of updating the mapping relationship in complex scenarios and improve the adaptability of the system.
[0114] To improve the reliability and efficiency of mapping relationship updates, the system adopts an incremental update strategy, that is, only the changed parts are modified, and the unchanged mapping relationships are retained. For example, in Figure 2B the situation shown, the system only needs to exchange the target object and audio channel mapping relationships corresponding to region 211 and region 212, without affecting the mapping relationship of region 213; in Figure 2C the situation shown, the system only needs to add the mapping relationship of region 214 - target object D204 - audio channel 4, while keeping the mapping relationships of the other three regions unchanged; in Figure 2D the situation shown, the system only needs to update the target object of region 213 from C203 to E205, and correspondingly update the target object characteristic parameters of audio channel 3, without affecting the mapping relationships of other regions. This targeted update method can reduce the consumption of computing resources and speed up the update speed, and is particularly suitable for large-scale audio processing scenarios. By adopting the embodiments of the present disclosure, corresponding mapping relationship update strategies are adopted according to different position information change situations, so as to ensure that the system maintains stable audio signal recognition performance in a complex and changeable environment.
[0115] In actual application scenarios, the audio signal processing system may face the situation of failed recognition of the second audio signal. Usually, it includes the following situations: when the voice characteristics of the target object change suddenly and significantly, such as the pitch significantly increases due to excitement; when the acoustic characteristics change due to sudden changes in environmental conditions, such as the reverberation change caused by the opening and closing of the doors and windows in the meeting room; when multiple target objects move positions simultaneously, causing the system to be unable to accurately track; or when part of the microphone array fails, resulting in a decrease in spatial positioning ability. The above situations may all cause the system to be unable to correctly recognize the second audio signal based on the current mapping relationship, thus affecting the continuity and stability of the overall audio processing.
[0116] To solve the above problems, the embodiments of the present disclosure provide a backtracking mechanism based on the mapping relationship of historical versions. On the basis of the above embodiments, as an optional embodiment, the above audio signal processing method may further include the following operations:
[0117] In the case of failed recognition of the second audio signal, obtain the mapping relationship of the historical version, where the mapping relationship of the historical version is the mapping relationship of a preset number of rounds before the current mapping relationship;
[0118] Recognize the second audio signal based on the mapping relationship of the historical version.
[0119] Among them, the historical version of the mapping relationship refers to the data record of the correspondence between the target object voice signal and the audio channel previously established and saved by the system during the audio processing process. In the embodiment of the present disclosure, it can be understood as a copy of the original mapping relationship automatically saved by the system before performing the mapping relationship update operation. It is used to provide a fallback option when the current mapping relationship fails or the recognition result is not ideal, ensuring that the system has the ability to restore to a previously verified valid state.
[0120] The preset round refers to the tracing depth or number of version intervals when the system traces back the historical version mapping relationship, indicating the number of versions traced back from the current mapping relationship.
[0121] For example, in different application scenarios, the setting method of preset rounds may be different: in a high-frequency interactive conference scenario, because the position and speech status of the participants change quickly, the system may set the preset rounds to a smaller value (such as 1-2 rounds) to ensure that the backtracked mapping relationship still has a high relevance to the current scenario; in a relatively stable classroom teaching environment, because the position of the participants is relatively fixed and changes slowly, the system can set the preset rounds to a medium value (such as 3-5 rounds) to provide more backtracking options while maintaining the relevance of the scene; in occasions such as long-term recording of professional meetings or court hearings, in order to cope with possible complex changes in the situation, the system may adopt a dynamic preset round strategy, that is, automatically adjust the backtracking depth according to the update frequency of the mapping relationship and the stability of the environment, and use a larger round (such as 5-10 rounds) when the environment is stable, and a smaller round (such as 1-3 rounds) when the environment changes frequently. In addition, for specific key scenes (such as the entry of important speakers or the start of topic discussions), the system may set marked versions and specially retain the mapping relationships of these versions, so that even if it exceeds the conventional preset round range, it can be traced back to the mapping status of these key moments when necessary.
[0122] Specifically, during the mapping relationship backtracking process, the system first attempts to use the most recent historical version of the mapping relationship to identify the second audio signal. If the recognition still fails, the system will gradually backtrack to earlier historical versions until it finds a mapping relationship version that can successfully identify the signal, or tries all available historical versions. This progressive backtracking strategy can maximize the possibility of restoring normal processing capabilities while maintaining system response speed.
[0123] By adopting the embodiment of the present disclosure, through the historical version mapping relationship backtracking mechanism, it is possible to effectively deal with the situation where the second audio signal recognition fails, ensuring that the system maintains continuous and stable audio processing capabilities in a complex and changing environment.
[0124] In actual application scenarios, audio signal processing systems usually face complex and diverse user compositions, and different target objects may have different identity recognition states. For example, in a corporate meeting environment, regular employees may have registered their voiceprint features in the system, while temporary visitors or newly joined members may not have registered; in an educational environment, teachers may have completed voiceprint registration, while students may not have registered voiceprints due to their large numbers or frequent changes; in a home scenario, family members may have registered voiceprints, while visitors may not have registered. The differences in the identity recognition states of the above target objects result in the system being unable to adopt a unified recognition and recognition strategy for all target objects, thus affecting the integrity and accuracy of audio signal processing.
[0125] To solve the above problems, the embodiments of the present disclosure classify target objects into two categories: the first object and the second object, and adopt different processing strategies for different types of target objects. Among them, the first object is the object that has registered the voiceprint, and the second object is the object that has not registered the voiceprint.
[0126] For these two different types of target objects, the role played by location information in the audio signal processing process also has significant differences.
[0127] Based on the above embodiments, as an alternative embodiment, for the first object, that is, the registered voiceprint feature of the target object, the above audio signal processing method may further include the following operations:
[0128] If it is determined that the location information of the target object changes, obtain the registered voiceprint feature of the target object, and update the mapping relationship based on the voiceprint feature of the target object, so that the mapping relationship includes the corresponding relationship between the voiceprint feature of the target object and the audio channel.
[0129] Specifically, when the system confirms the change in location information, the system will retrieve the voiceprint feature model of the target object from a pre-established voiceprint database. Exemplarily, the voiceprint feature model can usually be composed of one or more of parameters such as Mel Frequency Cepstral Coefficients, Linear Prediction Coefficients, and fundamental frequency features of the voice.
[0130] After the system extracts the voiceprint feature of the target object from the voiceprint database, it establishes a corresponding relationship with the new audio channel and updates this corresponding relationship to the mapping relationship. The updated mapping relationship not only includes the basic corresponding relationship between the target object and the audio channel, but also adds the voiceprint feature, a highly recognizable identity identifier.
[0131] Exemplarily, such as Figure 2BIn the scene shown, when the target object A201 moves from area 211 to area 212, the system not only updates the basic mapping relationship of area 212 - target object A201 - audio channel 2 to the mapping table, but also extracts the voiceprint feature of target object A201 from the voiceprint database, and updates the corresponding relationship of voiceprint feature - target object A201 - audio channel 2 to the mapping table at the same time.
[0132] By integrating the voiceprint features of the registered target objects into the mapping relationship, the system can maintain accurate recognition of their voice signals in an environment where the positions of the target objects change frequently. Even if the location information of the target object is temporarily unavailable or inaccurate, the system can still allocate its voice signal to the correct audio channel through voiceprint matching.
[0133] In addition, integrating the voiceprint features into the mapping relationship can also reduce the call frequency of the system for complete voiceprint recognition during audio processing. Traditional voiceprint recognition requires complete feature extraction and model matching for each piece of speech, with high computational complexity and difficulty in meeting the requirements of real-time processing. In the embodiments of the present disclosure, the system pre-associates the voiceprint features with specific audio channels. When processing the second audio signal, it only needs to quickly match the voice signal with the pre-allocated voiceprint features, improving the processing efficiency.
[0134] Adopting the embodiments of the present disclosure, by integrating the registered voiceprint features into the mapping relationship when the location information changes, the system can improve the processing efficiency while maintaining the recognition accuracy, providing more stable technical support for the audio recognition of multiple target objects in a complex dynamic environment. For the second object, that is, the target object without a registered voiceprint, the location information becomes the main basis for voice recognition and personnel differentiation. Since the system lacks the voiceprint feature model of the second object and cannot accurately distinguish its identity through voiceprint recognition technology, it can only rely on the location information to distinguish different second objects and allocate their voice signals to the corresponding audio channels. In the above situation, the location information provides stable spatial differentiation ability. The system can effectively distinguish the second objects from different locations by analyzing the spatial direction features of the voice signals, even if their voice features are highly similar. Especially in the case where multiple second objects speak simultaneously and form voice overlaps, the system can still effectively identify the voice signals of different second objects through the spatial direction differences, preventing chaos. For example, in a conference room scenario, even if several first-time visitors speak simultaneously, the system can allocate their respective voice signals to different audio channels based on their different positions, maintaining the orderliness of voice processing. Further, in the mapping relationships obtained at multiple time nodes, if the number of times the position of a certain second object does not change, and / or the number of times it is recognized as a second object is greater than a specific threshold, the corresponding voice signal in the second object can be stored as a voiceprint feature.
[0135] It should be noted that the location information also plays different roles for the two types of target objects during the mapping relationship update process. For the first object, the change in location information mainly triggers the update of the corresponding relationship between the spatial coordinates and the audio channels in the mapping relationship, while the voiceprint feature part remains stable; for the second object, the change in location information not only involves the update of the spatial coordinates, but may also cause the system to need to reconstruct the temporary sound feature model to adapt to the acoustic environment characteristics of the new location. Therefore, when the system detects a change in location information, it will adopt corresponding mapping relationship update strategies according to the type of target object to ensure a continuous and stable audio signal recognition effect.
[0136] By adopting the embodiments of the present disclosure and distinguishing the different roles of location information for the first object with registered voiceprint and the second object without registered voiceprint, the system can flexibly adapt to different user compositions in various complex application scenarios, effectively balance the relationship between computing resources and recognition accuracy, and provide a more stable and efficient audio signal processing ability.
[0137] The following introduces the audio signal processing methods for different target objects in the embodiments of the present disclosure.
[0138] Based on the above embodiments, the operation of determining whether the location information of each target pair changes based on the target voice signal may specifically further include the following operations:
[0139] Operation S301, obtaining the target features in the target voice signal;
[0140] Operation S302, if it is determined that the target features match the features indicated by the mapping relationship, it is determined that the location information has not changed;
[0141] Operation S303, if it is determined that the target features do not match the features indicated by the mapping relationship, it is determined that the location information has changed;
[0142] Among them, the target features include at least one of voice features and semantic features.
[0143] Specifically, for the first object, the system can use its registered voiceprint feature as the main judgment basis. The system extracts the voiceprint feature vector from the target voice signal and calculates the similarity with the registered voiceprint model stored by the first object in the mapping relationship. If the similarity exceeds the preset threshold, it is considered that the voice signal indeed comes from the first object and the location information has not changed; otherwise, if the similarity is lower than the threshold, it indicates that the voice signal received by the current audio channel may come from other target objects and the location information has changed.
[0144] In addition, the system can also assist in comprehensively judging by combining the voice features formed by the first object in the first audio signal, including but not limited to: voice frequency, voice amplitude, timbre characteristics, speech rate pattern, etc.; semantic features, including but not limited to: lexical expression, sentence structure, etc., to further improve the recognition accuracy. For example, when the voiceprint matching degree is in a critical state, the system can assist in confirming the identity by analyzing the semantic coherence between the speech content and the historical speech of the first object.
[0145] For the second object, due to the lack of registered voiceprint features, the system mainly relies on the voice features and semantic features formed by it in the first audio signal for judgment. The system extracts voice features and semantic features from the target voice signal and matches them with the reference features of the second object recorded in the mapping relationship.
[0146] In practical applications, the system will also dynamically adjust the strategies for feature extraction and matching according to the requirements of different scenarios. For example, in an environment with high noise, the system may rely more on the voiceprint features of the first object and reduce the weight of voice features; in a long meeting scenario, the system may gradually increase the influence of semantic features because as the content accumulates, semantic coherence can provide more judgment basis; in a fast interaction scenario, the system may preferentially use voice features with lower computational complexity to improve the response speed.
[0147] By adopting the embodiments of the present disclosure and using different target feature extraction and matching strategies for the first object and the second object respectively, the system can accurately judge the change situation of the position information of different types of target objects, thereby accurately updating the mapping relationship and ensuring a continuous and stable audio signal recognition effect.
[0148] In practical application scenarios, for the second object without a registered voiceprint, the system mainly relies on the voice features and semantic features extracted from the first audio signal for identity recognition and audio recognition. However, as the usage time extends, the voice features of the second object may drift. For example, during a long continuous speech by a person, due to emotional changes, acoustic parameters such as pitch, speech rate, and volume gradually change; or changes in environmental conditions (such as changes in air density caused by room temperature changes, changes in sound wave propagation characteristics caused by humidity fluctuations, etc.) will also affect the collection and processing of sounds; over time, the impact of the above-mentioned feature drift phenomenon on the second object is obvious, and the originally established mapping relationship may gradually lose accuracy.
[0149] To solve the above-mentioned feature drift problem, the embodiments of the present disclosure provide a dynamic update mechanism for the mapping relationship. On the basis of the above embodiments, as an optional embodiment, when the target object is the second object, the above audio signal processing method may further include the following operations:
[0150] In the case where the difference degree between the target feature of the second object and the target feature recorded in the mapping relationship is greater than a preset threshold, update the target feature recorded for the second object in the mapping relationship.
[0151] Specifically, during the process of the system processing the second audio signal, it regularly extracts the latest acoustic features and semantic features from the recognized second object speech signal, and then calculates the difference degree between these new features and the reference features already recorded in the mapping relationship. Exemplarily, the calculation method can adopt feature vector distance measurement algorithms such as Euclidean distance and cosine similarity.
[0152] When the calculated difference degree exceeds the preset update threshold, the system will trigger an update operation of the mapping relationship, use the newly extracted features to replace or fuse the original features, form updated reference features, and store them in the corresponding second object record in the mapping relationship.
[0153] As an optional embodiment, to ensure the smoothness and continuity of feature updates, the system can adopt a weighted fusion method to combine the new features and the old features according to a certain weight ratio to generate updated reference features. This avoids feature instability caused by single-sampling fluctuations and at the same time retains the continuity of historical feature information. The weight allocation can be flexibly adjusted according to the application scenario and system requirements. For example, in an environment with high noise, the weight of historical features can be increased to improve stability, while in an environment with good acoustic conditions, the weight of new features can be increased to improve adaptability. In addition, the system can also set multiple levels of preset thresholds, and adopt different update strategies and weight configurations when the difference degree is in different threshold intervals.
[0154] Adopting the embodiments of the present disclosure, through the dynamic update mechanism, the system can effectively address the problem of drift of the acoustic features of the second object and maintain the continuous accuracy of the mapping relationship.
[0155] In actual application scenarios, audio signal processing systems often face the problem of speech overlap caused by multiple people speaking simultaneously. When two or more target objects speak at the same time, traditional audio recognition technologies are difficult to accurately distinguish each sound source, resulting in a significant decrease in the speech recognition rate, which in turn affects the accuracy of subsequent content extraction and information processing. Especially in scenarios such as meeting discussions, classroom interactions, or multi-person remote collaborations, the phenomenon of speech overlap occurs frequently, seriously reducing the effectiveness of audio signal processing systems. Related technologies use conventional single-channel audio inputs and traditional noise reduction methods and cannot effectively solve the problem of recognition accuracy in the scenario of multi-person overlapping speech.
[0156] To solve the above problems, on the basis of the above embodiments, as an optional embodiment, in the case where the target object is the first object, the audio signal processing method may further include the following operations:
[0157] Operation S401, when it is determined that there is an overlapping audio signal in the second audio signal, obtain the timestamp of the overlapping audio signal, where the overlapping audio signal is an audio signal in which the speech signals of at least two objects overlap;
[0158] Operation S402, extract the speech signal of the first object in the overlapping audio signal based on the voiceprint feature of the first object;
[0159] Operation S403, based on the timestamp of the overlapping audio signal, replace the overlapping audio signal with the speech signal of the first object.
[0160] In operation S401, the overlapping audio signal refers to a mixed speech signal generated by multiple target objects speaking simultaneously, manifested as an audio segment with superimposed sound waves, complex energy distribution, and mixed acoustic features. The system identifies the audio segments that may have overlapping speech by analyzing the acoustic features of the second audio signal, such as methods of energy distribution, spectral structure, and sound source number estimation. When it is determined that there is an overlapping audio signal, the system records the timestamp of the overlapping audio signal, that is, the start time and end time of the overlapping speech segment in the overall audio stream.
[0161] In operation S402, the system will use the registered voiceprint feature of the first object to accurately extract the speech content of the target object from the overlapping audio signal.
[0162] In a feasible implementation manner, the voiceprint feature extraction process not only considers static voiceprint features but also combines dynamic context information. The system analyzes the speech patterns of the first object before and after the overlap, and uses the continuity features of the speech to assist in the identification of the overlapping part, further improving the extraction accuracy.
[0163] In operation S403, the system replaces the original overlapping audio signal with the speech signal of the first object extracted in operation S402 based on the previously recorded timestamp of the overlapping audio signal.
[0164] Exemplarily, the execution of the replacement operation requires precise time alignment and smooth transition processing. First, locate the exact position in the original audio stream that needs to be replaced according to the recorded timestamp, and then insert the extracted speech signal of the first object into this position.
[0165] In some application scenarios, the system may need to retain the original overlapping audio signal as a backup so that users can view the original content retrospectively when necessary. To this end, the system can, while performing the replacement operation, save the original overlapping audio signal and its timestamp information in an auxiliary storage area and provide an interface to allow users to switch and view when needed.
[0166] By adopting the embodiments of the present disclosure, through overlapping speech recognition and replacement, the problem of speech overlap caused by multiple people speaking simultaneously can be effectively addressed. By precisely extracting the speech content of the first object using its registered voiceprint feature and replacing the original overlapping audio signal, the accuracy of subsequent speech recognition and content analysis is significantly improved.
[0167] The embodiments of the present disclosure also disclose an electronic device, including:
[0168] An audio device for collecting a first audio signal, where the first audio signal carries at least the speech signals of two different target objects;
[0169] A processor for determining a mapping relationship based on the speech signals of the respective target objects in the first audio signal, where the mapping relationship represents the correspondence between the speech signals of the respective target objects and audio channels; and recognizing the second audio signal based on the mapping relationship to obtain the speech signals corresponding to at least one target object, where the second audio signal is determined after the first audio signal.
[0170] Figure 3 The block diagram of an electronic device provided by the embodiments of the present disclosure is schematically shown. Figure 3 The shown electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0171] As Figure 3 shown, the electronic device 300 according to the embodiments of the present disclosure includes an audio device (not shown in the figure) and a processor 301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the memory 308 into the random access memory (RAM) 303. The processor 301 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), and so on. The processor 301 may also include on-board memory for caching purposes. The processor 301 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiments of the present disclosure.
[0172] In the RAM 303, various programs and data required for the operation of the electronic device 300 are stored. The processor 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. The processor 301 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 302 and / or the RAM 303. It should be noted that the programs may also be stored in one or more memories other than the ROM 302 and the RAM 303. The processor 301 may also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.
[0173] According to an embodiment of the present disclosure, the electronic device 300 may further include an input / output (I / O) interface 304, and the input / output (I / O) interface 304 is also connected to the bus 304. The system 300 may further include one or more of the following components connected to the input / output (I / O) interface 304: an input device 306 including a keyboard, a mouse, etc.; an output device 307 including a cathode ray tube (CRT), a liquid crystal display (LCD), a display screen, etc. and a speaker, etc.; a memory 308 including a hard disk, etc.; and a communication part 309 including a network interface card such as a LAN card, a modem, etc. The communication part 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the input / output (I / O) interface 304 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read from it can be installed into the memory 308 as needed.
[0174] According to an embodiment of the present disclosure, the method flow according to the embodiments of the present disclosure may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network via the communication part 309, and / or installed from the removable medium 311. When the computer program is executed by the processor 301, the above-mentioned functions defined in the system according to the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described system, device, apparatus, module, unit, etc. may be implemented by computer program modules.
[0175] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the methods according to the embodiments of the present disclosure are implemented.
[0176] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.
[0177] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 302 and / or RAM 303 and / or one or more memories other than ROM 302 and RAM 303.
[0178] Embodiments of the present disclosure further include a computer program product, which includes a computer program that contains program code for executing the methods provided by the embodiments of the present disclosure. When the computer program product runs on an electronic device, the program code is used to cause the electronic device to implement the methods provided by the embodiments of the present disclosure.
[0179] When the computer program is executed by the processor 301, the above functions defined in the system / apparatus of the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0180] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 309, and / or be installed from the removable medium 311. The program code included in the computer program may be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0181] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the above-mentioned module, segment of a program, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art can understand that the features recited in various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0182] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. An audio signal processing method, comprising: Obtaining a first audio signal, where the first audio signal carries at least two different target objects' speech signals; Based on the speech signals of the respective target objects in the first audio signal, determining a mapping relationship, where the mapping relationship represents the corresponding relationship between the speech signals of the respective target objects and audio channels; Based on the mapping relationship, identifying the second audio signal to obtain the speech signals corresponding to at least one target object, where the second audio signal is obtained after the first audio signal.
2. The method according to claim 1, before determining the mapping relationship based on the speech signals of the respective target objects in the first audio signal, further comprising: Obtaining the position information of each of the target objects; Determining the audio channels corresponding to the position information matching each of the target objects.
3. The method according to claim 2, the identifying the second audio signal based on the mapping relationship includes: Obtaining a second audio signal, and processing the second audio signal to obtain at least one target speech signal; For any one of the target speech signals, based on the target speech signal, determining whether the position information of each of the target objects changes; If it is determined that the position information of the target object changes, updating the mapping relationship, and identifying the second audio signal based on the updated mapping relationship; If it is determined that the position information of the target object does not change, identifying the second audio signal based on the mapping relationship.
4. The method according to claim 3, further comprising: If it is determined that the position information of the target object changes, obtaining the registered voiceprint feature of the target object, and updating the mapping relationship based on the voiceprint feature of the target object, so that the mapping relationship includes the corresponding relationship between the voiceprint feature of the target object and the audio channel.
5. The method according to claim 3, the determining whether the position information of each of the target objects changes based on the target speech signal includes: Obtaining the target feature in the target speech signal; If it is determined that the target feature matches the feature indicated by the mapping relationship, determining that the position information does not change; If it is determined that the target feature does not match the feature indicated by the mapping relationship, determining that the position information changes; Wherein, the target feature includes at least one of a voice feature and a semantic feature.
6. The method according to claim 5, the target object is a second object, and the voiceprint feature of the second object is not registered. After obtaining the target feature in the target speech signal, further comprising: In the case where the difference degree between the target feature of the second object and the target feature recorded in the mapping relationship is greater than a preset threshold, updating the target feature recorded for the second object in the mapping relationship.
7. The method according to claim 4, the target object is the first object, and the voiceprint feature of the first object is registered. The method further comprises: When it is determined that there is an overlapping audio signal in the second audio signal, obtain the timestamp of the overlapping audio signal, where the overlapping audio signal is an audio signal in which the voice signals of at least two objects overlap; Extract the voice signal of the first object in the overlapping audio signal based on the voiceprint feature of the first object; Replace the overlapping audio signal with the voice signal of the first object based on the timestamp of the overlapping audio signal.
8. The method according to claim 3, wherein the change in the position information of each of the target objects includes at least one of the following: The relative positions between the target objects change; One or more other target objects are added; One or more of the target objects are replaced by other target objects.
9. The method according to claim 1, further comprising: When the recognition of the second audio signal fails, obtain the mapping relationship of the historical version, where the mapping relationship of the historical version is the mapping relationship of a preset number of rounds before the current mapping relationship; Recognize the second audio signal based on the mapping relationship of the historical version.
10. An electronic device, comprising: An audio device for collecting a first audio signal, where the first audio signal carries at least voice signals of two different target objects; A processor for determining a mapping relationship based on the voice signals of the target objects in the first audio signal, where the mapping relationship represents the correspondence between the voice signals of the target objects and the audio channels; Recognize a second audio signal based on the mapping relationship to obtain the voice signals corresponding to at least one target object, where the second audio signal is obtained after the first audio signal.