Conference voice recognition method and device, electronic equipment and storage medium

Through the collaborative acquisition of multi-sound pickup devices and the fusion of multi-modal information, the problem of reduced speech recognition accuracy in multi-person meetings is solved, and high-quality speech recognition results are achieved.

CN120472883APending Publication Date: 2025-08-12HANVON CORP

Patent Information

Application Number
CN202510686745.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Existing voice recognition systems are susceptible to noise interference in multi-person meetings, resulting in a decrease in speech recognition accuracy.

Method used

Multiple sound pickup devices collaborately collect conference audio, conduct conference scene consistency judgment, filter high-quality audio bands and perform multi-difference splicing, and combine visual information to fusion for voice recognition.

Benefits of technology

It improves the accuracy and robustness of speech recognition, effectively solving the signal attenuation problem of a single device in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472883A_ABST
    Figure CN120472883A_ABST
Patent Text Reader

Abstract

The invention discloses a conference voice recognition method and device and electronic equipment, and belongs to the technical field of voice recognition. The method comprises the following steps: performing conference scene consistency judgment on conference audios collected by a plurality of pickup devices, and obtaining conference audios matched with a target conference scene in the conference audios; performing segmented screening and multi-device splicing processing on the conference audio matched with the target conference scene to obtain a spliced audio of the target conference scene; performing multi-modal information fusion on the pre-collected visual information and the spliced audio of the target conference scene to obtain multi-modal fusion information; and performing voice recognition based on the multi-modal fusion information to obtain a conference voice recognition result of the target conference scene. According to the method, conference audios of a single conference scene are cooperatively acquired by using multiple pickup devices, and high-quality voice signals are ensured to be obtained; the multi-modal information is fused in the audio signal for speech recognition, various data are comprehensively captured and processed, and the accuracy and robustness of speech recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition technology, and in particular to conference speech recognition methods, devices, electronic devices, and computer-readable storage media. Background Art

[0002] With the development of speech recognition technology, its application scenarios are expanding, and it has become a key means of recording and transcribing content in conference settings. However, existing speech recognition systems typically rely on a single device for audio pickup. In noisy environments or scenarios with multiple sound sources, such as multi-person conferences, the collected audio signals are susceptible to noise interference, resulting in reduced speech recognition accuracy.

[0003] It can be seen that the conference speech recognition method in the existing technology still needs to be improved. Summary of the Invention

[0004] The embodiments of the present application provide a conference speech recognition method and device, which help to improve the accuracy of conference speech recognition.

[0005] In a first aspect, an embodiment of the present application provides a conference speech recognition method, comprising:

[0006] Performing conference scene consistency judgment on conference audio collected by multiple sound pickup devices, and obtaining conference audio that matches the target conference scene in the conference audio;

[0007] Performing segmented screening and multi-device splicing processing on the conference audio matching the target conference scene to obtain the spliced audio of the target conference scene;

[0008] Performing multimodal information fusion on the pre-collected visual information of the target conference scene and the spliced audio to obtain multimodal fusion information;

[0009] Speech recognition is performed based on the multimodal fusion information to obtain a conference speech recognition result of the target conference scene.

[0010] In a second aspect, an embodiment of the present application provides a conference speech recognition device, comprising:

[0011] A conference scene consistency judgment module is used to judge the conference scene consistency of the conference audio collected by multiple sound pickup devices, and obtain the conference audio that matches the target conference scene in the conference audio;

[0012] A spliced audio acquisition module is used to perform segmented screening and multi-device splicing processing on the conference audio matching the target conference scene to obtain the spliced audio of the target conference scene;

[0013] A multimodal fusion module, configured to perform multimodal information fusion on the pre-collected visual information of the target conference scene and the spliced audio to obtain multimodal fusion information;

[0014] The speech recognition module is used to perform speech recognition based on the multimodal fusion information to obtain a conference speech recognition result of the target conference scene.

[0015] In a third aspect, an embodiment of the present application further discloses an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the conference speech recognition method described in the embodiment of the present application when executing the computer program.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the conference speech recognition method disclosed in the embodiment of the present application are implemented.

[0017] The conference speech recognition method disclosed in the embodiment of the present application performs conference scene consistency judgment on the conference audio collected by multiple sound pickup devices to obtain the conference audio matching the target conference scene in the conference audio; performs segmented screening and multi-device splicing processing on the conference audio matching the target conference scene to obtain the spliced audio of the target conference scene; performs multimodal information fusion on the pre-collected visual information of the target conference scene and the spliced audio to obtain multimodal fusion information; performs speech recognition based on the multimodal fusion information to obtain the conference speech recognition result of the target conference scene. This method introduces multiple sound pickup devices to collaboratively collect the conference audio of a single conference scene, making the sound source information more comprehensive; combines the segmented alignment technology to perform quality scoring on the conference audio collected by multiple sound pickup devices, thereby screening high-quality audio segments for audio splicing, effectively utilizing the conference audio collected by multiple sound pickup devices, and ensuring that the final result is a high-quality speech signal; introduces visual information into the audio signal, fuses multimodal information for speech recognition, and comprehensively captures and processes multiple data, further improving the accuracy and robustness of speech recognition. By combining the above technical means, this method effectively solves the signal attenuation caused by factors such as distance and noise when a single device picks up sound, and effectively improves the accuracy and robustness of conference speech recognition results.

[0018] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0020] Figure 1 This is a flowchart of the steps of the conference speech recognition method disclosed in the embodiment of the present application;

[0021] Figure 2 This is a schematic diagram of the entire process of the conference speech recognition method disclosed in the embodiment of this application;

[0022] Figure 3 This is a schematic diagram of the structure of the conference speech recognition device disclosed in the embodiment of this application;

[0023] Figure 4 A block diagram schematically shows an electronic device for executing the method according to the present application; and

[0024] Figure 5 The figure schematically shows a storage unit for storing or carrying a program code for implementing the method according to the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0026] The conference speech recognition method disclosed in the embodiments of this application is applicable to scenarios where multiple sound pickup devices are configured to collect conference audio. The multiple sound pickup devices include one or more sound pickup devices located in different locations in the conference room, configured to collect conference speech from multiple angles, thereby ensuring the integrity of the conference speech collection. The types of sound pickup devices include, but are not limited to, one or more of the following: recording devices, mobile terminals, desktop computers, laptop computers, and other electronic devices with recording capabilities.

[0027] In an embodiment of the present application, the sound pickup device can upload the collected conference audio to the server, and the server will judge and process the received conference audios to obtain all conference audios belonging to the same conference scene. Afterwards, for each conference scene, all conference audios belonging to the conference scene are segmented and aligned according to the timestamps of the audios to obtain multiple audio segments corresponding to each timestamp. Then, the quality of each audio segment in each conference audio is further judged, so as to select the audio segment with the highest quality corresponding to each timestamp from the multiple conference audios of the conference scene as the retained audio segment. Finally, the retained audio segments are spliced into the audio of the conference scene to perform speech recognition based on the audio.

[0028] The following combination Figure 2 The flowchart shown illustrates the specific implementation of the present application in detail.

[0029] Reference Figure 1 , a conference speech recognition method disclosed in an embodiment of the present application includes: steps 110 to 140.

[0030] Step 110 , performing conference scene consistency judgment on the conference audios collected by multiple sound pickup devices, and obtaining conference audios matching the target conference scene from the conference audios.

[0031] The conference audio collected by the multiple sound pickup devices refers to conference audio collected in the same time period, or conference audio collected with an overlapping time greater than or equal to a preset overlapping time condition. The preset overlapping time condition is set according to application requirements.

[0032] After obtaining multiple conference audios collected by multiple sound pickup devices, the scene consistency of these multiple conference audios is first judged from multiple dimensions, so as to obtain all conference audios belonging to the same conference scene.

[0033] In some optional embodiments, the conference scene consistency judgment is performed on the conference audio collected by multiple sound pickup devices, and the conference audio matching the target conference scene in the conference audio is obtained, including: obtaining a multi-dimensional matching score of the sound pickup device, wherein the multi-dimensional matching score includes one or more of the following: a conference scene identification matching score, a sound pickup device distance matching score, a conference audio background noise matching score, and a participant matching score; based on the multi-dimensional matching score, the conference audio matching the target conference scene is obtained.

[0034] Optionally, the sound pickup device distance matching score is determined based on propagation properties of wireless communication signals received by the sound pickup device, wherein the wireless communication signals include but are not limited to one or more of the following: WIFI (Wireless Fidelity) signals and Bluetooth signals.

[0035] Among them, the conference scene identification matching score is used to indicate whether the sound pickup devices are configured to access the same conference scene; the sound pickup device distance matching score is used to indicate the distance between the sound pickup devices, and is negatively correlated with the distance between the sound pickup devices; the conference audio background noise matching score is used to indicate the similarity of the background noise in the conference audio collected by the sound pickup device, and the conference audio background noise matching score is positively correlated with the similarity of the background noise in the conference audio; the participant matching score is used to indicate whether the conference audio collected by the sound pickup device contains the same participant.

[0036] The following is a specific example to illustrate how to obtain each multi-dimensional matching score.

[0037] (1) Meeting scene identification matching score

[0038] In some optional embodiments, participants can use the conference application to scan a QR code or click a link to actively join the conference. The QR code contains the conference scene identifier and related information. After the participant scans the QR code, the conference application is triggered to execute the conference joining operation. In other optional embodiments, the participant can manually enter the conference scene identifier and related information to trigger the conference application to execute the conference joining operation.

[0039] After the conference application detects the operation of joining the conference, it automatically verifies the validity of the conference and assigns a unique identity to the participant after verification. After the conference application activates the audio pickup device, the unique identity will be used as the conference scene identifier corresponding to the activated audio pickup device.

[0040] After the participants join the meeting, the conference application will upload the conference audio collected by the pickup device started by the application to the server in real time. The server associates and stores the conference audio and the conference scene identifier. In the conference voice recognition stage, by comparing the conference scene identifier stored in association with the conference audio, it can be determined whether the conference audio matches the same conference scene. For example, when the conference scene identifiers associated with two conference audios are the same, it can be preliminarily determined that the two conference audios belong to the same conference scene. In this case, the conference scene identifier matching scores of the two conference audios can be set to the first score; when the conference scene identifiers associated with two conference audios are different, it can be preliminarily determined that the two conference audios belong to different conference scenes. In this case, the conference scene identifier matching scores of the two conference audios can be set to the second score, wherein the first score is greater than the second score. In the embodiment of the present application, the specific values of the first score and the second score can be set according to specific needs.

[0041] (2) Pickup device distance matching score

[0042] In an embodiment of the present application, the sound pickup device distance matching score includes: a first distance matching score of the sound pickup device and / or a second distance matching score of the sound pickup device.

[0043] Wi-Fi signal strength can be used to determine the relative positions of multiple devices. Devices that are close together typically indicate they are in the same conference room. In some optional embodiments, the sound pickup device distance matching score can be obtained by: using Wi-Fi positioning technology to determine the distance of the sound pickup device relative to the same Wi-Fi access point; and obtaining a first sound pickup device distance matching score for the sound pickup device based on the distance difference between the sound pickup device and the same Wi-Fi access point.

[0044] In specific implementations, if the sound pickup device has WiFi signal transceiver capabilities, after each device connects to the conference room's WiFi network, it calls the device's network card driver interface to obtain the current RSSI (Received Signal Strength Indicator) value. The RSSI value reflects the received signal strength of the device, typically expressed in dBm. The RSSI value is related to the distance between the device and the wireless access point.

[0045] In some optional embodiments, a path loss model can be used to estimate the distance between the sound pickup device and the access point. Then, based on the difference in distance between each sound pickup device and the same access point, the first distance matching score between the sound pickup devices is calculated. Optionally, the path loss model can be expressed as:

[0046] PL(d)=PL0+10·n·log 10 (d) = transmit power - RSSI;

[0047] Among them, transmit power represents the transmit power of a specified Wi-Fi access point, RSSI represents the signal strength indicator value of the specified access point received by a specified pickup device, PL(d) represents the path loss at a distance d, PL0 is the path loss at a reference distance d0 (usually measured at 1 meter), n is the path loss factor, which depends on environmental conditions and is usually between 2 and 4 in indoor environments, and d is the distance from the pickup device to the access point.

[0048] Because wireless signals attenuate with distance as they propagate through the air, this attenuation pattern can be approximated by an inverse square law: signal strength is inversely proportional to the square of the distance. The above formula uses a logarithmic form because wireless signal strength often changes exponentially. Using logarithms can linearize exponential relationships, facilitating calculation and analysis.

[0049] Based on the above formula, after knowing the transmission power of the access point and the RSSI value of the WiFi signal received by the sound pickup device, the distance d between the sound pickup device and the wireless access point can be solved through inverse operation. Furthermore, by comparing the distances between different sound pickup devices and the wireless access point, it can be determined whether the sound pickup devices are located in the same physical space (for example, the same conference room). For example, first solve the distance d of all sound pickup devices from the specified access point i , respectively calculate the distance difference between each two sound pickup devices and the designated wireless access point.

[0050] In some embodiments of the present application, a distance difference threshold can also be pre-set based on the size of the conference room to indicate the maximum allowable distance difference between the pickup devices in the same conference room. If the distance difference between two pickup devices and the designated wireless access point is greater than the distance difference threshold, it can be considered that the probability that the two pickup devices are in the same conference room is greater, that is, the probability that the conference audio collected by the two pickup devices belongs to different conference scenes is greater. Conversely, if the distance difference between two pickup devices and the designated wireless access point is less than or equal to the distance difference threshold, it can be considered that the probability that the two pickup devices are in the same conference room is smaller, that is, the probability that the conference audio collected by the two pickup devices belongs to the same conference scene is smaller.

[0051] In some optional embodiments, the distance difference between the two sound pickup devices and the designated wireless access point can be converted to a first distance matching score for the sound pickup devices. For example, the first distance matching score for the sound pickup devices corresponding to the distance difference being equal to the distance difference threshold can be set to be equal to zero, and the first distance matching score for the sound pickup devices corresponding to the distance difference being greater than the distance difference threshold can be set to be less than zero. Based on this principle, a negative correlation function between the distance difference and the first distance matching score for the sound pickup devices is established, and the distance difference is converted to the first distance matching score for the sound pickup devices through the established negative correlation function. In the embodiments of the present application, there is no limitation on the specific conversion scheme for converting the distance difference into the first distance matching score for the sound pickup devices.

[0052] In other optional embodiments, the first distance matching score of the sound pickup device can be directly obtained based on the RSSI value of the designated access point received by the sound pickup device. For example, the sound pickup device whose difference between the RSSI values received from the same wireless access point is less than a preset strength difference threshold can be considered to have a greater probability of belonging to the same conference room, while the sound pickup device whose difference between the RSSI values received from the same wireless access point is greater than or equal to the preset strength difference threshold can be considered to have a lower probability of belonging to the same conference room. Based on this, a negative correlation function between the RSSI value difference and the first distance matching score of the sound pickup device is established, and the RSSI value difference is converted into the first distance matching score of the sound pickup device through the established negative correlation function. In the embodiments of the present application, there is no limitation on the specific conversion scheme for converting the RSSI value difference into the first distance matching score of the sound pickup device.

[0053] In other embodiments, for scenarios where the sound pickup device can receive Bluetooth signals, Bluetooth beacons can be deployed in the conference room, and the second distance matching score of the sound pickup device is obtained by the following method: using Bluetooth positioning technology to determine the position of each sound pickup device; and obtaining the second distance matching score of the sound pickup device between the sound pickup devices based on the position of the sound pickup device and the position of the Bluetooth beacon.

[0054] For example, first obtain the RSSI value of each Bluetooth beacon scanned by the sound pickup device; then, based on the RSSI value and the deployment location of the corresponding Bluetooth beacon, use the triangulation method to determine the position coordinates of the sound pickup device relative to the Bluetooth beacon; then, calculate the distance between the sound pickup devices based on the position of the sound pickup device relative to the Bluetooth beacon, and determine whether the sound pickup devices are located in the same conference room based on the distance between the sound pickup devices, that is, determine whether the conference audio collected by the sound pickup devices matches the same conference scene based on the distance between the sound pickup devices. For example, a distance difference threshold can be preset. If the distance between the two sound pickup devices is less than the distance difference threshold, it can be considered that the two sound pickup devices are located in the same conference room. Conversely, if the distance between the two sound pickup devices is greater than or equal to the distance difference threshold, it can be considered that the two sound pickup devices are located in different conference rooms, and the judgment result is converted into an example matching score.

[0055] In other optional embodiments, a conversion function can be defined to convert the distance between the pickup devices into a matching distance score, such that the distance between the pickup devices is negatively correlated with the matching distance score. The conversion function can be defined based on Bluetooth signal transmission characteristics or the requirements of a specific conference scenario. The specific form of the conversion function is not limited in the embodiments of this application.

[0056] The specific implementation methods of the above WIFI positioning technology and Bluetooth positioning technology can be found in the prior art and will not be repeated in the embodiments of this application.

[0057] Optionally, the sound pickup device distance matching score may be converted into a discrete value based on the judgment result or may be converted into a continuous value through a conversion function. The specific form of the sound pickup device distance matching score is not limited in the embodiments of the present application.

[0058] (3) Meeting audio background noise matching score

[0059] Optionally, the conference audio background noise matching score is obtained by analyzing the similarity of spectral features of the audio.

[0060] Optionally, obtaining the conference audio background noise matching score by analyzing the similarity of the audio spectrum features includes the following steps: dividing the conference audio collected by the pickup device into several frames of audio signals; calculating the Mel-frequency cepstral coefficient features of each frame of audio signal respectively, and using splicing or other fusion methods based on the Mel-frequency cepstral coefficient features of the single-frame audio signal in the conference audio collected by each pickup device to obtain the Mel-frequency cepstral coefficient features of each conference audio; calculating the cosine similarity between the Mel-frequency cepstral coefficient features of different conference audios as the conference audio background noise matching score of the corresponding conference audio. Among them, the method for obtaining the Mel-frequency cepstral coefficient features of the conference audio can be referred to the prior art and will not be repeated in the embodiments of this application.

[0061] In some optional embodiments, a conference audio background noise matching score threshold can be pre-set based on experience. If the conference audio background noise matching scores of two conference audios exceed the conference audio background noise matching score threshold, that is, the cosine similarity between the Mel-frequency cepstral coefficient features of the two conference audios exceeds a certain threshold, then the two conference audios are considered to match the same conference scene, that is, the pickup devices that collected the two conference audios are in the same conference scene. Conversely, if the conference audio background noise matching scores of the two conference audios are lower than the conference audio background noise matching score threshold, then the two conference audios can be considered to match different conference scenes, that is, the pickup devices that collected the two conference audios are in different conference scenes.

[0062] Optionally, the conference audio background noise matching score can be converted into a discrete value based on the judgment result, or it can be converted into a continuous value based on the cosine similarity between the Mel-frequency cepstral coefficient features. The specific form of expression of the conference audio background noise matching score is not limited in the embodiments of the present application.

[0063] (IV) Participant Matching Score

[0064] Optionally, the participant matching score can be obtained by performing voiceprint recognition on the conference audio based on a preset voiceprint library to obtain the identity information of the participant whose voiceprint features match those in each conference audio; and obtaining the participant matching scores for different conference audios based on the participant identity information. The preset voiceprint library is generated by pre-registering the voiceprints of the participants and stores the voiceprint information of different participants.

[0065] Optionally, before the meeting begins, participants need to complete the pre-registration of their role identities, including but not limited to: registering voiceprint features, registering identity information (such as name and / or personnel number, etc.). After the meeting, for each conference audio, the conference audio can be divided into multiple audio segments, and the voiceprint features of each audio segment are extracted separately. Afterwards, the extracted voiceprint features are matched with the pre-registered voiceprint features in the voiceprint library, and the identity information of the participants corresponding to the successfully matched voiceprint features is obtained as the identity information of the participants matched by the corresponding audio segment. Then, based on the identity information of the participants matched by each audio segment, the identity information of the participants matched by each conference audio, that is, the number of participants, is obtained. Finally, based on the identity information of the participants matched by different conference audios, the target conference audio that matches the same participant and the number of participants that the target conference audio matches can be determined, and the participant matching score of the target conference audio can be determined based on the number of participants that match together. For example, the participant matching score of the two conference audios can be determined based on the proportion of the number of participants that match together in the two conference audios. For example, for conference audio A and conference audio B with overlapping meeting time periods, after voiceprint feature matching, it is determined that conference audio A includes participants a, b and c, and conference audio B includes participants a and c. Then it can be considered that the participant matching score of conference audio A and conference audio B is equal to 2 / 3.

[0066] After obtaining the above-mentioned multiple matching scores, further, conference audio matching the same conference scene can be determined based on one matching score or multiple matching scores.

[0067] In some optional embodiments, when a conference scene identifier is set for a conference, obtaining the conference audio that matches the target conference scene based on the multi-dimensional matching score includes: determining the conference audio that matches the same conference scene according to the conference scene identifier matching score. For example, if the conference scene identifier matching score indicates that two conference audios match (such as the two conference audios match the same conference scene identifier), then the two conference audios can be considered to match the same conference scene; if the conference scene identifier matching score indicates that two conference audios do not match (such as the two conference audios match different conference scene identifiers), then the two conference audios can be considered to match different conference scenes.

[0068] In some optional embodiments, when a conference scene identifier is not set, obtaining the conference audio that matches the target conference scene based on the multi-dimensional matching score includes: determining conference audio that matches the same conference scene based on the participant matching score. For example, if the participant matching score indicates that two conference audios match (e.g., the participants of the two conference audios are exactly the same), then the two conference audios can be considered to match the same conference scene; if the participant matching score indicates that two conference audios do not match (e.g., the participants of the two conference audios are completely different), then the two conference audios can be considered to match different conference scenes.

[0069] In some optional embodiments, when multiple sound pickup devices use different methods to collect conference audio, some conference audio matches with conference scene identifiers, while some conference audio does not have matching conference scene identifiers. Accordingly, multiple conference audios that match the target conference scene are obtained based on the multi-dimensional matching scores, including: when the target conference audio is associated with the conference scene identifier, determining whether the target conference audio matches the same conference scene based on the conference scene identifier matching score; when the target conference audio is not associated with the conference scene identifier, determining the conference audio that matches the same conference scene based on the participant matching score. Wherein, the target conference audio is any two of the conference audios.

[0070] In some optional embodiments, obtaining the conference audio that matches the target conference scene based on the multi-dimensional matching scores includes: performing a weighted summation of one or more scores of the target conference audio's conference scene identification matching score, the participant matching score, the first distance matching score of the sound pickup device, the second distance matching score of the sound pickup device, and the conference audio background noise matching score to obtain a comprehensive matching score for the target conference audio; and determining whether the target conference audio matches the target conference scene based on the comprehensive matching score. The target conference audio is selected from any two conference audios from the conference audio, and the target conference scene refers to the same conference scene. When performing the weighted summation, the weights corresponding to the conference scene identification matching score, the participant matching score, the first distance matching score of the sound pickup device, the second distance matching score of the sound pickup device, and the conference audio background noise matching score are determined based on the reliability of the specific scores. For example, the weight corresponding to the conference scene identification matching score can be set to be greater than the weight corresponding to the participant matching score, and the weights corresponding to the conference scene identification matching score and the participant matching score can be set to be greater than the weights corresponding to the other scores.

[0071] Through the above operations, based on multi-dimensional signal feature analysis, it can be ensured that the sound source information collected by multiple pickup devices belongs to the same conference scene, effectively improving the accuracy and effectiveness of collaborative collection of multiple pickup devices, and avoiding the interference of mixed data from different scenes on the recognition results.

[0072] After obtaining all the conference audios for the same conference scene, for each conference scene, all the conference audios that match the conference scene are segmented and screened, high-quality audio segments at each moment are screened from different conference audios, and the screened high-quality audio segments are integrated into the complete audio for speech recognition of the conference scene.

[0073] Step 120 : performing segmented screening and multi-device splicing processing on the conference audio matching the target conference scene to obtain the spliced audio of the target conference scene.

[0074] Due to differences in acquisition equipment and environments, the quality of conference audio signals collected from multiple pickup devices (such as mobile phones, microphones, far-field devices, etc.) also varies. It is necessary to perform a quality score on the conference audio collected by different pickup devices, and prioritize high-quality signals as input for speech recognition. In the embodiments of this application, the quality of the audio segments is scored by segmenting different conference audios. Afterwards, high-quality audio ends are selected from different conference audios and fused to obtain the complete conference audio.

[0075] In some optional embodiments, each of the conference audios matching the target conference scene is separately segmented and aligned to obtain several audio segments corresponding to each of the conference audios, and before the timestamp corresponding to each of the audio segments, it also includes: using pre-emphasis processing technology to pre-process each of the conference audios matching the target conference scene to enhance the high-frequency signals in the conference audio.

[0076] The quality of speech signals has a direct impact on semantic communication, sound recognition, and auditory perception. Loss of useful information and interference from redundant information can hinder communication and speech signal processing. Therefore, preprocessing conference audio data is essential. Most of the energy in a speech signal is concentrated in the low-frequency portion, which reduces the signal-to-noise ratio in the high-frequency portion. This requires filtering the signal with a first-order high-pass filter, a technique known as pre-emphasis. Pre-emphasis compensates for the suppressed high-frequency portion of the speech signal, flattens the signal spectrum, and removes spectral tilt. The pre-emphasis formula is as follows: y(t) = x(t) - α · x(t-1), where x(t) is the input speech signal (i.e., the conference audio), y(t) is the pre-emphasized signal, and α is the pre-emphasis coefficient, typically ranging from 0.9 to 1.0. This step enhances the high-frequency signal by removing the low-frequency portion.

[0077] Optionally, the segmented screening and multi-device splicing processing of the conference audio matching the target conference scene based on audio quality to obtain the spliced audio of the target conference scene includes: sub-step 1201 and sub-step 1202.

[0078] Sub-step 1201 : performing segment alignment processing on each of the conference audios matching the target conference scene, and obtaining a plurality of audio segments corresponding to each of the conference audios, and a timestamp corresponding to each of the audio segments.

[0079] In some embodiments of the present application, each of the conference audios matching the target conference scene is subjected to segmented alignment processing to obtain a number of audio segments corresponding to each of the conference audios, as well as a timestamp corresponding to each of the audio segments, including: performing frame alignment processing on each of the conference audios matching the target conference scene based on a system time alignment method to obtain a number of audio segments corresponding to each of the conference audios, as well as a timestamp corresponding to each of the audio segments.

[0080] When the local system time of the pickup device is synchronized, the conference audio can be segmented and aligned separately using a system time alignment method. Since the timestamps of the conference audio collected by each pickup device may be slightly different, the conference audio collected by each pickup device needs to be divided into small frames, i.e., audio segments. The length of each audio segment is determined according to the sampling rate of the audio signal. For example, conference audio of a fixed length is sequentially intercepted from the conference audio and normalized. Taking the 10-minute conference audio X1 collected by pickup device A and the conference audio X2 collected by pickup device B as an example, the audio is intercepted every 1 minute and the conference audio X1 and X2 are divided into 10 segments respectively. In this way, the timeline is unified, i.e., the standardized timeline.

[0081] Furthermore, a timestamp corresponding to each audio segment may be determined based on the system time corresponding to the conference audio and the duration corresponding to the audio segment.

[0082] Preferably, the multiple audio segments corresponding to the same conference audio have overlapping intervals. In the case of partial overlap between audio segments, signal continuity can be maintained. The conference audio signal collected by each pickup device is segmented to standardize the timeline, ensuring that conference audio collected by multiple pickup devices can be analyzed synchronously.

[0083] In some embodiments of the present application, each of the conference audios matching the target conference scene is subjected to segmented alignment processing to obtain a number of audio segments corresponding to each of the conference audios, as well as a timestamp corresponding to each of the audio segments, including: based on a dynamic time warping alignment algorithm, each of the conference audios matching the target conference scene is subjected to frame alignment processing to obtain a number of audio segments corresponding to each of the conference audios, as well as a timestamp corresponding to each of the audio segments.

[0084] When the local system time of the pickup device is not synchronized, frame alignment cannot be achieved directly through the system time. Instead, a signal alignment method based on dynamic time warping (DTW) can be used to achieve alignment between different time axes. DTW is mainly used to address the problem of aligning time series with mismatched lengths. DTW allows for flexible matching between frames of different speech signals, meaning that it does not require each frame to be strictly aligned, but rather dynamically adjusts the matching path to achieve optimal alignment of the overall sequence. DTW can be used to align input speech and reference speech templates to improve recognition accuracy.

[0085] The specific implementation scheme of the signal alignment method based on dynamic time warping is as follows: First, it is necessary to ensure that at least one pickup device fully participates in the entire process of conference voice recording, and the time axis of the conference audio collected by the pickup device is used as the reference time axis. For other devices that may not have fully participated in the meeting, the corresponding position of their audio segments on the reference time axis is determined by calculating the signal similarity. During the specific implementation process, the system segments the conference audio collected by each pickup device. In order to maintain signal continuity, segmentation is performed using segment overlap. After the segmentation is completed, the similarity between the conference audio collected by each pickup device and the conference audio corresponding to the reference time axis is calculated based on the dynamic time warping algorithm, thereby determining the corresponding position of the audio segments in the conference audio collected by each pickup device on the reference time axis, and finally achieving precise alignment of the time axes of multiple pickup devices.

[0086] Segment alignment involves segmentation at fixed time intervals, while DTW alignment is an alignment method based on dynamic similarity adjustments. Segment alignment alone can lead to time step mismatches. DTW, however, utilizes nonlinear time warping to make audio segment alignment more robust, making it particularly suitable for comparing, matching, and aligning speech signals.

[0087] For the specific implementation of segmenting and processing conference audio using the dynamic time warping method, please refer to the prior art and will not be repeated in the embodiments of this application.

[0088] After segment alignment, the conference audio collected by each pickup device will be divided into multiple audio segments. For example, the conference audio collected by pickup device 1 can be represented as The conference audio collected by the sound pickup device 2 can be expressed as in, and Represents the audio segments before alignment. The dynamic time warping (DTW) algorithm is used to find the optimal alignment path, which includes the audio segments and their corresponding timestamps. The aligned frame sequences are X1′ and X2′.

[0089] Based on the dynamic time warping method, the similarity between different conference audios is measured and the timing of the frames is adjusted to ensure the consistency of the conference audio on the timeline.

[0090] Sub-step 1202 : performing multi-pickup device splicing processing on the audio segments based on the audio quality and the timestamp to obtain spliced audio of the target conference scene.

[0091] Next, high-quality signals are screened based on the conference audio of the same conference scene collected by multiple sound pickup devices.

[0092] Optionally, before performing multi-pickup device splicing processing on the retained audio segment based on the timestamp to obtain the spliced audio of the target conference scene, the method further includes: post-processing the retained audio segment, wherein the post-processing includes one or more of the following processing methods: signal compensation processing on the audio segment, spectrum analysis and denoising processing on the audio segment, and voice enhancement processing on the audio segment.

[0093] The signal compensation processing performed on the audio segments is used to compensate for signal loss or omission that may occur during the segment alignment process by using an interpolation algorithm or a prediction model to compensate for lost signal frames, thereby maintaining the continuity and integrity of the audio as much as possible.

[0094] The spectrum analysis and denoising processing of the audio segment includes: using short-time Fourier transform to perform frequency domain conversion on the segmented audio signal to detect interference factors such as background noise and static noise in the signal; using noise suppression techniques such as adaptive filtering or spectrum subtraction to reduce background noise in the signal while retaining the main frequency band characteristics of the voice signal to improve voice clarity.

[0095] The speech enhancement processing of the audio segment includes: introducing a deep learning model for speech enhancement, such as a speech enhancement model that combines a convolutional neural network and a recurrent neural network architecture, extracting spatiotemporal features in the speech signal, enhancing the main frequency components in the speech, and reducing the impact of environmental noise.

[0096] Next, based on indicators such as the signal quality and background noise level of different pickup devices, the conference audio collected by multiple pickup devices is weighted and fused to generate the final high-quality conference audio.

[0097] In some optional embodiments, based on the audio quality and the timestamps, the audio segments are spliced together using multiple pickup devices to obtain the spliced audio of the target conference scene, including: filtering the audio segments based on the audio quality to obtain the retained audio segments corresponding to each of the timestamps; and splicing the retained audio segments using multiple pickup devices based on the timestamps to obtain the spliced audio of the target conference scene. That is, first, the audio segments of the conference audio collected by different pickup devices are scored based on the audio quality. Then, based on the quality scoring results, high-quality audio segments are selected from the audio segments collected by different pickup devices, and the selected high-quality audio segments are spliced together to generate high-quality complete conference audio for speech recognition.

[0098] Optionally, the audio segments are screened based on audio quality to obtain retained audio segments corresponding to each of the timestamps, including: based on preset audio indicators of each of the audio segments, respectively obtaining first quality features of each of the audio segments, wherein the preset audio indicators include one or more of the following: signal-to-noise ratio, clarity index and Mel-cephalogram distortion; performing a reference-free quality assessment on the audio segment sequence of each of the conference audios through a pre-trained neural network model to obtain a second quality feature of each of the audio segments; using the first quality feature and the second quality feature as input of a pre-trained multi-device fusion model, and obtaining a comprehensive quality score of each of the audio segments through the multi-device fusion model; screening the audio segments in each of the conference audios based on the comprehensive quality score to obtain retained audio segments corresponding to each of the timestamps.

[0099] The first quality feature of the audio segment can be obtained by weighting one or more audio indicators, such as the signal-to-noise ratio, clarity index, and Mel-cepstrum distortion of the audio segment. Methods for obtaining audio indicators, such as the signal-to-noise ratio, clarity index, and Mel-cepstrum distortion of the audio segment, are described in the prior art and will not be further described in the embodiments of this application.

[0100] Optionally, a reference-free quality assessment is performed on each audio segment sequence of the conference audio using a pre-trained neural network model to obtain a second quality feature of each audio segment, including: using each audio segment sequence of the conference audio as input to the pre-trained neural network model, so that the neural network model performs a quality assessment on each audio segment sequence of the conference audio based on the local time-frequency features and long-term dependency features of each audio segment in the input to obtain a second quality feature of each audio segment.

[0101] The neural network model can adopt the CBLSTM model (a neural network model that combines a convolutional neural network with a bidirectional long short-term memory network). The CBLSTM model combines a convolutional neural network and a bidirectional long short-term memory network to extract the time-frequency features of speech signals in conference audio and model time series information. The model structure of the CBLSTM model includes a convolutional layer for extracting local time-frequency features, a bidirectional long short-term memory network for modeling the long-term dependency features of the sequence, an attention module for enhancing the weights of important features, and an output layer for the final predicted quality score.

[0102] Optionally, the loss function of the neural network model is configured to be based on a signal-to-loss ratio. For example, the loss function can be defined as: Where S represents the reference signal and D represents the degraded signal. During model training, the network parameters are adjusted by measuring the signal loss between the reference and degraded signals. This loss function effectively measures the loss ratio of speech signals in conference audio, thereby optimizing the training process of the neural network model.

[0103] The network structure of the CBLSTM model is referred to in the prior art and will not be described in detail in the present embodiment. The training method of the CBLSTM model is referred to in the prior art and will not be described in detail in the present embodiment.

[0104] Next, for each conference audio, a set of feature vectors is assembled based on the first and second quality features of each audio segment in the conference audio. Furthermore, for each conference audio, the feature vectors of each audio segment are assembled into a sequence of feature vectors according to the timestamp corresponding to the audio segment, which serves as the feature vector of the corresponding conference audio. The feature vectors of the conference audio collected by each pickup device are then used as input to the multi-device fusion model, which is used to obtain a comprehensive quality score for each audio segment in each conference audio.

[0105] Optionally, the multi-device fusion model is implemented based on the Transformer framework. Its core is to integrate voice signals from different pickup devices through the self-attention mechanism to obtain a fused signal, thereby effectively enhancing the quality and integrity of the signal. The multi-device fusion model calculates the attention weight of the audio segment of each pickup device through a multi-head self-attention module, fuses the audio information from different pickup devices, and improves the accuracy of the fusion result. Afterwards, based on the fused signal output by the Transformer, the final signal quality, that is, the comprehensive quality score of each audio segment in each conference audio, can be further predicted through a fully connected layer or other task-specific layers.

[0106] The network structure and training method of the multi-device fusion model refer to the network structure and training method of the fusion model in the prior art, and will not be repeated in the embodiments of this application.

[0107] After obtaining the comprehensive quality scores of the audio segments in each conference audio, the audio segment with the best signal quality corresponding to each timestamp is selected as the retained audio segment corresponding to the corresponding timestamp according to the comprehensive quality scores.

[0108] Next, a multi-pickup device splicing process is performed based on each retained audio segment to obtain the spliced audio of the target conference scene.

[0109] In some optional embodiments, performing multi-pickup device splicing processing on the reserved audio segments based on the timestamps to obtain the spliced audio of the target meeting scenario includes: splicing the reserved audio segments in the order of the timestamps to obtain the spliced audio of the target meeting scenario. For example, after screening out the reserved audio segments corresponding to each timestamp, the reserved audio segments corresponding to each timestamp (i.e., the optimal audio signals for each time period) are spliced from front to back in the order of the timestamps from earliest to latest to obtain a complete and high-quality meeting audio.

[0110] Optionally, during the splicing process, to ensure the continuity of the audio, weighted smoothing is used for transition processing to reduce the abruptness caused by the differences between the signals of different pickup devices. The specific implementation of weighted smoothing processing for audio signals can be found in the prior art and will not be elaborated in this embodiment of the present application.

[0111] In some other optional embodiments, there is an overlapping interval between the audio segments. Performing multi-pickup device splicing processing on the reserved audio segments based on the timestamps to obtain the spliced audio of the target meeting scenario includes: obtaining the overlapping interval audio of adjacent reserved audio segments; performing weighted smoothing transition processing on the overlapping interval audio to obtain the superimposed audio of adjacent reserved audio segments; splicing the non-overlapping interval audio and the superimposed audio of adjacent reserved audio segments in the order of the timestamps to obtain the spliced audio of the target meeting scenario.

[0112] Specifically, during implementation, the overlapping interval can be determined according to the sampling results of each audio segment. For example, audio segment S t ends at timestamp t_end, and audio segment S t+1 starts at t_start. If t_start < t_end, the overlapping interval audio of adjacent audio segments S t and S t+1 is the meeting audio from t_start to t_end.

[0113] When performing weighted smoothing transition processing on the overlapping interval audio, fusion weights can be assigned to the audio signals in the overlapping region according to the time position, and the audio signals in the overlapping region are weighted and summed according to the assigned fusion weights to obtain the superimposed audio corresponding to the overlapping region.

[0114] Taking the superimposed region determined above as an example, the fusion weight W t of the audio signal at the time point in the overlapping region of audio segment S t is calculated by the formula W t = 1 - (t - t_start) / overlap_duration, and audio segment S t+1The fusion weight W of the audio signal at a given time point in the overlapping area t+1 By formula W t+1 =(t-t_start) / overlap_duration, where overlap_duration represents the duration of the overlapping region and t represents the current time of the overlapping region. Next, the weights calculated for each adjacent audio segment are used to calculate the weights using the formula W t ×S t +W t+1 ×S t+1 The audio signals in the overlapping areas are weightedly fused to obtain the superimposed audio in the overlapping areas.

[0115] Finally, the previous audio segment S t The audio signal before the overlapping area, the superimposed audio, and the subsequent audio segment S t+1 The audio signals after the overlapping area are spliced in sequence to obtain the spliced audio at time t. The spliced audio at all conference moments constitutes the complete audio of the target conference scene.

[0116] By using this method, the fusion weights change linearly over time. When fusing adjacent audio segments, the audio signal of the previous segment is fully used at the beginning of the overlapping area. As time passes, the audio signal of the subsequent segment is gradually adopted, ensuring a smooth and continuous signal transition. This weighted smooth transition method effectively reduces signal abrupt changes during splicing and produces high-quality audio signals.

[0117] In the embodiments of this application, by adopting a timestamp-based segmented optimal signal selection and splicing method, a high-quality voice signal containing the entire conference content can be generated. By using multiple signal quality scoring methods based on signal quality, the accuracy of signal quality judgment is improved, thereby improving the quality of the spliced audio.

[0118] Step 130 : Perform multimodal information fusion on the pre-collected visual information of the target conference scene and the spliced audio to obtain multimodal fusion information.

[0119] In addition to voice information, conference scenarios can also include slides or other document information displayed during the meeting. Existing technologies often ignore this visual information, resulting in insufficient feature information, which in turn affects speech recognition accuracy. In some embodiments of the present application, a multi-dimensional information fusion method is used to enhance voice feature information and improve speech recognition accuracy. This not only considers voice information but also incorporates image information such as slides and documents displayed during the meeting, enriching the feature information and enhancing speech recognition effectiveness.

[0120] In some optional embodiments, the visual information includes: a presentation document video, and the pre-collected visual information of the target conference scene and the spliced audio are subjected to multimodal information fusion to obtain multimodal fusion information, including: synchronizing the pre-collected presentation document video and the spliced audio of the target conference scene to obtain a synchronized conference video; performing image recognition on the synchronized conference video to obtain a conference text; and associating the conference text and the spliced audio according to the conference time to obtain multimodal fusion information at each time point in the target conference scene.

[0121] During a meeting, devices with video capabilities (such as mobile phones and webcams) are automatically or manually activated to continuously capture video of the meeting scene as visual information. For example, a video camera can capture video of a slide show or document being played in a meeting. Each image frame in the video is timestamped. Because each device may have a different time offset, the time of the audio and video capture devices in the same meeting scene needs to be synchronized, mapping the times of different devices to a main timeline.

[0122] For example, the conference information collection time of the device with the longest conference recording time can be used as the main timeline. Then, the audio and video synchronization method in the existing technology is used to synchronize the conference video to the time space of the spliced audio to obtain a synchronized conference video.

[0123] For the method of synchronizing conference video with conference audio, please refer to the prior art and will not be described in detail in the embodiments of this application.

[0124] Then, the synchronous conference video is preprocessed to extract key information and enhance the background information of the voice features. In some optional embodiments, for example, image processing technology such as OCR (optical character recognition) technology is used to identify the text content displayed in the slides or documents in each video image frame of the synchronous conference video, and then the consistency of the text content in adjacent video image frames is judged to obtain video image frames corresponding to the same text content, and the video time period corresponding to the same text content is determined based on the timestamps of all video image frames corresponding to the same text content. Furthermore, the spliced audio corresponding to each video time period is obtained, and the text content and spliced audio corresponding to each video time period are associated. According to this method, the text content in each frame of the slide or each page of the document played during the meeting will be associated with a piece of spliced audio, respectively, to obtain multimodal information fusion of multiple consecutive time points during the meeting. These text contents can help understand the keywords and contextual information in the conference audio.

[0125] In some optional embodiments, facial image information of the participants can also be collected, and facial recognition technology can be used to identify the identities of the participants, so as to match the voiceprint information with the speaker, thereby enhancing the adaptability of the speech recognition model under different speaker conditions.

[0126] In some optional embodiments, it is also possible to set pictures and voice annotations for the meeting minutes of the target meeting scene based on the timestamps corresponding to the images of slides or documents in the collected meeting video and the audio clips of the spliced audio, and finally generate meeting minutes with pictures and voice annotations, providing multi-dimensional information for subsequent retrieval and archiving.

[0127] Step 140: Perform speech recognition based on the multimodal fusion information to obtain a conference speech recognition result of the target conference scene.

[0128] In some optional embodiments, the speech recognition is performed based on the multimodal fusion information to obtain the conference speech recognition result of the target conference scene, including: performing feature extraction on the conference text included in the multimodal fusion information through the text processing branch of a pre-trained multimodal fusion model to obtain text features; performing feature extraction on the spliced audio included in the multimodal fusion information through the audio processing branch of the multimodal fusion model to obtain audio features; using the text features and the audio features as multimodal inputs of the feature fusion branch of the multimodal fusion model to obtain fusion features output by the multimodal fusion model; and performing speech recognition based on the fusion features to obtain the conference speech recognition result of the target conference scene.

[0129] Among them, the pre-trained multimodal fusion model can use a multimodal deep learning model based on the Transformer framework to splice audio and text information. It combines the pre-trained Wav2Vec 2.0 model to build an audio processing branch to extract audio features. It combines the BERT model to construct a text processing branch to process text features. Feature fusion is achieved through a shared multimodal Transformer layer (i.e., the feature fusion branch), and the cross-modal attention mechanism is used to enhance the interaction and association between audio and text information. During the training process of the multimodal fusion model, the multi-task learning ability of the model is optimized through joint training, thereby improving the comprehensive understanding of audio and text information, and ultimately obtaining rich signal features of the spliced audio and text information.

[0130] Finally, speech recognition is performed based on the fusion features output by the multimodal fusion model, thereby obtaining a conference speech recognition result for the target conference scene.

[0131] In summary, the conference speech recognition method disclosed in the embodiment of the present application judges the consistency of the conference scene of the conference audio collected by multiple sound pickup devices to obtain the conference audio matching the target conference scene in the conference audio; performs segmented screening and multi-device splicing processing on the conference audio matching the target conference scene to obtain the spliced audio of the target conference scene; performs multimodal information fusion on the pre-collected visual information of the target conference scene and the spliced audio to obtain multimodal fusion information; performs speech recognition based on the multimodal fusion information to obtain the conference speech recognition result of the target conference scene. This method introduces multiple sound pickup devices to collaboratively collect the conference audio of a single conference scene, making the sound source information more comprehensive; combines the segmented alignment technology to perform quality scoring on the conference audio collected by multiple sound pickup devices, thereby screening high-quality audio segments for audio splicing, effectively utilizing the conference audio collected by multiple sound pickup devices, and ensuring that the final result is a high-quality speech signal; introduces visual information into the audio signal, fuses multimodal information for speech recognition, and comprehensively captures and processes a variety of data, further improving the accuracy and robustness of speech recognition. By combining the above technical means, this method effectively solves the signal attenuation caused by factors such as distance and noise when a single device picks up sound, and effectively improves the accuracy and robustness of conference speech recognition results.

[0132] On the other hand, by using multi-dimensional signal features to judge the consistency of the conference scene, the interference of inconsistent information collected by different sound pickup devices on the speech recognition results is effectively avoided, ensuring the consistency and effectiveness of the conference audio data. In complex scenes with multiple sound sources and multiple noises, the stable performance of conference audio fusion can still be maintained, greatly enhancing the application scenario adaptability of the conference speech recognition method.

[0133] Reference Figure 3 , the embodiment of the present application further discloses a conference speech recognition device, the device comprising:

[0134] A conference scene consistency judgment module 310 is configured to perform conference scene consistency judgment on conference audios collected by multiple sound pickup devices, and obtain conference audios that match a target conference scene from the conference audios;

[0135] The spliced audio acquisition module 320 is used to perform segmented screening and multi-device splicing processing on the conference audio matching the target conference scene to obtain the spliced audio of the target conference scene;

[0136] A multimodal fusion module 330 is configured to perform multimodal information fusion on the pre-collected visual information of the target conference scene and the spliced audio to obtain multimodal fusion information;

[0137] The speech recognition module 340 is configured to perform speech recognition based on the multimodal fusion information to obtain a conference speech recognition result of the target conference scene.

[0138] Optionally, the conference scene consistency judgment module 310 is further configured to:

[0139] Obtaining a multi-dimensional matching score of the sound pickup device, wherein the multi-dimensional matching score includes one or more of the following: a conference scene identification matching score, a sound pickup device distance matching score, a conference audio background noise matching score, and a participant matching score;

[0140] Based on the multi-dimensional matching score, the conference audio that matches the target conference scene is obtained.

[0141] Optionally, the spliced audio acquisition module 320 is further configured to:

[0142] Performing segment alignment processing on each of the conference audios matching the target conference scene to obtain a plurality of audio segments corresponding to each of the conference audios and a timestamp corresponding to each of the audio segments;

[0143] Based on the audio quality and the timestamp, the audio segments are spliced by multiple sound pickup devices to obtain spliced audio of the target conference scene.

[0144] Optionally, performing multi-pickup device splicing processing on the audio segments based on the audio quality and the timestamp to obtain spliced audio of the target conference scene includes:

[0145] Filtering the audio segments based on audio quality to obtain a retained audio segment corresponding to each of the timestamps;

[0146] The retained audio segments are spliced by multiple sound pickup devices based on the timestamps to obtain spliced audio of the target conference scene.

[0147] Optionally, the filtering the audio segments based on audio quality to obtain a retained audio segment corresponding to each of the timestamps includes:

[0148] Based on a preset audio indicator of each of the audio segments, respectively obtain a first quality feature of each of the audio segments, wherein the preset audio indicator includes one or more of the following: signal-to-noise ratio, clarity index, and Mel-cepstrum distortion;

[0149] Performing a no-reference quality assessment on each audio segment sequence of the conference audio using a pre-trained neural network model to obtain a second quality feature of each audio segment;

[0150] Using the first quality feature and the second quality feature as inputs of a pre-trained multi-device fusion model, and obtaining a comprehensive quality score for each of the audio segments through the multi-device fusion model;

[0151] The audio segments in each of the conference audios are screened based on the comprehensive quality score to obtain a retained audio segment corresponding to each of the timestamps.

[0152] Optionally, the visual information includes: a presentation document video, and the multimodal fusion module 330 is further configured to:

[0153] Synchronously processing the pre-collected presentation document video of the target conference scene and the spliced audio to obtain a synchronized conference video;

[0154] Performing image recognition on the synchronous conference video to obtain conference text;

[0155] The conference text and the spliced audio are associated according to the conference time to obtain multimodal fusion information of each time point in the target conference scene.

[0156] The conference speech recognition device disclosed in the embodiment of the present application is used to implement the conference speech recognition method described in the embodiment of the present application. The specific implementation methods of each module of the device will not be repeated here. Please refer to the specific implementation methods of the corresponding steps in the method embodiment.

[0157] The conference speech recognition device disclosed in the embodiment of the present application performs conference scene consistency judgment on the conference audio collected by multiple sound pickup devices to obtain the conference audio matching the target conference scene in the conference audio; performs segmented screening and multi-device splicing processing on the conference audio matching the target conference scene to obtain the spliced audio of the target conference scene; performs multimodal information fusion on the pre-collected visual information of the target conference scene and the spliced audio to obtain multimodal fusion information; performs speech recognition based on the multimodal fusion information to obtain the conference speech recognition result of the target conference scene. This device introduces multiple sound pickup devices to collaboratively collect the conference audio of a single conference scene, making the sound source information more comprehensive; combines the segmented alignment technology to perform quality scoring on the conference audio collected by multiple sound pickup devices, thereby screening high-quality audio segments for audio splicing, effectively utilizing the conference audio collected by multiple sound pickup devices, and ensuring that the final result is a high-quality speech signal; introduces visual information into the audio signal, fuses multimodal information for speech recognition, comprehensively captures and processes multiple data, and further improves the accuracy and robustness of speech recognition. By combining the above technical means, this device effectively solves the signal attenuation caused by factors such as distance and noise when a single device picks up sound, and effectively improves the accuracy and robustness of conference speech recognition results.

[0158] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between the various embodiments can be referred to in conjunction with each other. For the device embodiments, since they are generally similar to the method embodiments, their description is relatively simple, and for relevant parts, reference can be made to the description of the method embodiments.

[0159] The above is a detailed introduction to a conference speech recognition method and device provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

[0160] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0161] The various component embodiments of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It will be appreciated by those skilled in the art that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the electronic device according to the embodiment of the present application. The application can also be implemented as a device or apparatus program (for example, a computer program and a computer program product) for performing a part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0162] For example, Figure 4An electronic device that can implement the method according to the present application is shown. The electronic device can be a PC, a mobile terminal, a personal digital assistant, a tablet computer, etc. The electronic device conventionally includes a processor 410 and a memory 420, and program code 430 stored on the memory 420 and executable on the processor 410. When the processor 410 executes the program code 430, the method described in the above embodiments is implemented. The memory 420 can be a computer program product or a computer-readable medium. The memory 420 can be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. The memory 420 has a storage space 4201 for program code 430 of a computer program for executing any of the method steps described above. For example, the storage space 4201 for program code 430 can include individual computer programs for implementing various steps in the above method. The program code 430 is computer-readable code. These computer programs can be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards, or floppy disks. The computer program includes a computer-readable code, and when the computer-readable code is run on an electronic device, the electronic device is caused to execute the method according to the above embodiment.

[0163] An embodiment of the present application also discloses a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the conference speech recognition method as described in the embodiment of the present application are implemented.

[0164] Such a computer program product may be a computer-readable storage medium having a computer program product. Figure 4 The memory 420 in the electronic device shown is similarly arranged as a storage segment, storage space, etc. The program code can be compressed and stored in the computer readable storage medium in an appropriate form. The computer readable storage medium is generally as shown in FIG. Figure 5 The portable or fixed storage unit generally includes computer-readable code 430', which is a code read by a processor and implements the steps of the above-described method when executed by the processor.

[0165] References herein to "one embodiment," "an embodiment," or "one or more embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. Furthermore, please note that instances of the phrase "in one embodiment" do not necessarily all refer to the same embodiment.

[0166] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.

[0167] In the claims, any reference signs placed between brackets shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.

[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A conference speech recognition method, characterized in that: The method comprises: Performing conference scene consistency judgment on conference audio collected by multiple sound pickup devices, and obtaining conference audio that matches the target conference scene in the conference audio; Performing segmented screening and multi-device splicing processing on the conference audio matching the target conference scene to obtain the spliced audio of the target conference scene; Performing multimodal information fusion on the pre-collected visual information of the target conference scene and the spliced audio to obtain multimodal fusion information; Speech recognition is performed based on the multimodal fusion information to obtain a conference speech recognition result of the target conference scene.

2. The method according to claim 1, characterized in that The performing conference scene consistency judgment on the conference audios collected by the multiple sound pickup devices and obtaining the conference audios matching the target conference scene in the conference audios includes: Obtaining a multi-dimensional matching score of the sound pickup device, wherein the multi-dimensional matching score includes one or more of the following: a conference scene identification matching score, a sound pickup device distance matching score, a conference audio background noise matching score, and a participant matching score; Based on the multi-dimensional matching score, the conference audio that matches the target conference scene is obtained.

3. The method according to claim 1, characterized in that The performing segmented screening and multi-device splicing processing on the conference audio matching the target conference scene to obtain the spliced audio of the target conference scene includes: Performing segment alignment processing on each of the conference audios matching the target conference scene to obtain a plurality of audio segments corresponding to each of the conference audios and a timestamp corresponding to each of the audio segments; Based on the audio quality and the timestamp, the audio segments are spliced by multiple sound pickup devices to obtain spliced audio of the target conference scene.

4. The method according to claim 3, characterized in that The performing multi-pickup device splicing processing on the audio segments based on the audio quality and the timestamp to obtain the spliced audio of the target conference scene includes: Filtering the audio segments based on audio quality to obtain a retained audio segment corresponding to each of the timestamps; The retained audio segments are spliced by multiple sound pickup devices based on the timestamps to obtain spliced audio of the target conference scene.

5. The method according to claim 4, characterized in that The filtering the audio segments based on audio quality to obtain a retained audio segment corresponding to each of the timestamps includes: Based on a preset audio indicator of each of the audio segments, respectively obtain a first quality feature of each of the audio segments, wherein the preset audio indicator includes one or more of the following: signal-to-noise ratio, clarity index, and Mel-cepstrum distortion; Performing a no-reference quality assessment on each audio segment sequence of the conference audio using a pre-trained neural network model to obtain a second quality feature of each audio segment; Using the first quality feature and the second quality feature as inputs of a pre-trained multi-device fusion model, and obtaining a comprehensive quality score for each of the audio segments through the multi-device fusion model; The audio segments in each of the conference audios are screened based on the comprehensive quality score to obtain a retained audio segment corresponding to each of the timestamps.

6. The method according to claim 1, characterized in that The visual information includes: a presentation document video, and the pre-collected visual information of the target conference scene and the spliced audio are subjected to multimodal information fusion to obtain multimodal fusion information, including: Synchronously processing the pre-collected presentation document video of the target conference scene and the spliced audio to obtain a synchronized conference video; Performing image recognition on the synchronous conference video to obtain conference text; The conference text and the spliced audio are associated according to the conference time to obtain multimodal fusion information of each time point in the target conference scene.

7. A conference speech recognition device, characterized in that: The device comprises: A conference scene consistency judgment module is used to judge the conference scene consistency of the conference audio collected by multiple sound pickup devices, and obtain the conference audio that matches the target conference scene in the conference audio; A spliced audio acquisition module is used to perform segmented screening and multi-device splicing processing on the conference audio matching the target conference scene to obtain the spliced audio of the target conference scene; A multimodal fusion module, configured to perform multimodal information fusion on the pre-collected visual information of the target conference scene and the spliced audio to obtain multimodal fusion information; The speech recognition module is used to perform speech recognition based on the multimodal fusion information to obtain a conference speech recognition result of the target conference scene.

8. The device according to claim 7, characterized in that The spliced audio acquisition module is further used to: Performing segment alignment processing on each of the conference audios matching the target conference scene to obtain a plurality of audio segments corresponding to each of the conference audios and a timestamp corresponding to each of the audio segments; Based on the audio quality and the timestamp, the audio segments are spliced by multiple sound pickup devices to obtain spliced audio of the target conference scene.

9. An electronic device comprising a memory, a processor, and a program code stored in the memory and executable on the processor, wherein: When the processor executes the program code, the method according to any one of claims 1 to 6 is implemented.

10. A computer-readable storage medium having program code stored thereon, characterized in that: When the program code is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Meeting hot spot processing method and device, terminal equipment and storage medium

    CN110211590A

  • Conference auxiliary method and device, electronic device and storage medium

    CN112053691A

  • Multi-channel audio data processing method and system

    CN115643242A

  • Multi-modal speech recognition method based on visual scene, electronic equipment and medium

    CN118155624A

  • Conference summary generation method and device based on voice recognition technology, and medium

    CN120015034A

Cited By

  • 3D digital human system integration method and system based on credential environment

    CN120894472A

  • Speech recognition method and system based on AIOT edge device and cloud LLM large model

    CN121789677A