A processing method and a processing device

By comparing the text contents of the first audio and the second audio during the audio processing, the sound quality loss problem is determined and corresponding measures are taken, the problem of sound quality loss in the audio processing is solved and the audio quality is improved.

CN114582362BActive Publication Date: 2025-05-27LENOVO (BEIJING) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210189603.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-05-27
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

There are often sound quality loss problems during audio processing, and it is difficult to determine the specific link of the problem for targeted processing.

Method used

By converting the first audio and the second audio (objected by the first audio) into text, and comparing the similarity of the text contents of the two, it is determined whether to perform an operation of undoing or optimizing the target processing on the first audio, or to issue an audio abnormality prompt.

Benefits of technology

The sound quality loss problem and its links are effectively determined, and the audio quality is improved through corresponding operations or prompts to avoid sound quality loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114582362B_ABST
    Figure CN114582362B_ABST
Patent Text Reader

Abstract

The present application discloses a processing method and a processing device. In response to obtaining a first text converted from a first audio and a second text converted from a second audio (obtained by performing a target process on the first audio), the first text and the second text are subjected to a comparison process, and based on the comparison result of the first text and the second text, it is determined whether to perform a first operation on the first audio and / or issue a first prompt. Through the above text comparison process, the present application can effectively determine the sound quality loss problem and the location of the problem in the process of performing the target process on the first audio, and perform corresponding problem processing by performing the first operation on the first audio and / or issuing the first prompt.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of audio processing, and particularly relates to a processing method and a processing device. Background Art

[0002] There are often problems of sound quality loss in the audio processing process. How to determine the problem and the link where the problem lies to process the corresponding problem has become a technical problem that needs to be solved urgently in this field. Summary of the Invention

[0003] For this reason, this application discloses the following technical solutions:

[0004] A processing method, including:

[0005] In response to obtaining a first text converted from a first audio and a second text converted from a second audio, performing a comparison process on the first text and the second text;

[0006] According to the comparison result of the first text and the second text, determining whether to perform a first operation on the first audio and / or issue a first prompt;

[0007] Wherein, the second audio is obtained by performing a target process on the first audio.

[0008] Optionally, the performing a comparison process on the first text and the second text includes:

[0009] Determining the text content similarity between the first text and the second text.

[0010] Optionally, the first text and the second text respectively correspond to corresponding timestamps, and the timestamps are timestamps calibrated based on the positions of audio byte streams;

[0011] Wherein, by aligning the text contents of the first text and the second text according to the timestamps, and based on the text contents aligned according to the timestamps, determining the text content similarity between the first text and the second text.

[0012] Optionally, the first operation includes: canceling the operation of the target process or optimizing the operation of the target process;

[0013] The determining whether to perform a first operation on the first audio according to the comparison result of the first text and the second text includes:

[0014] If the text content similarity does not meet the corresponding similarity condition, canceling or optimizing the target process of the first audio.

[0015] Optionally, the first prompt includes one or more of the following: audio anomaly prompt, corresponding optional operation prompt, and anomaly cause prompt;

[0016] Determining whether to issue the first prompt according to the comparison result of the first text and the second text includes:

[0017] If the text content similarity does not meet the corresponding similarity condition, then prompt for audio anomaly and / or anomaly cause and / or corresponding optional operation according to the corresponding situation.

[0018] Optionally, the method further includes:

[0019] In response to the text content similarity between the first text and the second text meeting the corresponding similarity condition, when the voice audio of the call counterpart is not received at the sending end of the first audio, insert a target noise that meets the preset perception condition into the audio listened to by the sending end of the first audio.

[0020] Optionally, the target processing includes one or more of the following: propagation, compression, gain, and noise reduction.

[0021] Optionally, the first audio and the second audio include one or more of the following:

[0022] Call audio, recording audio, broadcast audio, and conference audio.

[0023] Optionally, when the first audio and the second audio are call audio, the first audio is closer to the audio input end in the propagation path than the second audio.

[0024] A processing device includes:

[0025] A comparison module, configured to, in response to obtaining a first text converted from a first audio and a second text converted from a second audio, perform a comparison process on the first text and the second text;

[0026] A determination module, configured to determine whether to perform a first operation on the first audio and / or issue a first prompt according to the comparison result of the first text and the second text;

[0027] Wherein, the second audio is obtained by performing target processing on the first audio.

[0028] An embodiment of the present application also discloses an electronic device, including:

[0029] A memory, configured to store a computer instruction set;

[0030] The computer instruction set can be implemented in the form of a computer program.

[0031] A processor for implementing the processing method disclosed in any one of the above by executing a computer instruction set.

[0032] As can be seen from the above solutions, the processing method and device disclosed in this application, in response to obtaining the first text converted from the first audio and the second text converted from the second audio (obtained by performing target processing on the first audio), perform a comparison process on the first text and the second text, and determine whether to perform a first operation on the first audio and / or issue a first prompt according to the comparison result of the first text and the second text. Through the above text comparison process, this application can effectively determine the sound quality loss problem and the location of the problem in the process of performing target processing on the first audio, and perform corresponding problem processing by performing a first operation on the first audio and / or issuing a first prompt. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0034] Figure 1 It is a schematic flowchart of a processing method provided by the present application;

[0035] Figure 2 It is another schematic flowchart of a processing method provided by the present application;

[0036] Figure 3 It is yet another schematic flowchart of a processing method provided by the present application;

[0037] Figure 4 It is a schematic diagram of the processing process of call audio in a dual - end call application provided by the present application;

[0038] Figure 5 It is a composition structure diagram of a processing device provided by the present application;

[0039] Figure 6 It is a composition structure diagram of an electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0041] The present application discloses a processing method and a processing device, which are applicable to any multi-node processing scenario of audio, and are used to improve the audio quality loss problem of audio in the multi-node processing scenario of audio and / or help users automatically identify the reasons for audio quality loss and facilitate anomaly troubleshooting. The processing method can be applied to an electronic device. The electronic device applying the processing method can be an independent device, such as a voice recorder, a conference recorder, an audio playback device, etc. Or, the electronic device applying the processing method can also be multiple devices that cooperate to participate in audio processing, such as the two devices in a voice call / conference call scenario, etc., and there is no limitation thereto.

[0042] See Figure 1 As shown in the flowchart of the processing method, the processing method provided by the embodiments of the present application at least includes the following processing flows:

[0043] Step 101, in response to obtaining the first text converted from the first audio and the second text converted from the second audio, perform a comparison process on the first text and the second text.

[0044] The first audio can be, but is not limited to, any one or more of call audio, recording audio, broadcast audio, and conference audio. The second audio is obtained by performing target processing on the first audio.

[0045] The target processing can be any processing involved in any path node in the audio flow path from audio acquisition, transmission to audio storage, playback, and even to secondary acquisition and secondary playback, including but not limited to any one or more of audio propagation, compression, gain, and noise reduction processing.

[0046] Here, the secondary acquisition refers to playing the acquired / received audio and then acquiring the played audio. Correspondingly, the secondary playback refers to playing the audio obtained through secondary acquisition.

[0047] In the embodiments of the present application, according to actual requirements, different target path nodes on the audio flow path of the audio are selected: the first target path node and the second target path node. During the audio flow process, the audio at each target path node is acquired, and text conversion processing is performed on the audio to obtain corresponding text information. The two selected target path nodes can be any different nodes on the audio flow path, depending on actual requirements.

[0048] Among them, the time corresponding to the second target path node is later than the time corresponding to the first target path node. The first audio is obtained by acquiring the audio at the first target path node, and the first text is obtained by performing text conversion processing on the first audio based on technologies such as ASR (Automatic Speech Recognition). Correspondingly, the second audio is obtained by acquiring the audio at the second target path node, and the second text is obtained by performing text conversion processing on the second audio based on ASR and the like.

[0049] For example, in the scenario of recording application, during the process of the conference recorder collecting the original audio, performing noise reduction on the original audio, and then saving the noise-reduced audio, the original audio collected by the conference recorder is obtained as the first audio, the noise-reduced audio is obtained as the second audio, and the first text and the second text are respectively obtained by performing ASR processing on the first audio and the second audio; or, in the scenario of call application, the voice of the audio initiator collected by the audio initiator device is obtained as the first audio, the audio transmitted to the communication peer through the communication channel is obtained as the second audio, and the first text and the second text are respectively obtained by performing ASR processing on the first audio and the second audio, etc.

[0050] In response to obtaining the first text converted from the first audio and the second text converted from the second audio, the first text and the second text are compared.

[0051] Specifically, in this embodiment, the text content similarity between the first text and the second text is determined, and the comparison between the two is realized by determining the similarity between the two texts. The first text and the second text respectively correspond to corresponding timestamps. By aligning the text content of the first text and the second text according to the timestamps, and based on the text content aligned according to the timestamps, the text content similarity between the first text and the second text is determined.

[0052] Optionally, during the process of performing speech recognition on the first audio or the second audio, or after completing speech recognition to obtain the first text and the second text, timestamps are respectively calibrated for each text content of the first text and the second text to obtain the timestamps of each text content of the first text and the second text. This timestamp is not a natural time, but a timestamp calibrated by the position of the audio byte stream.

[0053] When determining the similarity between the first text and the second text, first align the text content of the first text and the second text according to the timestamps, and then compare the content aligned according to the timestamps in the first text and the second text to determine the text content similarity between the first text and the second text.

[0054] Step 102: Determine whether to perform the first operation on the first audio and / or issue the first prompt according to the comparison result between the first text and the second text.

[0055] Among them, the first operation includes: an operation of canceling the target processing of the first audio or optimizing the target processing. Correspondingly, according to the comparison result of the first text and the second text, it is determined whether to perform the first operation on the first audio, which can be further implemented as:

[0056] 11) Determine whether the text content similarity between the first text and the second text meets the corresponding similarity condition.

[0057] The similarity condition here can be, but is not limited to, a similarity threshold preset to represent a relatively high similarity, or a similarity value range / interval.

[0058] Correspondingly, the text content similarity between the first text and the second text meeting the corresponding similarity condition can be, but is not limited to, that the text content similarity between the first text and the second text reaches the set similarity threshold, or is within the set similarity value range / interval.

[0059] 12) If not satisfied, cancel or optimize the target processing of the first audio.

[0060] If the text content similarity between the first text and the second text does not meet the corresponding similarity condition, it indicates that the text content similarity between the first text and the second text is relatively low. Correspondingly, it indicates that in the process of performing target processing on the first audio to obtain the second audio, the target processing of the first audio causes a loss of audio quality (audio quality loss) in terms of audio content in the second audio compared to the first audio, and the loss degree exceeds the allowable loss degree defined based on the set similarity condition.

[0061] Based on this, in response to the determination result that the text content similarity between the first text and the second text does not meet the corresponding similarity condition, in one embodiment, the target processing of the first audio is canceled. By canceling the target processing of the first audio, serious damage / distortion of the audio content is avoided.

[0062] For example, in the process of collecting the original audio, performing noise reduction on the original audio and then saving the noise-reduced audio, or in the process of performing noise reduction optimization on the audio stream of the audio and playing the noise-reduced audio stream in real time, based on the monitoring that the text content similarity between the audio before and after noise reduction does not meet the corresponding configured similarity condition, the noise reduction processing of the audio is canceled, and the collected original audio is directly saved, or the non-noise-reduced audio is directly played to avoid losses / distortion of the audio content beyond the allowable degree during the noise reduction process.

[0063] However, without limitation, in other embodiments, in response to the determination result that the text content similarity between the first text and the second text does not meet the similarity condition, the target processing of the first audio can also be optimized. By optimizing the target processing of the first audio, the audio quality of the obtained second audio can be improved, and serious damage / distortion of the audio content can be avoided.

[0064] Continuing with the above example, based on the fact that the text content similarity of the audio before and after noise reduction does not meet the similarity condition configured correspondingly, the audio quality of the obtained second audio can be improved by, but not limited to, optimizing the filter parameters of the noise reduction filter. Or, in the compression processing of the audio, based on the fact that the text content similarity of the audio before and after compression does not meet the similarity condition configured correspondingly, by optimizing parameters such as the sampling rate, quantization accuracy, compression ratio, bit rate, etc. of the compression processing, such as increasing the sampling rate and reducing the compression ratio, the quality of the compressed audio can be improved.

[0065] The first prompt includes, but is not limited to, one or more of the following: audio anomaly prompt, corresponding optional operation prompt, and anomaly cause prompt.

[0066] Correspondingly, determining whether to issue the first prompt to the first audio according to the comparison result between the first text and the second text can be further implemented as:

[0067] 21) Determine whether the text content similarity between the first text and the second text meets the corresponding similarity condition.

[0068] Similarly, the similarity condition here can also be, but not limited to, a similarity threshold representing a relatively high similarity set in advance, or a similarity value range / interval. That the text content similarity between the first text and the second text meets the corresponding similarity condition can be, but not limited to, that the text content similarity between the first text and the second text reaches the set similarity threshold, or is within the set similarity value range / interval.

[0069] The similarity condition here can be the same as or different from the similarity condition in step 11), and no limitation is made on this.

[0070] 22) If not satisfied, prompt for audio anomaly and / or anomaly cause and / or corresponding optional operation according to the corresponding situation.

[0071] If the text content similarity between the first text and the second text does not meet the corresponding similarity condition, the corresponding representation indicates that the target processing of the first audio results in a loss of audio quality (audio quality loss) in the audio content of the second audio compared to the first audio, and the degree of loss exceeds the allowable loss degree defined based on the corresponding similarity condition. In response to this situation, this embodiment issues a first prompt for the first audio, including but not limited to any one or more of a prompt for an audio anomaly event, a prompt for the cause of the anomaly, and a prompt for corresponding optional operations.

[0072] Optionally, the corresponding optional operation corresponds to the cause of the anomaly. By prompting the optional operation corresponding to the cause of the anomaly, the user can perform the corresponding operation according to the prompt information to overcome the anomaly in audio processing or at least reduce the degree of the anomaly, thereby improving the quality of the second audio obtained by performing target processing on the first audio.

[0073] As can be seen from the above solution, the processing method disclosed in this application, in response to obtaining the first text converted from the first audio and the second text converted from the second audio (obtained by performing target processing on the first audio), performs a comparison process on the first text and the second text, and determines whether to perform a first operation on the first audio and / or issue a first prompt based on the comparison result of the first text and the second text. Through the above text comparison process, this application can effectively determine the problem of audio quality loss and the link where the problem lies during the target processing of the first audio, and perform corresponding problem processing by performing a first operation on the first audio and / or issuing a first prompt.

[0074] Optionally, in one embodiment, the first audio and the second audio are call audios, such as call audios in scenarios such as voice calls and conference calls, and the first audio and the second audio are respectively audios of the same call voice at different propagation nodes on the call path.

[0075] In this embodiment, refer to Figure 2 The processing method provided, the processing method of this application can be specifically implemented as:

[0076] Step 201, obtain texts converted from audios of the same call voice at different propagation nodes on the call path, and use them as the first text and the second text respectively.

[0077] Among them, the first audio is closer to the audio input end on the call path (the propagation path of the call) than the second audio.

[0078] In this embodiment, the devices of the two call parties are the first electronic device and the second electronic device respectively, and the first electronic device is the current initiator of the voice audio. Correspondingly, the first text and the second text can be obtained specifically by performing any one or more of the following processes, but not limited to:

[0079] 31) Obtain the initiating - end text of the voice - audio conversion currently initiated by the first electronic device as the first text, and obtain the receiving - end first sub - text of the voice - audio conversion received by the second electronic device as the second text;

[0080] 32) Obtain the initiating - end text of the voice - audio conversion currently initiated by the first electronic device as the first text, and obtain the receiving - end second sub - text obtained by the second electronic device by playing the received voice - audio, collecting the played audio, and performing text conversion on the collected audio as the second text;

[0081] 33) Obtain the receiving - end first sub - text of the voice - audio conversion received by the second electronic device as the first text, and obtain the receiving - end second sub - text obtained by the second electronic device by playing the received voice - audio, collecting the played audio, and performing text conversion on the collected audio as the second text.

[0082] Step 202: Determine the text - content similarity between the first text and the second text.

[0083] Among them, the first text and the second text respectively correspond to corresponding timestamps, and the timestamps corresponding to the first text and the second text are timestamps calibrated based on the positions of audio byte streams.

[0084] In implementation, the text contents of the first text and the second text can be aligned by timestamps first, and after alignment, based on the text contents aligned by timestamps, determine the text - content similarity between the first text and the second text.

[0085] For the above several situations of the first text and the second text, specifically, based on the timestamp alignment method, by performing one or more of the following processes, determine the text - content similarity between the first text and the second text:

[0086] 41) Determine the first text - content similarity between the initiating - end text and the receiving - end first sub - text;

[0087] 42) Determine the second text - content similarity between the initiating - end text and the receiving - end second sub - text;

[0088] 43) Determine the third text - content similarity between the receiving - end first sub - text and the receiving - end second sub - text.

[0089] Step 203: Determine whether the text - content similarity between the first text and the second text meets the corresponding similarity condition.

[0090] After that, further determine whether the text - content similarity between the first text and the second text meets the corresponding similarity condition.

[0091] For example, determine whether the text content similarity between the first text and the second text reaches a set similarity threshold or falls within a set similarity range / interval. If it reaches the set similarity threshold or falls within the set similarity range / interval, then the text content similarity between the first text and the second text meets the corresponding similarity condition; otherwise, it does not.

[0092] In implementation, for one or more of the first text content similarity, the second text content similarity, and the third text content similarity in 41)-43), it is possible to determine whether they meet the similarity conditions by setting corresponding one or more similarity conditions (similarity threshold / similarity range) respectively. And when setting multiple similarity conditions, the multiple similarity conditions set can be the same or different, and there is no restriction on this.

[0093] Step 204, if not met, issue a first prompt on the first electronic device and / or the second electronic device that are the two parties of the call, and / or perform a first operation on the second electronic device.

[0094] Optionally, it is possible to determine whether at least one of the above-mentioned first text content similarity, second text content similarity, and third text content similarity does not meet the corresponding similarity condition. If so, it is considered that the text content of the first text and the second text does not meet the similarity condition. In this case, issue a first prompt on the first electronic device and / or the second electronic device, and / or perform a first operation on the second electronic device.

[0095] Issuing a first prompt on the first electronic device and / or the second electronic device can be further implemented as: prompting at least one of an abnormal event of abnormal voice audio listening, the cause of the abnormality, and the corresponding voice recognition rate when listening abnormally on the first electronic device and / or the second electronic device;

[0096] Among them, if the first text content similarity does not meet the corresponding similarity condition, the cause of the abnormality is related to the quality of the communication channel, such as abnormal communication quality of the communication channel, etc.; if the third text content similarity does not meet the corresponding similarity condition, or the text content of the second sub-text at the receiving end is empty, it is determined that the cause of the abnormality is related to the voice listening environment at the second electronic device end.

[0097] The voice listening environment at the second electronic device end can be, but is not limited to, the hardware environment or the surrounding environment of the second electronic device. For example, the speaker / mic of the second electronic device itself is not turned on, or is damaged, or the volume is too small, or there is noise interference to the second electronic device, etc.

[0098] The voice recognition rate corresponding to abnormal listening is determined by the text content similarity between the first text and the second text. Specifically, the voice recognition rate corresponding to abnormal listening can be determined according to the first text content similarity between the initiating text and the first sub-text of the receiving end, or according to the second text content similarity between the initiating text and the second sub-text of the receiving end, or according to the determination of the first text content similarity between the initiating text and the first sub-text of the receiving end and the third text content similarity between the first sub-text of the receiving end and the second sub-text of the receiving end, to determine the voice recognition rate corresponding to abnormal listening, etc., and there is no limitation on this.

[0099] In the traditional technology, in remote calls such as conferences, in order to know whether their own speaking voice is heard by the other party, people often need to determine it through manual inquiry. For example, the speaker asks the other party: "Can you hear my voice? Clearly?" And when the listener does not hear the other party's voice, they often cannot determine whether the other party did not speak or there is a problem with the call. The two parties in the call can only confirm through manual means relying on mutual language communication, which affects the service quality of voice communication and at the same time results in a poor user experience.

[0100] In this embodiment, by comparing the audio texts of different propagation nodes on the call path and automatically detecting and analyzing the reasons for abnormal listening to the call audio based on the text differences, the continuous tracking of the call state during the call is realized, and by automatically giving relevant abnormal prompts such as abnormal listening and root cause when the audio listening is abnormal, the above problems of the traditional technology are solved, without the two parties in the call having to determine the call state through manual conversation, and it is convenient for the two parties in the call to troubleshoot abnormal listening.

[0101] In one embodiment, refer to Figure 3 the provided flowchart of the processing method, Figure 2 After step 204 of the processing method corresponding to the corresponding embodiment, it may further include:

[0102] Step 205, in response to the text content similarity between the first text and the second text satisfying the corresponding similarity condition, insert a target noise that meets the preset perception condition into the audio listened to by the sending end of the first audio when the voice audio of the call peer is not received at the sending end of the first audio.

[0103] This embodiment also targets the scenario of a call application. In this scenario, if the text content similarity between the first text and the second text meets the corresponding similarity condition, such as the similarity between the two reaches a set similarity threshold or falls within a set similarity range, it indicates that there is no abnormality in the audio listening on the second electronic device for the audio sent from the first electronic device. In this case, when the first electronic device sends audio to the second electronic device and the first electronic device does not receive the audio of the second electronic device, that is, the sending end of the first audio does not receive the voice audio of the call counterpart (for example, in a two-party call, when the user of the first electronic device speaks to the other party but the user of the second electronic device does not speak), a target noise that meets the preset perception condition is inserted into the audio listened to by the sending end of the first audio (that is, the audio listened to by the first electronic device). Optionally, the target noise that meets the preset perception condition is comfort noise, such as pink noise, so that the user of the first electronic device can perceive a certain amount of comfort noise in addition to their own voice during the speaking process, rather than being in a completely silent state.

[0104] In this embodiment, when there is no abnormality in audio listening during a two-party call and in the single-user speaking scenario where the user of the first electronic device speaks to the other party but the user of the second electronic device does not speak, a target noise that meets the preset perception condition, such as pink noise, is inserted into the audio listened to by the audio sending end, that is, the first electronic device, which can prevent the user at the audio sending end from mistakenly thinking that the other party cannot hear or has hung up the phone due to excessive quietness on the other side in this single-user speaking scenario.

[0105] The following details the processing process of the first audio and the second audio as call audio with an example.

[0106] Refer to Figure 4 In the example, it is a typical two-party call system. The two parties in the call are Endpoint A and Endpoint B respectively, and Endpoint A is the current voice initiator. When one party speaks, the other party will play the sound from the speaker and record it at its mic end, and then the part belonging to the echo will be removed through AEC (echo cancellation). In this example, ASR (automatic speech recognition) is performed at three audio passing points on both ends to obtain the text and timestamp of the audio conversion. This timestamp is not the natural time but the timestamp calibrated by the position of the audio byte stream.

[0107] Combined with reference to Figure 4 In this example, the implementation process of performing quality determination and prompting and other processing on the voice audio during the call based on the method of this application includes:

[0108] When the audio of Endpoint A enters the MIC, perform ASR processing (ASR1) on the audio to obtain the corresponding text ASRText1, calibrate the timestamp for the text, and record the text ASRText1 and the timestamp locally;

[0109] Among them, the ASR processing of the audio of Endpoint A entering the MIC can be before or after the echo cancellation of the audio, and there is no restriction on this. Preferably, after the echo cancellation of the audio of Endpoint A entering the MIC, perform ASR processing on the audio.

[0110] 52) After the audio arrives at Endpoint B via the communication channel, before playing through the speaker, perform ASR processing (i.e., ASR2) and timestamp calibration on the audio received by Endpoint B, and record the corresponding text ASRText2 and the timestamp at the Endpoint B side;

[0111] 53) After the audio is externally played at Endpoint B and enters the MIC of Endpoint B, after entering the MIC of Endpoint B, perform ASR processing (i.e., ASR3) and timestamp calibration on the audio entering the MIC of Endpoint B, and record the corresponding text ASRText3 and the timestamp at the Endpoint B side;

[0112] Preferably, after the audio enters the MIC of Endpoint B and before the echo cancellation of the audio, perform ASR processing on the audio.

[0113] 54) Endpoint B sends back the text ASRText2, ASRText3 obtained based on ASR2, ASR3 and their respective corresponding timestamps to Endpoint A;

[0114] 55) Endpoint A receives the text ASRText2, ASRText3 and their timestamps sent back by Endpoint B, and analyzes and evaluates the listening quality of the audio of Endpoint A at Endpoint B based on the text and the corresponding timestamps obtained through three times of ASR processing. When the listening is abnormal, a first prompt is issued on the first electronic device and / or the second electronic device, and / or a first operation is performed on the second electronic device.

[0115] Optionally, a Voice Quality Monitor is set at Endpoint A. The three texts and corresponding timestamps obtained based on three ASRs are analyzed in real time through the set Voice Quality Monitor to evaluate the listening quality of the audio at Endpoint A on Endpoint B. When the listening is abnormal, a first prompt is issued on the first electronic device and / or the second electronic device, and / or a first operation is performed on the second electronic device, specifically as follows:

[0116] a) Align the three ASR texts ASRText1, ASRText2, and ASRText3 based on the timestamps;

[0117] b) Use standard algorithms such as WER (Word Error Rate) to score the similarity of the three texts ASRText1, ASRText2, and ASRText3 obtained through ASR processing, and obtain the similarity scores between every two texts;

[0118] c) Determine the listening quality of the audio based on the similarity scores of the corresponding texts in ASRText1, ASRText2, and ASRText3, and analyze the cause of the abnormality when the listening is abnormal. Specifically as follows:

[0119] If the text similarities of the 3 ASRTexts all reach the corresponding similarity thresholds, it indicates that the listening quality of the audio at Endpoint A on Endpoint B is good and there is no abnormality in audio listening;

[0120] If ASRText1 and ASRText2 reach the corresponding similarity thresholds, it is considered that the communication channel quality between Endpoint A and Endpoint B is good. Otherwise, if the corresponding similarity thresholds are not reached, the listening is abnormal, and the cause of the abnormality is the communication channel quality problem between Endpoint A and Endpoint B;

[0121] If ASRText2 and ASRText3 do not reach the corresponding similarity thresholds, the listening is abnormal, and the cause of the abnormality is that the mic acquisition quality of Endpoint B is poor or there are interferences such as speech / noise;

[0122] If ASRText2 has content but ASRText3 has no content, that is, the content is empty, the listening is abnormal, and the cause of the abnormality is that the MIC of Endpoint B is not turned on, or the mic input source is selected incorrectly, or there are serious problems in audio acquisition (such as mic hardware damage, driver damage, software failure, etc.), or the speaker volume is too low / speaker is damaged.

[0123] After evaluation and analysis, if the audio of Endpoint A is abnormally heard at Endpoint B, at least one of the abnormal listening event, the reason for the abnormality (such as channel quality problems or noise interference, or the MIC of Endpoint B is not turned on, etc.), the corresponding selectable operations, and the speech recognition rate during abnormal listening is prompted at Endpoint A and / or Endpoint B, or the corresponding first operation is executed on the corresponding electronic device to improve the listening quality of the audio. Optionally, information can be prompted by means such as, but not limited to, beeps, pictures, and / or text to help users at both ends troubleshoot the reasons for the abnormality.

[0124] Among them, the above "selectable operations" are related to the reasons for the abnormal audio listening. For example, if after analysis, the reason for the abnormality is that the MIC of Endpoint B is not turned on, or there are serious problems with audio acquisition, the selectable operations prompted can be: please turn on the MIC of Endpoint B, or please check whether there are hardware damages, driver damages, software failures, etc. in the MIC of Endpoint B.

[0125] The first operation executed on the corresponding electronic device is also related to the reason for the abnormality. In this example, the first operation executed can be, but not limited to, performing self-check and listening quality restoration processing based on the analyzed reason for the abnormal listening at Endpoint B. For example, for the reason that the speaker is not turned on / has a low volume, the speaker is automatically turned on or the volume is adjusted, and for the reason that the input source of the MIC is selected incorrectly, the input source is automatically changed, etc.

[0126] 56) When it is confirmed that there is no problem with the listening quality at Endpoint B, that is, there is no abnormal listening, and no voice audio from Endpoint B is received during the process of Endpoint A sending audio, comfort noise, such as pink noise, is inserted into the audio played by the speaker of Endpoint A to prevent the speaker at Endpoint A from mistakenly thinking that the other party cannot hear due to excessive quietness at Endpoint B, or mistakenly thinking that the call has been hung up.

[0127] In other embodiments, the first audio can also be other forms of audio different from the call audio, such as recorded audio, broadcast audio, etc. In this embodiment, the target processing of the first audio can be any one or more of, but not limited to, compression, gain, and noise reduction, etc. Correspondingly, for recorded audio and broadcast audio, the processing quality of the target processing of the audio in the audio path can be detected based on the difference comparison of the audio texts at any different nodes in the audio path, and whether to execute the first operation and / or issue the first prompt for the first audio is determined based on the fact that the text similarity of the audio texts at different nodes does not meet the corresponding similarity conditions.

[0128] For example, for meeting recordings, when collecting meeting audio and saving the meeting audio after noise reduction processing, if the text content similarity of the audio text before and after noise reduction is monitored and does not meet the similarity conditions of the corresponding configuration, the noise reduction processing of the audio is revoked and the originally collected raw audio is directly saved. Or, the filter parameters during noise reduction are optimized, or a corresponding prompt is given, and the user manually operates to revoke the noise reduction processing of the audio or perform parameter optimization and other processing.

[0129] Or, during the process of noise reduction optimization of the audio stream of the audio and real-time playing of the noise-reduced audio stream, if the text content similarity of the audio text before and after noise reduction is monitored and does not meet the similarity conditions of the corresponding configuration, the noise reduction processing of the audio before playing is revoked, and the non-noise-reduced audio is directly played to avoid losses beyond the allowable degree to the played audio content during the noise reduction process. Or, the filter parameters during noise reduction are optimized to optimize the noise reduction quality and reduce content distortion caused by noise reduction.

[0130] In the compression processing of the audio, if the text content similarity of the audio text before and after compression is monitored and does not meet the corresponding similarity conditions, the quality of the compressed audio is improved by optimizing parameters such as the sampling rate, quantization precision, compression ratio, and bit rate of the compression processing, such as increasing the sampling rate and reducing the compression ratio. In the gain processing of the audio, if the text content similarity of the audio text before and after gain processing is monitored and does not meet the corresponding similarity conditions, the quality of the gain processing is improved by changing the amplification factor and reducing distortion, or the gain processing of the audio is stopped, etc.

[0131] By comparing the differences in the audio texts of any different nodes in the audio path for recorded audio and broadcast audio, the processing quality of the target processing of the audio in the audio path is detected. And if the text similarity of the audio texts of different nodes does not meet the corresponding similarity conditions, a first operation is performed on the first audio and / or a first prompt is issued to stop the target processing of the first audio or optimize the target processing of the first audio, thereby improving the losses / distortion caused by the target processing to the audio content.

[0132] Corresponding to the above processing method, an embodiment of the present application further provides a processing device, and the composition structure of the processing device is as Figure 5 shown, including at least:

[0133] A comparison module 501, configured to perform a comparison process on the first text and the second text in response to obtaining the first text converted from the first audio and the second text converted from the second audio;

[0134] A determination module 502, configured to determine whether to perform a first operation on the first audio and / or issue a first prompt according to the comparison result of the first text and the second text;

[0135] Wherein, the second audio is obtained by performing target processing on the first audio.

[0136] In one embodiment, the comparison module 501 is specifically configured to:

[0137] Determine the text content similarity between the first text and the second text.

[0138] In one embodiment, the first text and the second text respectively correspond to corresponding timestamps, and the timestamps are timestamps calibrated based on the positions of audio byte streams;

[0139] Wherein, by aligning the text contents of the first text and the second text according to the timestamps, and based on the text contents aligned according to the timestamps, the text content similarity between the first text and the second text is determined.

[0140] In one embodiment, the first operation includes: an operation to cancel the target processing or an operation to optimize the target processing;

[0141] When determining whether to perform the first operation on the first audio according to the comparison result between the first text and the second text, the determination module 502 is specifically configured to:

[0142] If the text content similarity does not meet the corresponding similarity condition, cancel or optimize the target processing of the first audio.

[0143] In one embodiment, the first prompt includes one or more of the following: audio anomaly prompt, corresponding optional operation prompt, and anomaly cause prompt;

[0144] When determining whether to issue the first prompt according to the comparison result between the first text and the second text, the determination module 502 is specifically configured to:

[0145] If the text content similarity does not meet the corresponding similarity condition, prompt for audio anomaly and / or anomaly cause and / or corresponding optional operation according to the corresponding situation.

[0146] In one embodiment, the above device further includes:

[0147] An insertion module, configured to insert target noise that meets the preset perception condition into the first audio when the receiving end of the second audio does not receive voice audio in response to the text content similarity between the first text and the second text meeting the corresponding similarity condition.

[0148] In one embodiment, the target processing includes one or more of the following: propagation, compression, gain, and noise reduction.

[0149] In one embodiment, the first audio and the second audio include one or more of the following:

[0150] Call audio, recorded audio, broadcast audio, and conference audio.

[0151] In one embodiment, when the first audio and the second audio are call audio, the first audio is closer to the audio input end in the propagation path than the second audio.

[0152] The embodiments of the present application also disclose an electronic device, which may be, but is not limited to, a recording pen, a conference recorder, a smart phone, a tablet computer, a personal computer, or other devices capable of providing computing / processing capabilities.

[0153] The composition structure of the electronic device, as Figure 6 shown, at least includes:

[0154] A memory 10 for storing a computer instruction set;

[0155] The computer instruction set can be implemented in the form of a computer program.

[0156] A processor 20 for implementing the processing method disclosed in any of the above method embodiments by executing the computer instruction set.

[0157] The processor 20 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, etc.

[0158] In addition, the electronic device may further include components such as a communication interface and a communication bus. The memory, the processor, and the communication interface complete communication with each other through the communication bus.

[0159] The communication interface is used for communication between the electronic device and other devices. The communication bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus may be divided into an address bus, a data bus, a control bus, etc.

[0160] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.

[0161] For the convenience of description, when describing the above system or device, it is divided into various modules or units according to functions for separate description. Of course, when implementing the present application, the functions of each unit can be realized in the same or multiple software and / or hardware.

[0162] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.

[0163] Finally, it should also be noted that in this article, relational terms such as first, second, third, and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0164] The above are only the preferred embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A processing method, comprising: responding to obtaining a first text converted from a first audio and a second text converted from a second audio, and performing a comparison process on the first text and the second text; realizing the comparison process by determining the similarity between the two texts; determining whether to perform a first operation on the first audio and / or issue a first prompt according to the comparison result of the first text and the second text, so as to avoid quality loss of the second audio compared with the first audio; the second audio is obtained by performing target processing on the first audio.

2. The method according to claim 1, wherein the performing a comparison process on the first text and the second text, comprising: determining the text content similarity between the first text and the second text.

3. The method according to claim 2, wherein the first text and the second text respectively correspond to corresponding timestamps, and the timestamps are timestamps calibrated based on the positions of audio byte streams; wherein, by aligning the text contents of the first text and the second text according to the timestamps, and based on the text contents aligned according to the timestamps, determining the text content similarity between the first text and the second text.

4. The method according to claim 2, wherein the first operation comprising: canceling the operation of the target processing or optimizing the operation of the target processing; the determining whether to perform a first operation on the first audio according to the comparison result of the first text and the second text, comprising: if the text content similarity does not meet the corresponding similarity condition, canceling or optimizing the target processing of the first audio.

5. The method according to claim 2, wherein the first prompt includes one or more of the following: audio anomaly prompt, corresponding optional operation prompt, and anomaly cause prompt; the determining whether to issue a first prompt according to the comparison result of the first text and the second text, comprising: if the text content similarity does not meet the corresponding similarity condition, prompting audio anomaly and / or anomaly cause and / or corresponding optional operation according to the corresponding situation.

6. The method according to claim 2, further comprising: responding to the text content similarity between the first text and the second text meeting the corresponding similarity condition, inserting target noise that meets a preset perception condition into the audio listened to by the sending end of the first audio when the voice audio of the call counterpart is not received at the sending end of the first audio.

7. The method according to claim 1, wherein the target processing includes one or more of the following: propagation, compression, gain, and noise reduction.

8. The method according to claim 5, wherein the first audio and the second audio include one or more of the following: call audio, recording audio, broadcast audio, and conference audio.

9. The method according to claim 6, when the first audio and the second audio are call audio, the first audio is closer to the audio input end in the propagation path than the second audio.

10. A processing device, comprising: A comparison module, configured to perform a comparison process on the first text converted from the first audio and the second text converted from the second audio in response to obtaining them; the comparison process is implemented by determining the similarity between the two texts; A determination module, configured to determine whether to perform a first operation on the first audio and / or issue a first prompt according to the comparison result of the first text and the second text, so as to avoid sound quality loss of the second audio compared with the first audio; The second audio is obtained by performing a target process on the first audio.

Citation Information

Patent Citations

  • Text acquisition and live broadcast methods and devices, and storage medium

    CN113761986A

  • Voice verification method and device, electronic equipment and medium

    CN113889145A