Audio data processing method, device, audio data processing equipment and storage medium

By using a combination of scene common thesaurus and target term thesaurus in audio data processing, the problem of inaccurate term word processing in audio data conversion is solved, and higher text conversion accuracy and efficiency are achieved.

CN113761118BActive Publication Date: 2025-08-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110437064.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-22
Publication Date
2025-08-26
Estimated Expiration
2041-04-22

AI Technical Summary

Technical Problem

In the prior art, when the audio data is directly converted through a general vocabulary, it is impossible to accurately process the specific terms and words of the scene, resulting in inaccurate text conversion results.

Method used

After the preliminary text conversion is performed using the scene common lexicon in the target recognition scenario, the abnormal conversion words are corrected using the target term lexicon associated with the target recognition scenario, replaced with the target term lexicon, and modified text data is generated.

Benefits of technology

It improves the accuracy of text conversion, reduces the typo rate, saves human resources, improves processing efficiency, and avoids dependence on the capabilities of audio data processing personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113761118B_ABST
    Figure CN113761118B_ABST
Patent Text Reader

Abstract

The present application discloses an audio data processing method, apparatus, audio data processing device, and storage medium. The method comprises: obtaining target audio data collected in a target recognition scenario and obtaining a scenario-wide vocabulary; the scenario-wide vocabulary comprises M scenario-wide words, where M is a positive integer; performing text conversion on the target audio data based on the M scenario-wide words to obtain converted text data; obtaining a target terminology vocabulary associated with the target recognition scenario and obtaining abnormal conversion words in the converted text data; the target terminology vocabulary comprises N terminology words in the target recognition scenario, where N is a positive integer; obtaining target terminology words corresponding to the abnormal conversion words from the N terminology words, replacing the abnormal conversion words in the converted text data with the target terminology words, and obtaining corrected text data of the converted text data. This method can improve the accuracy of text conversion of the target audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech processing technology, and in particular to an audio data processing method, an audio data processing apparatus, an audio data processing device, and a computer storage medium. Background Art

[0002] With the advancement of data processing technology and the rapid popularization of mobile Internet, speech recognition technology has been applied in more and more fields. Among them, speech recognition technology, also known as automatic speech recognition (ASR), allows computers to "dictate" speech, that is, to realize the technology of converting "audio data" into "text data".

[0003] In existing technologies, when converting audio data to text, the conversion is typically done directly using common words in a general vocabulary. However, since audio data can be acquired in any context, it is likely to contain specific words specific to that context, which are generally not common words. Therefore, directly converting audio data to text using a general vocabulary can result in inaccurate results. Summary of the Invention

[0004] The embodiments of the present application provide an audio data processing method, apparatus, audio data processing device, and storage medium, which can accurately convert target audio data into text data and improve the accuracy of text conversion.

[0005] In one aspect, an embodiment of the present application provides an audio data processing method, the audio data processing method comprising:

[0006] Obtain target audio data collected in a target recognition scenario and obtain a scenario-general vocabulary; the scenario-general vocabulary includes M scenario-general words, where M is a positive integer;

[0007] Perform text conversion on the target audio data based on M common scene words to obtain converted text data;

[0008] Obtain a target terminology lexicon associated with a target recognition scenario, and obtain abnormal conversion words in the converted text data; the target terminology lexicon includes N terminology words in the target recognition scenario, where N is a positive integer;

[0009] A target term corresponding to the abnormal conversion term is obtained from the N term words, and the abnormal conversion term in the conversion text data is replaced with the target term word to obtain corrected text data of the conversion text data.

[0010] In one aspect, an embodiment of the present application provides an audio data processing device, the audio data processing device comprising:

[0011] An acquisition unit, configured to acquire target audio data collected in a target recognition scenario and acquire a scene-general vocabulary; the scene-general vocabulary includes M scene-general words, where M is a positive integer;

[0012] a conversion unit, configured to perform text conversion on the target audio data based on the M scene-common words to obtain converted text data;

[0013] The acquisition unit is further configured to acquire a target terminology lexicon associated with the target recognition scenario and to acquire abnormal conversion words in the converted text data; the target terminology lexicon includes N terminology words in the target recognition scenario, where N is a positive integer;

[0014] The replacement unit is used to obtain the target term words corresponding to the abnormal conversion words from the N term words, replace the abnormal conversion words in the conversion text data with the target term words, and obtain the corrected text data of the conversion text data.

[0015] In one aspect, an embodiment of the present application provides an audio data processing device, the audio data processing device including an input interface and an output interface, and the audio data processing device further including:

[0016] a processor adapted to implement one or more instructions; and

[0017] A computer storage medium storing one or more instructions, wherein the one or more instructions are suitable for being loaded by a processor and executed by a processor:

[0018] Obtain target audio data collected in a target recognition scenario and obtain a scenario-general vocabulary; the scenario-general vocabulary includes M scenario-general words, where M is a positive integer;

[0019] Perform text conversion on the target audio data based on M common scene words to obtain converted text data;

[0020] Obtain a target terminology lexicon associated with a target recognition scenario, and obtain abnormal conversion words in the converted text data; the target terminology lexicon includes N terminology words in the target recognition scenario, where N is a positive integer;

[0021] A target term corresponding to the abnormal conversion term is obtained from the N term words, and the abnormal conversion term in the conversion text data is replaced with the target term word to obtain corrected text data of the conversion text data.

[0022] In one aspect, an embodiment of the present application provides a computer storage medium storing one or more instructions, wherein the one or more instructions are suitable for being loaded by a processor and executing the following steps:

[0023] Obtain target audio data collected in a target recognition scenario and obtain a scenario-general vocabulary; the scenario-general vocabulary includes M scenario-general words, where M is a positive integer;

[0024] Perform text conversion on the target audio data based on M common scene words to obtain converted text data;

[0025] Obtain a target terminology lexicon associated with a target recognition scenario, and obtain abnormal conversion words in the converted text data; the target terminology lexicon includes N terminology words in the target recognition scenario, where N is a positive integer;

[0026] A target term corresponding to the abnormal conversion term is obtained from the N term words, and the abnormal conversion term in the conversion text data is replaced with the target term word to obtain corrected text data of the conversion text data.

[0027] In an embodiment of the present application, when an audio data processing device obtains target audio data collected in a target recognition scenario, it can first perform text conversion on the target audio data based on the M scene-general words included in the scene-general word library to obtain converted text data, and then correct the abnormal conversion words in the converted text data according to the target terminology word library associated with the target recognition scenario, and replace the abnormal conversion words in the converted text data with the target terminology words in the target terminology word library to obtain corrected text data of the converted text data. Since the target terminology word library associated with the target recognition scenario is used to correct the converted text data of the target audio data in the target recognition scenario, compared with the solution in which the annotator manually performs text conversion on the target audio data in the target recognition scenario, there is no need to manually query the terminology words in the target recognition scenario, which can effectively save human resources and improve the processing efficiency of the target audio data; moreover, it is not limited by the ability of the audio data processing personnel, and the target terminology words corresponding to the target audio data can be accurately obtained through the target terminology word library, thereby reducing the error rate in the text data (corrected text data) and ensuring the accuracy of the text data (corrected text data) corresponding to the target audio data. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0029] Figure 1 1 is a schematic diagram of the system architecture of an audio data processing system provided in an embodiment of the present application;

[0030] Figure 2 This is a flowchart of an audio data processing method provided by an embodiment of the present application;

[0031] Figure 3 This is a flowchart of another audio data processing method provided by an embodiment of the present application;

[0032] Figure 4 This is a flowchart of another audio data processing method provided by an embodiment of the present application;

[0033] Figure 5 This is a schematic diagram of an interface of an information interaction platform provided in an embodiment of the present application;

[0034] Figure 6 This is a flow chart of generating a target terminology database provided by an embodiment of the present application;

[0035] Figure 7 This is a schematic diagram of the structure of a blockchain provided by an embodiment of the present application;

[0036] Figure 8 is a structural diagram of an audio data processing device provided in an embodiment of the present application;

[0037] Figure 9 It is a structural diagram of an audio data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0039] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0040] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0041] Among them, the embodiment of the present application proposes an audio data processing method based on speech processing technology, which can improve the accuracy of text conversion of target audio data in the target recognition scenario through the target terminology vocabulary in the target recognition scenario. Specifically, the target audio data collected in the target recognition scenario can be first converted into text using the M general words included in the scene general vocabulary to obtain converted text data, and then the target terminology vocabulary associated with the target recognition scenario is used to correct the abnormal conversion words in the converted text data to obtain corrected text data. Since the target audio data may contain some terminology words in the target recognition scenario, and the terminology words do not belong to the scene general words, there may be some abnormal conversion words in the converted text data obtained by text conversion based on the scene general vocabulary. The present application uses the terminology words in the target terminology vocabulary to correct the abnormal conversion words in the converted text data to obtain corrected text data, so that the accuracy of the text data converted from the target audio data is higher.

[0042] The accuracy of the text data can be evaluated by the error rate. The accuracy of the text data is negatively correlated with the error rate of the text data, that is, the lower the error rate of the text data, the higher the accuracy of the text data. The audio data processing device can calculate the error rate of the text data using the following expression:

[0043]

[0044] Among them, CER is used to represent the character error rate in text data, N is used to represent the total number of characters in text data, S is used to represent the number of replaced characters in text data, D is used to represent the number of deleted characters in text data, and I is used to represent the number of inserted characters in text data.

[0045] In one embodiment, the audio data processing method can be used to convert target audio data collected in a target recognition scenario into text, and determine the text data corresponding to the target audio data. When the audio data processing method is used to determine the text data corresponding to the target audio data, the audio data processing method can be applied in the following situations: Figure 1In the audio data processing system shown in FIG, the audio data processing system may include at least: an audio data acquisition device 11 and an audio data processing device 12. The audio data acquisition device 11 may be Figure 1 The independent microphone shown can also be other devices with microphones, such as smart phones, etc. The audio data processing device 12 can be Figure 1 The server shown in the figure can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, content delivery networks (CDN), middleware services, domain name services, security services, and basic cloud computing services such as big data and artificial intelligence platforms, etc. The audio data processing device 12 can also be a terminal device, which can include but is not limited to: smart phones, tablet computers, laptops, wearable devices, desktop computers, etc. Figure 1 In the embodiment, the audio data acquisition device 11 and the audio data processing device 12 are independent devices. It should be noted that in some feasible implementations, the audio data acquisition device 11 can also be a device embedded in the audio data processing device 12, and this application does not limit this.

[0046] See Figure 2 , is a flow chart of an audio data processing method proposed in an embodiment of the present application. Figure 2 As shown, the audio data processing method includes steps S201-S204:

[0047] S201, acquiring target audio data collected in a target recognition scene, and acquiring a scene-general vocabulary; the scene-general vocabulary includes M scene-general words, where M is a positive integer.

[0048] The target audio data may include one or more of the following types of data: audio or video files, a sentence of audio, or a real-time voice stream.

[0049] In one embodiment, the audio data processing device can directly obtain the target audio data in the target recognition scenario. In another embodiment, the audio data processing device can obtain the target audio data in the target recognition scenario from the audio data acquisition device. Specifically, a communication connection is established between the audio data acquisition device and the audio data processing device in the audio data processing system. The audio data acquisition device in the target recognition scenario can collect audio data in the target recognition scenario to obtain target audio data. Then, the audio data processing device can obtain the target audio data collected in the target recognition scenario through the communication connection between the audio data acquisition device and the audio data processing device. For example, in a game scenario, the game audio data of the user in the game interaction process can be collected by an audio data acquisition device (such as a game terminal device), and then the audio data acquisition device can use the game audio data of the user in the game interaction process as the target audio data. The audio data processing device can obtain the game audio data in the game scenario, such as Figure 3 The target audio data shown in .

[0050] The scenario-wide lexicon is used to convert audio data into text data. The scenario-wide lexicon can include M scenario-wide terms, where M is a positive integer. The scenario-wide lexicon can be a general dictionary, such as the Modern Chinese Grammar Information Dictionary, the Tsinghua University Dictionary, or the HowNet Dictionary. It can also be a custom dictionary.

[0051] S202 , performing text conversion on the target audio data based on M common scene words to obtain converted text data.

[0052] In one embodiment, the audio data processing device may perform an audio data segmentation operation on the target audio data to obtain L audio segments, where L is a positive integer and L is less than or equal to M. Then, the L audio segments are respectively converted into text based on the M scene-general words, and the scene-general words corresponding to each of the L audio segments are determined, thereby determining the converted text data, which includes the L scene-general words. Figure 3In the example shown, the target audio data can be divided into audio segment L1, audio segment L2, audio segment L3, and audio segment L4. Based on the M scene-general words in the scene-general vocabulary, audio segments L1, L2, L3, and L4 are respectively converted into text, and scene-general word 1 (he) corresponding to audio segment L1, scene-general word 2 (tax refund) corresponding to audio segment L2, scene-general word 3 (give) corresponding to audio segment L3, and scene-general word 4 (you) corresponding to audio segment L4 can be obtained. Then, the audio data processing device can determine that the converted text data is "He refunds the tax to you" based on scene-general word 1, scene-general word 2, scene-general word 3, and scene-general word 4.

[0053] The audio data processing device may determine the scene-general word corresponding to the audio segment in various ways. Optionally, for any one of the L audio segments, the audio data processing device may obtain the pinyin of the word indicated by the audio segment from the audio segment, and then obtain a scene-general word with a similar pinyin to the word indicated by the audio segment from the M scene-general words as the scene-general word corresponding to the audio segment.

[0054] Optionally, for any one of the L audio segments, the audio data processing device can obtain a scene-common word selection instruction for the M scene-common words; and determine the scene-common word selected from the M scene-common words as the scene-common word corresponding to the audio segment according to the scene-common word selection instruction.

[0055] S203, obtaining a target terminology lexicon associated with the target recognition scenario, and obtaining abnormal conversion words in the converted text data; the target terminology lexicon includes N terminology words in the target recognition scenario, where N is a positive integer.

[0056] A recognition scenario is associated with at least one terminology lexicon, which includes terminology words specifically used in the recognition scenario. The terminology words are generally not words included in a conventional dictionary (such as the aforementioned scenario-wide lexicon) (such as scenario-wide general terms in the scenario-wide lexicon).

[0057] Specifically, in the target recognition scenario, the audio data processing device can obtain a terminology library associated with the target recognition scenario, and refer to the terminology library associated with the target recognition scenario as a target terminology library. The target terminology library includes N terminology words in the target recognition scenario, where N is a positive integer and the specific value of N is determined according to the actual application scenario. Figure 3 As shown, when the target recognition scene is a game scene, a target term vocabulary associated with the game scene can be obtained, and the target term vocabulary includes: "retreat water", "step on wolf" and "share edge".

[0058] In order to obtain the target terminology library associated with the target recognition scenario, in one embodiment, the audio data processing device can directly obtain the target terminology library associated with the target recognition scenario. Specifically, the audio data processing device can pre-store terminology libraries under P recognition scenarios. The audio data processing device can select the target terminology library associated with the target recognition scenario from the P recognition scenarios based on the target recognition scenario, where P is a positive integer. In another embodiment, in the target recognition scenario, the audio data processing device can obtain the target terminology library associated with the target recognition scenario from the blockchain. The audio data processing device can send a query request carrying a scene identifier to the blockchain node, and the blockchain node can receive the query request carrying the scene identifier, select the target terminology library associated with the target recognition scenario from the blockchain based on the scene identifier, and send the target terminology library to the audio data processing device.

[0059] In one embodiment, the target terminology vocabulary associated with the target recognition scenario is pre-generated. The specific steps of the audio data processing device generating the target terminology vocabulary can be referred to the specific description of the subsequent related embodiments. No further details will be given here.

[0060] Among them, since the scene-general vocabulary cannot cover the target terminology in the target recognition scenario; for example, the scene-general vocabulary cannot cover the terminology in the dialect (such as Longmenzhen); for another example, the scene-general vocabulary cannot cover the terminology in the game scene (such as retreat), etc., when the target audio data is converted into text based on M scene-general words, it may not be possible to accurately convert the target audio data containing terminology, so there may be abnormal conversion words in the obtained converted text data.

[0061] The audio data processing device can obtain abnormal conversion words in the converted text data in different ways. In one embodiment, the audio data processing device can obtain L scene-general words contained in the converted text data; and obtain the text conversion confidence of each scene-general word in the converted text data in the L scene-general words, and then determine the scene-general words in the L scene-general words whose text conversion confidence is less than the confidence threshold as abnormal conversion words. Specifically, the converted text data may include L scene-general words, that is, the L scene-general words constitute the converted text data. When converting the target audio data into converted text data, each scene-general word in the L scene-general words can have a text conversion confidence. In which, the target audio data can be converted into converted text data by a text conversion model. When the text conversion model converts the target audio data into converted text data, it can obtain the conversion probability of each scene-general word in the converted text data. The higher the conversion probability, the higher the confidence that the corresponding scene-general word is converted. The lower the conversion probability, the lower the accuracy of the conversion of the corresponding scene-general word. The conversion probability can be used as the above-mentioned text conversion confidence. Therefore, the text conversion confidence of each scene-general word in the above-mentioned L scene-general words can be compared with the confidence threshold (which can be set according to the actual application scenario), and based on the comparison result, it is determined whether each scene-general word in the L scene-general words is an abnormal conversion word, that is, for scene-general word A in the L scene-general words, if the text conversion confidence of scene-general word A is greater than or equal to the confidence threshold, then scene-general word A is a normal conversion word; if the text conversion confidence of scene-general word A is less than the confidence threshold, then scene-general word A is an abnormal conversion word.

[0062] In another embodiment, the audio data processing device may obtain L scene-general words contained in the converted text data, and obtain a word selection instruction for the L scene-general words. Then, according to the word selection instruction, the scene-general word selected from the L scene-general words is determined as an abnormal conversion word. Among them, the word selection instruction may include multiple types. The word selection instruction is generated based on a user operation in the user interface, and the user operation may include one or more operations such as clicking, sliding, long pressing, and double-clicking in the user interface.

[0063] For example, assume that L scene-general words are scene-general word A1, scene-general word A2, and scene-general word A3. The audio data processing device can obtain a word selection instruction for these three scene-general words and determine the abnormal conversion words among the three scene-general words according to the word selection instruction. For example, if the word selection instruction is used to indicate the selection of scene-general word A1, the audio data processing device determines that scene-general word A1 is an abnormal conversion word; if the word selection instruction is used to indicate the selection of scene-general word A2, the audio data processing device determines that scene-general word A2 is an abnormal conversion word; if the word selection instruction is used to indicate the selection of scene-general word A3, the audio data processing device determines that scene-general word A3 is an abnormal conversion word; if the word selection instruction is used to indicate the selection of scene-general word A1 and scene-general word A2, the audio data processing device determines that scene-general word A1 and scene-general word A2 are abnormal conversion words. If the word selection instruction is used to indicate the selection of scene-general word A1 and scene-general word A3, the audio data processing device determines that scene-general word A1 and scene-general word A3 are abnormal conversion words. If the word selection instruction is used to instruct the selection of scene-general words A2 and scene-general words A3, the audio data processing device determines that scene-general words A2 and scene-general words A3 are abnormal conversion words. If the word selection instruction is used to instruct the selection of scene-general words A1, scene-general words A2, and scene-general words A3, the audio data processing device determines that scene-general words A1, scene-general words A2, and scene-general words A3 are abnormal conversion words.

[0064] There may be multiple abnormal conversion words, and one abnormal conversion word may correspond to one target term.

[0065] S204 , obtaining target terminology corresponding to the abnormal conversion term from the N terminology, replacing the abnormal conversion term in the converted text data with the target terminology, and obtaining corrected text data of the converted text data.

[0066] In one embodiment, the audio data processing device can obtain the pinyin of the abnormal conversion word according to the target audio data. Since this pinyin is obtained during the process of speech recognition of the target audio data followed by text conversion, it may be different from the actual pinyin of the abnormal conversion word itself. Optionally, the audio data processing device can obtain the term word with a similar pinyin to the abnormal conversion word from N term words according to the pinyin of the abnormal conversion word obtained through the target audio data, and use it as the target term word, and replace the abnormal conversion word in the converted text data with the target term word to obtain the corrected text data of the converted text data, that is, the corrected text data is the text data obtained after replacing the abnormal conversion word in the converted text data with the target term word. Among them, the audio data processing device can use a pinyin conversion tool to determine the pinyin of each word.

[0067] Optionally, the target term word can also be obtained through the actual pinyin of the abnormal conversion word itself. For example, in the target term word library, the term word with a pinyin similar to the actual pinyin of the abnormal conversion word itself can be used as the target term word corresponding to the abnormal conversion word.

[0068] Still following Figure 3 the example shown, assuming that "tax refund" in the converted text data is an abnormal conversion word, then the audio data processing device can determine that the pinyin of the abnormal conversion word "tax refund" is "tuishui" according to the target audio data. As described above, the "water discharge", "step on the wolf", and "common edge" included in the target term word library, then the audio data processing device can determine that the pinyin of the term word "water discharge" is "tuishui", determine that the pinyin of the term word "step on the wolf" is "cailang", and determine that the pinyin of the term word "common edge" is "gongbian". The audio data processing device can use the term word "water discharge" with a similar pinyin to the abnormal conversion word "tax refund" as the target term word, and replace the abnormal conversion word "tax refund" in the converted text data "He refunds the tax to you" with the target term word "water discharge" to obtain the corrected text data of the converted text data "He discharges water to you".

[0069] In one embodiment, the audio data processing device can determine the target audio data and the corrected text data as training sample data for the initial text conversion model; and train the initial text conversion model based on the training sample data to obtain a target text conversion model, wherein the target text conversion model is used to convert audio type data into text type data. In an embodiment of the present application, since the converted text data of the target audio data in the target recognition scenario can be corrected by the target terminology vocabulary associated with the target recognition scenario, more accurate text data of the target audio data (such as corrected text data) can be obtained. Therefore, when the target audio data and the corrected text data are used as training sample data for the initial text conversion model, the accuracy of the training sample data is higher, and thus the accuracy of the target text conversion model obtained by training with the more accurate training sample data will also be higher, and the target text conversion model obtained by subsequent training can also more accurately convert the audio type data that needs to be converted into text type data.

[0070] In an embodiment of the present application, when an audio data processing device obtains target audio data collected in a target recognition scenario, it can first perform text conversion on the target audio data based on the M scene-general words included in the scene-general word library to obtain converted text data, and then correct the abnormal conversion words in the converted text data according to the target terminology word library associated with the target recognition scenario, and replace the abnormal conversion words in the converted text data with the target terminology words in the target terminology word library to obtain corrected text data of the converted text data. Since the target terminology word library associated with the target recognition scenario is used to correct the converted text data of the target audio data in the target recognition scenario, compared with the solution in which the annotator manually performs text conversion on the target audio data in the target recognition scenario, there is no need to manually query the terminology words in the target recognition scenario, which can effectively save human resources and improve the processing efficiency of the target audio data; moreover, it is not limited by the ability of the audio data processing personnel, and the target terminology words corresponding to the target audio data can be accurately obtained through the target terminology word library, thereby reducing the error rate in the text data (corrected text data) and ensuring the accuracy of the text data (corrected text data) corresponding to the target audio data.

[0071] See above Figure 2 It can be seen from the relevant description of the embodiment of the method shown that Figure 2 The audio data processing method shown can modify the converted text data of the target audio data through the target terminology word library. The target terminology word library can be generated based on the scene interaction text content associated with the target recognition scene in the information interaction platform. Figure 4 As shown in FIG, the embodiment of the present application proposes a flowchart of another audio data processing method. Figure 4As shown in FIG. 4 , the audio data processing method may include steps S401 to S403:

[0072] S401: Acquire scene interaction text content associated with the target recognition scene in the information interaction platform.

[0073] An information exchange platform refers to any platform that supports online communication between users. Examples include search engines and instant messaging platforms. Within an information exchange platform, multiple users can discuss or exchange ideas on a single topic or multiple topics. For example, multiple users can discuss and exchange ideas in a search engine's discussion area.

[0074] In one embodiment, because the information interaction platform disseminates a large amount of data, it can disseminate scene interaction text content from various recognition scenarios. For each recognition scenario, the scene interaction text content typically includes terminology specific to that recognition scenario. Therefore, the audio data processing device can obtain scene interaction text content associated with the target recognition scenario from the information interaction platform.

[0075] Among them, the audio data processing device can obtain the scene interaction text content associated with the target recognition scene in the information interaction platform based on the web crawler technology. Specifically, the audio data processing device can enter the access URL of the information interaction platform in the web crawler framework, and enter the search field associated with the target recognition scene in the information interaction platform, so that the audio data processing device can view the user interface associated with the target recognition scene in the information interaction platform. Figure 5As shown, when the information interaction platform is a search engine and the target identification scene is a Werewolf game scene, the audio data processing device can enter the search field "Werewolf" under the Werewolf game scene in the search engine to view pages containing communication content related to the Werewolf game scene. Then, a web crawling framework written in a programming language uses an article extractor to initiate a request to the Uniform Resource Locator (URL) in the information interaction platform to obtain the scene interaction text content associated with the target identification scene in the Hypertext Markup Language (HTML) page corresponding to the URL of the information interaction platform. It should be noted that the audio data processing device can obtain the scene interaction text content associated with the target identification scene through the page content in the HTML page. If the page content is empty, the audio data processing device initiates a request to the next URL of the information interaction platform to obtain the scene interaction text content associated with the target identification scene in the HTML page corresponding to the next URL of the information interaction platform. When the page content is not empty, the text data in the page content is preprocessed (such as cleaning the punctuation marks and Unicode encoding areas in the text data of the page content, retaining the text data of the page content (such as numbers, Chinese and English) to obtain the scene interaction text content associated with the target recognition scene.

[0076] S402 : Acquire scene-general words in the scene-interaction text content according to the scene-general word library, and filter the scene-general words in the scene-interaction text content to obtain first filtered text data.

[0077] In one embodiment, since terminology is usually not included in the scene-general vocabulary, the scene-general words in the scene interaction text content can be filtered based on the scene-general vocabulary to obtain first filtered text data that does not contain the scene-general words in the scene-general vocabulary.

[0078] Specifically, the audio data processing device can perform natural language processing (NLP) operations on the scene interaction text content, identify the scene-general words in the scene interaction text content that belong to the scene-general vocabulary, filter the scene-general words in the scene interaction text content, and obtain the first filtered text data. Optionally, in order to facilitate the distinction between scene-general words in the scene interaction text content, after the audio data processing device determines the scene-general words in the scene interaction text content according to the scene-general vocabulary, the audio data processing device can output prompt information. The prompt information here can be display information or voice information. Display information can refer to the audio data processing device identifying the scene-general words in the scene interaction text content through color markings (including highlighting or gray display, etc.) on the screen; for example, Figure 6 As shown in 601, the interactive text content of this scene is "He retreated the water to give you the police badge, and only others can step on you. The two wolves are stepping on each other. If you go to deal with him, he is a wolf, aren't you asking for a fight?" According to the scene common vocabulary, the scene common words included in the interactive text content of this scene are "police badge", "only", "others", "you go", "he is" and "is not". Figure 6 In step 602, the gray content in the scene interaction text content is the above-mentioned scene common words. Voice information can refer to identifying the scene common words in the scene interaction text content through voice broadcast, such as a voice prompt carrying the scene common words.

[0079] After identifying the common scene words in the scene interaction text content, the audio data processing device can filter the common scene words in the scene interaction text content to obtain the first filtered text data. Figure 6 In the example shown, the audio data processing device identifies the common scene words "badge", "only", "others", "you go", "he is" and "is not" in the scene interaction text content "He retreats the water to give you the police badge. Only others can step on you. The two wolves step on each other. If you go to check if he is a wolf, aren't you asking for a fight?" and filters them to obtain the following: Figure 6 The first filtered text data shown in 603 includes "He retreats and gives you a break", "Step on you two wolves step on wolves on the same side", "plate", "wolf" and "Are you looking for a fight?"

[0080] S403: Generate a target term vocabulary based on the first filtered text data.

[0081] In one embodiment, the audio data processing device can obtain the connection attribute words in the first filtered text data, and then filter the connection attribute words in the first filtered text data to obtain the second filtered text data; a target term library is generated according to the second filtered text data. Since there may still be some connection attribute words indicating grammatical formats in the filtered first filtered text content, and term words are usually not connection attribute words, in order to obtain term words, it is also necessary to filter the connection attribute words in the first filtered text data to obtain the second filtered text data that does not contain connection attribute words, and a target term library is generated from this second filtered text data. Among them, the connection attribute words may include one or more of the following: pronouns, auxiliary words, numeral-classifiers, and prepositions.

[0082] For example, still continuing with the above Figure 6 example, the audio data processing device can identify the pronouns in the first filtered text data and filter the pronouns in the first filtered text data. For pronouns, for example, the Figure 6 pronouns "he" and "you" in the first filtered text data shown in 603, such as "He draws water for you to yield", "Step on you two wolves step on wolves and be on the same side", "Discuss", "Wolf", "Should I be beaten?", are filtered, and text data without pronouns can be obtained. The obtained text data includes "draws water for", "yield", "step on", "two wolves step on wolves and be on the same side", "discuss", "wolf", and "should I be beaten?".

[0083] For auxiliary words, for example, the Figure 6 auxiliary word "ma" in the first filtered text data shown in 603, such as "He draws water for you to yield", "Step on you two wolves step on wolves and be on the same side", "Discuss", "Wolf", "Should I be beaten?", can be filtered, and text data without auxiliary words can be obtained. The obtained text data includes "He draws water for you to yield", "Step on you two wolves step on wolves and be on the same side", "Discuss", "Wolf", and "Should I be beaten?".

[0084] For numeral-classifiers, for example, the Figure 6 [[ID=1...]]numeral-classifier "two" in the first filtered text data shown in 603, such as "He draws water for you to yield", "Step on you two wolves step on wolves and be on the same side", "Discuss", "Wolf", "Should I be beaten?", can be filtered, and text data without numeral-classifiers can be obtained, including "He draws water for you to yield", "Step on you", "wolves step on wolves and be on the same side", "discuss", "wolf", and "should I be beaten?". <...000206>

[0085] For prepositions, for example, the Figure 6 prepositions "for", "to", and "seek" in the first filtered text data shown in 603, such as "He draws water for you to yield", "Step on you two wolves step on wolves and be on the same side", "Discuss", "Wolf", "Should I be beaten?", can be filtered, and text data without prepositions can be obtained, including "He draws water", "you", "Step on you two wolves step on wolves and be on the same side", "discuss", "wolf", and "beaten?". [[ID=...]]

[0086] After filtering the connection attribute words in the first filtered text data, the second filtered text data can be obtained. The words included in the second filtered text are used as term words. For example, Figure 6 in the first filtered text data of 603 in Figure 6 , pronouns, auxiliary words, numeral-classifier words, and prepositions can be filtered, and the second filtered text data is "drawdown", "step on", "wolves step on wolves and share a side", "disc", "wolf", "beat", and the second filtered text also includes three words. Therefore, the audio data processing device can obtain three term words "drawdown", "step on wolf", and "share a side" according to the second filtered text, such as Figure 6 the words within the rectangular box in 605 of Figure 6 . Therefore, the generated target term library can include three term words "drawdown", "step on wolf", and "share a side".

[0087] In another embodiment, the audio data processing device can obtain a selection instruction for the first filtered text content. Then, determine the term words in the first filtered text content according to the selection instruction, and generate a target term library based on the term words. For example, when the first filtered text data is "He gives you a drawdown", "step on you two, wolves step on wolves and share a side", "disc", "wolf", "looking for a beating", if the selection instruction indicates that the three words "drawdown", "step on wolf", and "share a side" in the first filtered text content are term words, then these three words are used as term words, and a target term library is generated.

[0088] In one embodiment, the audio data processing device can encapsulate the target term library in the target recognition scenario into a block and store the block on the blockchain. The blockchain is a chained data structure formed by combining data blocks in sequence according to time order, and a distributed ledger that ensures data cannot be tampered with and forged in a cryptographic manner. Multiple independent distributed nodes save the same records. Blockchain technology has achieved decentralization and has become the cornerstone of reliable digital asset storage, transfer, and trading.

[0089] Taking Figure 7 the structural schematic diagram of the blockchain shown as an example, when writing the target term library in the target recognition scenario into the blockchain, the target term library in the target recognition scenario can be encapsulated into a block and added to the end of the existing blockchain. Through the consensus algorithm, it is ensured that the newly added blocks for each node are exactly the same. Each block records the target term library and also contains the hash value of the previous block. All blocks save the hash value in the previous block in this way and are connected in sequence to form the blockchain. The hash value of the previous block will be stored in the block header of the next block in the blockchain. When the target term library in the previous block changes, the hash value of this block will also change accordingly. Therefore, the target term library uploaded to the blockchain is difficult to be tampered with, improving the reliability of the data.

[0090] In an embodiment of the present application, the audio data processing device can obtain the scene interaction text content associated with the target recognition scene in the information interaction platform, so that the target terminology vocabulary associated with the target recognition scene can be constructed based on the scene interaction content associated with the target recognition scene. When the audio data processing device obtains the converted text data of the target audio data in the target recognition scene based on the scene general vocabulary, the target terminology vocabulary associated with the target recognition scene can be used to correct the abnormal conversion words in the converted text data, and more accurate text data of the target audio data (i.e., corrected text data) can be obtained. Since the target terminology vocabulary constructed in this way can cover the terminology words in the target recognition scene to a greater extent, when the target terminology vocabulary is used to correct the abnormal conversion words in the converted text data, the error rate in the text data of the target audio data can be effectively reduced, and the accuracy of the text data corresponding to the target audio data can be guaranteed.

[0091] Based on the description of the above-mentioned audio data processing method embodiment, the present application embodiment also discloses an audio data processing device, which can be a computer program (including program code) running on the above-mentioned audio data processing device. The audio data processing device can execute Figure 2 or Figure 4 See the method shown in Figure 8 , the audio data processing device can run the following units:

[0092] The acquisition unit 801 is configured to acquire target audio data collected in a target recognition scenario and acquire a scenario-general vocabulary; the scenario-general vocabulary includes M scenario-general words, where M is a positive integer;

[0093] A conversion unit 802 is configured to perform text conversion on the target audio data based on the M scene-common words to obtain converted text data;

[0094] The acquisition unit 801 is further configured to acquire a target terminology lexicon associated with the target recognition scenario and to acquire abnormal conversion words in the converted text data; the target terminology lexicon includes N terminology words in the target recognition scenario, where N is a positive integer;

[0095] The replacing unit 803 is configured to obtain a target term corresponding to the abnormal conversion term from the N term words, and replace the abnormal conversion term in the converted text data with the target term word to obtain corrected text data of the converted text data.

[0096] In one embodiment, the acquisition unit 801 acquires abnormal conversion words in the converted text data, including:

[0097] Obtain L common scene words contained in the converted text data; L is a positive integer, and L is less than or equal to M;

[0098] Obtaining a text conversion confidence of each of the L scene-general words in the converted text data;

[0099] The scene-general words whose text conversion confidences are less than the confidence threshold among the L scene-general words are determined as abnormal conversion words.

[0100] In another embodiment, the acquisition unit 801 acquires abnormal conversion words in the converted text data, including:

[0101] Obtain L common scene words contained in the converted text data; L is a positive integer, and L is less than or equal to M;

[0102] Obtaining word selection instructions for L common words in different scenarios;

[0103] According to the word selection instruction, the scene-general word selected from the L scene-general words is determined as the abnormal conversion word.

[0104] In another embodiment, the audio data processing device further includes a filtering unit 804, the filtering unit 804 being configured to obtain scene interaction text content associated with the target recognition scene in the information interaction platform;

[0105] Acquire scene-general words in the scene interaction text content according to the scene-general word library, and filter the scene-general words in the scene interaction text content to obtain first filtered text data;

[0106] A target term vocabulary is generated based on the first filtered text data.

[0107] In another embodiment, the filtering unit 804 generates a target term vocabulary based on the first filtered text data, including:

[0108] Obtaining connection attribute words in the first filtered text data;

[0109] Filtering the connection attribute words in the first filtering text data to obtain second filtering text data;

[0110] A target term vocabulary is generated based on the second filtered text data.

[0111] In another embodiment, the replacing unit 803 obtains the target term corresponding to the abnormal conversion term from the N term words, including:

[0112] Obtain the pinyin of the abnormal conversion words according to the target audio data;

[0113] According to the pinyin corresponding to the abnormal conversion word, a term word with a similar pinyin to the abnormal conversion word is obtained from the N term words as the target term word.

[0114] In another embodiment, the audio data processing apparatus further includes a training unit 805, wherein the training unit 805 is configured to determine the target audio data and the modified text data as training sample data for the initial text conversion model;

[0115] An initial text conversion model is trained based on the training sample data to obtain a target text conversion model; the target text conversion model is used to convert audio type data into text type data.

[0116] According to one embodiment of the present application, Figure 2 or Figure 4 Each step involved in the method shown can be performed by Figure 8 The various units in the audio data processing device shown are executed. For example, Figure 2 Steps S201 and S203 shown are represented by Figure 8 The acquisition unit 801 shown in FIG is executed, and step S202 is performed by Figure 8 The conversion unit 802 shown in FIG is executed, and step S204 is performed by Figure 8 803 is used to perform the replacement. Figure 4 Steps S401-S403 are composed of Figure 8 The filtering unit 804 shown in FIG.

[0117] According to another embodiment of the present application, Figure 8 The various units in the audio data processing device shown can be individually or all combined into one or several other units to constitute, or one (or some) of the units can be further divided into multiple smaller units in function to constitute, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of a unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, other units can also be included based on the audio data processing device. In actual applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.

[0118] According to another embodiment of the present application, the system can be implemented by including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM) and other processing elements and storage elements. For example, a general computing device such as a computer can execute the following operations: Figure 2 or Figure 4A computer program (including program code) for each step involved in the corresponding method shown in Figure 8 The audio data processing device shown in the figure and the audio data processing method of the embodiment of the present application are implemented. The computer program can be recorded on a computer-readable recording medium, for example, and loaded into the above-mentioned audio data processing device through the computer-readable recording medium and run therein.

[0119] In an embodiment of the present application, when the acquisition unit acquires the target audio data collected in the target recognition scenario, the conversion unit can perform text conversion on the target audio data based on the M scene-general words included in the scene-general vocabulary to obtain converted text data, and then the acquisition unit can acquire the target terminology vocabulary associated with the target recognition scenario, and acquire abnormal conversion words in the converted text data; the replacement unit then corrects the abnormal conversion words in the converted text data according to the target terminology vocabulary associated with the target recognition scenario, and replaces the abnormal conversion words in the converted text data with the target terminology words in the target terminology vocabulary to obtain corrected text data of the converted text data. Since the target terminology vocabulary associated with the target recognition scenario is used to correct the converted text data of the target audio data in the target recognition scenario, compared with the solution in which the labeling personnel manually convert the target audio data in the target recognition scenario into text, there is no need to manually query the terminology in the target recognition scenario, which can effectively save human resources and improve the efficiency of target audio data processing; moreover, it is not limited by the ability of the audio data processing personnel. The target terminology vocabulary can be used to accurately obtain the corresponding target terminology in the target audio data, reduce the error rate in the text data (corrected text data), and ensure the accuracy of the text data (corrected text data) corresponding to the target audio data.

[0120] Based on the description of the above audio data processing method embodiment, the present application embodiment also discloses an audio data processing device. Figure 9 The audio data processing device at least includes a processor 901, an input interface 902, an output interface 903 and a computer storage medium 904 which can be connected via a bus or other means.

[0121] The computer storage medium 904 is a memory device in the audio data processing device, which is used to store programs and data. It is understandable that the computer storage medium 904 here can include both the built-in storage medium of the audio data processing device and, of course, the extended storage medium supported by the audio data processing device. The computer storage medium 904 provides a storage space, which stores the operating system of the audio data processing device. In addition, one or more instructions suitable for being loaded and executed by the processor 901 are also stored in the storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer storage medium here can be a high-speed RAM memory; optionally, it can also be at least one computer storage medium away from the aforementioned processor. The processor can be called a central processing unit (CPU), which is the core and control center of the audio data processing device, suitable for implementing one or more instructions, specifically loading and executing one or more instructions to achieve the corresponding method flow or function.

[0122] In one embodiment, the processor 901 may load and execute one or more instructions stored in the computer storage medium 904 to implement the following operations: Figure 2 or Figure 4 In the specific implementation of the steps involved in the corresponding method shown in FIG, one or more instructions in the computer storage medium 904 are loaded by the processor 901 and the following steps are executed:

[0123] Obtain target audio data collected in a target recognition scenario and obtain a scenario-general vocabulary; the scenario-general vocabulary includes M scenario-general words, where M is a positive integer;

[0124] Perform text conversion on the target audio data based on M common scene words to obtain converted text data;

[0125] Obtain a target terminology lexicon associated with a target recognition scenario, and obtain abnormal conversion words in the converted text data; the target terminology lexicon includes N terminology words in the target recognition scenario, where N is a positive integer;

[0126] A target term corresponding to the abnormal conversion term is obtained from the N term words, and the abnormal conversion term in the conversion text data is replaced with the target term word to obtain corrected text data of the conversion text data.

[0127] In one embodiment, the processor 901 obtains abnormal conversion words in the converted text data, including:

[0128] Obtain L common scene words contained in the converted text data; L is a positive integer, and L is less than or equal to M;

[0129] Obtaining a text conversion confidence of each of the L scene-general words in the converted text data;

[0130] The scene-general words whose text conversion confidences are less than the confidence threshold among the L scene-general words are determined as abnormal conversion words.

[0131] In another embodiment, the processor 901 obtains abnormal conversion words in the converted text data, including:

[0132] Obtain L common scene words contained in the converted text data; L is a positive integer, and L is less than or equal to M;

[0133] Obtaining word selection instructions for L common words in different scenarios;

[0134] According to the word selection instruction, the scene-general word selected from the L scene-general words is determined as the abnormal conversion word.

[0135] In another embodiment, the processor 901 is further configured to:

[0136] Obtaining scene interaction text content associated with the target recognition scene in the information interaction platform;

[0137] Acquire scene-general words in the scene interaction text content according to the scene-general word library, and filter the scene-general words in the scene interaction text content to obtain first filtered text data;

[0138] A target term vocabulary is generated based on the first filtered text data.

[0139] In another embodiment, the processor 901 generates a target term vocabulary based on the first filtered text data, including:

[0140] Obtaining connection attribute words in the first filtered text data;

[0141] Filtering the connection attribute words in the first filtering text data to obtain second filtering text data;

[0142] A target term vocabulary is generated based on the second filtered text data.

[0143] In another embodiment, the processor 901 obtains a target term corresponding to the abnormal conversion term from the N term words, including:

[0144] Obtain the pinyin of the abnormal conversion words according to the target audio data;

[0145] According to the pinyin corresponding to the abnormal conversion word, a term word with a similar pinyin to the abnormal conversion word is obtained from the N term words as the target term word.

[0146] In another embodiment, the processor 901 is further configured to: determine the target audio data and the modified text data as training sample data for the initial text conversion model;

[0147] An initial text conversion model is trained based on the training sample data to obtain a target text conversion model; the target text conversion model is used to convert audio type data into text type data.

[0148] In an embodiment of the present application, when an audio data processing device obtains target audio data collected in a target recognition scenario, it can first perform text conversion on the target audio data based on the M scene-general words included in the scene-general word library to obtain converted text data, and then correct the abnormal conversion words in the converted text data according to the target terminology word library associated with the target recognition scenario, and replace the abnormal conversion words in the converted text data with the target terminology words in the target terminology word library to obtain corrected text data of the converted text data. Since the target terminology word library associated with the target recognition scenario is used to correct the converted text data of the target audio data in the target recognition scenario, compared with the solution in which the annotator manually performs text conversion on the target audio data in the target recognition scenario, there is no need to manually query the terminology words in the target recognition scenario, which can effectively save human resources and improve the processing efficiency of the target audio data; moreover, it is not limited by the ability of the audio data processing personnel, and the target terminology words corresponding to the target audio data can be accurately obtained through the target terminology word library, thereby reducing the error rate in the text data (corrected text data) and ensuring the accuracy of the text data (corrected text data) corresponding to the target audio data.

[0149] It should be noted that the embodiments of the present application further provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the audio data processing device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the audio data processing device performs the above-mentioned audio data processing method embodiment. Figure 2 or Figure 4 The steps performed in .

[0150] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made in accordance with the claims of this application are still within the scope covered by the application.

Claims

1. A method for processing audio data, characterized in that: include: Obtaining scene interaction text content associated with the target recognition scene in the information interaction platform; Acquire scene-general words in the scene interaction text content according to a scene-general word library, and filter the scene-general words in the scene interaction text content to obtain first filtered text data; the scene-general word library includes M scene-general words, where M is a positive integer; Obtaining connection attribute words in the first filtered text data; Filtering the connection attribute words in the first filtering text data to obtain second filtering text data; generating a target terminology lexicon associated with the target recognition scenario based on the second filtered text data; the target terminology lexicon includes N terminology words in the target recognition scenario, where N is a positive integer; Acquire target audio data collected in a target recognition scenario, and obtain a common vocabulary for the scenario; Performing text conversion on the target audio data based on the M scene-general words using a preset text conversion model to obtain converted text data, wherein the converted text data includes a conversion probability of each of the L scene-general words corresponding to the target audio data, where L is a positive integer and is less than or equal to M; The conversion probability of each of the L common scene words is used as the corresponding text conversion confidence; Determining the scene-general words whose text conversion confidences are less than a confidence threshold among the L scene-general words as abnormal conversion words in the converted text data; Obtaining the target term database; Obtaining a target term corresponding to the abnormal conversion term from the N term words, and replacing the abnormal conversion term in the converted text data with the target term word to obtain corrected text data of the converted text data; Determining the target audio data and the revised text data as training sample data for an initial text conversion model; The initial text conversion model is trained based on the training sample data to obtain a target text conversion model; the target text conversion model is used to convert audio type data into text type data.

2. The method according to claim 1, wherein The method further comprises: Obtain L common scene words contained in the converted text data; L is a positive integer, and L is less than or equal to M; Obtaining a word selection instruction for the L common words in the scenarios; According to the word selection instruction, the scene-general word selected from the L scene-general words is determined as the abnormal conversion word.

3. The method according to claim 1, wherein The step of obtaining a target term corresponding to the abnormal conversion term from the N terminologies includes: Acquire the pinyin of the word corresponding to the abnormal conversion word according to the target audio data; According to the pinyin corresponding to the abnormal conversion word, a term word having a similar pinyin to the abnormal conversion word is obtained from the N term words as the target term word.

4. An audio data processing device, characterized in that: include: A filtering unit, configured to obtain scene interaction text content associated with a target recognition scene in the information interaction platform; Obtaining scene-general words in the scene interaction text content according to a scene-general vocabulary, filtering the scene-general words in the scene interaction text content to obtain first filtered text data; obtaining connection attribute words in the first filtered text data; filtering the connection attribute words in the first filtered text data to obtain second filtered text data; generating a target term vocabulary associated with the target recognition scene according to the second filtered text data; An acquisition unit, configured to acquire target audio data collected in a target recognition scenario and acquire a scenario-wide vocabulary; The scene-general vocabulary includes M scene-general words, where M is a positive integer; a conversion unit, configured to perform text conversion on the target audio data based on the M scene-general words using a preset text conversion model to obtain converted text data, wherein the converted text data includes a conversion probability of each of the L scene-general words corresponding to the target audio data, where L is a positive integer and is less than or equal to M; The acquisition unit is further configured to acquire a target terminology lexicon associated with the target recognition scenario and to acquire abnormal conversion words in the converted text data; the target terminology lexicon includes N terminology words in the target recognition scenario, where N is a positive integer; a replacing unit, configured to obtain a target term corresponding to the abnormal conversion term from the N term words, and replace the abnormal conversion term in the converted text data with the target term word to obtain corrected text data of the converted text data; The acquisition unit is specifically configured to, when acquiring the abnormal conversion words in the converted text data, use the conversion probability of each of the L scene-general words as the corresponding text conversion confidence; and determine the scene-general words among the L scene-general words whose text conversion confidence is less than the confidence threshold as the abnormal conversion words in the converted text data; A training unit is used to determine the target audio data and the corrected text data as training sample data for an initial text conversion model; train the initial text conversion model based on the training sample data to obtain a target text conversion model; and the target text conversion model is used to convert audio type data into text type data.

5. An audio data processing device, comprising an input interface and an output interface, characterized in that: Also includes: a processor adapted to implement one or more instructions; as well as, A computer storage medium storing one or more instructions, wherein the one or more instructions are suitable for being loaded by the processor and executing the audio data processing method according to any one of claims 1 to 3.

6. A computer storage medium, characterized in that The computer storage medium stores one or more instructions, and the one or more instructions are suitable for being loaded by a processor and executing the audio data processing method according to any one of claims 1 to 3.

7. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium and are suitable for being read and executed by a processor of an audio data processing device, so that the audio data processing device implements the audio data processing method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Text processing method and device

    CN112395863A

  • Speech to text conversion of non-supported technical language

    CN113678196A

  • Key term extraction

    US20070022115A1