Keyword extraction method and related device

By automatically comparing the transcribed text with the original text, extracting the keywords that were wrong in the translation engine, solving the problem of time-consuming and labor-intensive manually configuring keywords, improving the accuracy of the transcription and saving labor costs.

CN120199233APending Publication Date: 2025-06-24ANHUI IFLYREC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510410780.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art relies on manual configuration of keywords when the transfer engine transcribing audio into text, which leads to time-consuming and labor-intensive, easy to miss or configuration errors, affecting the accuracy of the translation.

Method used

Provide a keyword extraction method, by obtaining original text and audio data, using the target transfer engine to process the audio data into transcribed text, and comparing the transcribed text with the original text, and automatically extracting the target keywords.

Benefits of technology

It realizes automatic and accurate extraction of target keywords from the original text, improves the accuracy of transcription recognition, reduces manual participation, and saves time and labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199233A_ABST
    Figure CN120199233A_ABST
Patent Text Reader

Abstract

The invention discloses a keyword extraction method and a related device, and relates to the technical field of computers, and the keyword extraction method comprises the following steps: obtaining an original text and audio data of the original text, processing the audio data into a transliteration text through a target transliteration engine, and comparing the transliteration text with the original text to obtain a target keyword. According to the method and the device, the target transfer engine can perform real audio transfer, the target keyword with the transfer error of the target transfer engine can be extracted from the original text according to the transfer text and the original text, and the accuracy is higher. The whole process does not need manual participation, is more time-saving, labor-saving and efficient, and saves labor cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a keyword extraction method and related device. Background Art

[0002] In some scenarios where a transcribing engine transcribes audio, the speaker pronounces based on the original text, and the transcribing engine then transcribes the audio at the pronunciation site into text materials.

[0003] In order to improve the accuracy of the transcribing engine in transcribing audio into text, it is usually necessary to configure target keywords that are prone to recognition errors in the original text into the transcribing engine.

[0004] Currently, it is usually necessary to manually find target keywords from the original text. However, this method relies on manual experience, is time-consuming and laborious, and is very prone to problems such as omission of target keywords or inaccurate target keywords such as non-essential keywords being configured into the transcribing engine. Summary of the Invention

[0005] In view of this, this application provides a keyword extraction method and related device for automatically and accurately extracting target keywords from the original text. The technical solution is as follows:

[0006] The first aspect of this application provides a keyword extraction method, including:

[0007] Obtain the original text and the audio data of the original text;

[0008] Process the audio data into a transcribed text through a target transcribing engine;

[0009] Compare the transcribed text with the original text to obtain target keywords, where the target keywords refer to specific words that the target transcribing engine makes transcription errors when transcribing the audio data.

[0010] In a possible implementation manner, after comparing the transcribed text with the original text to obtain target keywords, it further includes:

[0011] Configure the target keywords into the target transcribing engine, and return to process the audio data into a transcribed text through the target transcribing engine until a preset loop end condition is met.

[0012] In a possible implementation manner, the comparing the transcribed text with the original text to obtain target keywords includes:

[0013] Compare the transcribed text with the original text sentence by sentence to obtain the different sentences in the original text and the incorrect sentences in the transcribed text, where the different sentences and the incorrect sentences correspond one by one;

[0014] Compare the different sentences and the incorrect sentences word by word to obtain the target keywords.

[0015] In a possible implementation, the step of comparing the different sentences and the incorrect sentences word by word to obtain the target keywords includes:

[0016] Compare the different sentences and the incorrect sentences word by word to obtain the different word information in the different sentences, where the different word information includes different words and / or the different positions of the different words;

[0017] Extract the target keywords from the different sentences according to the different word information.

[0018] In a possible implementation, the step of extracting the target keywords from the different sentences according to the different word information includes:

[0019] Segment the different sentences to obtain a segmentation sequence;

[0020] Extract target segments from the segmentation sequence according to the different word information, and determine the target segments as the target keywords.

[0021] In a possible implementation, after comparing the different sentences and the incorrect sentences word by word to obtain the different word information in the different sentences, the method further includes:

[0022] Determine the number of incorrect words and the total number of words in the original text;

[0023] Calculate the transcription recognition accuracy according to the number of incorrect words and the total number of words in the original text.

[0024] In a possible implementation, the preset loop end condition is any one or a combination of the following conditions: the transcription recognition accuracy reaches a preset accuracy threshold, the transcribed text is exactly the same as the original text, and a preset number of loops is reached.

[0025] In a possible implementation, the process of obtaining the audio data of the original text includes:

[0026] Perform audio synthesis on the original text to obtain synthesized audio, and use the synthesized audio as the audio data of the original text.

[0027] The second aspect of the present application provides a keyword extraction device, including: a data acquisition module, an audio transcription module, and a keyword extraction module;

[0028] The data acquisition module is configured to acquire the original text and the audio data of the original text;

[0029] The audio transcription module is configured to process the audio data into a transcribed text through a target transcription engine;

[0030] The keyword extraction module is configured to compare the transcribed text with the original text to obtain target keywords, where the target keywords refer to specific words with transcription errors when the target transcription engine transcribes the audio data.

[0031] The third aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, where:

[0032] The memory is used to store a computer program;

[0033] The processor is configured to execute the computer program so that the electronic device can implement the steps of any one of the above-mentioned keyword extraction methods.

[0034] The fourth aspect of the present application provides a computer storage medium, where the storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the steps of any one of the above-mentioned keyword extraction methods.

[0035] The fifth aspect of the present application provides a computer program product, including computer-readable instructions, and when the computer-readable instructions run on an electronic device, the electronic device can implement the steps of any one of the above-mentioned keyword extraction methods.

[0036] By means of the above technical solutions, the keyword extraction method provided by the present application first acquires the original text and the audio data of the original text, then processes the audio data into a transcribed text through a target transcription engine, and finally compares the transcribed text with the original text to obtain target keywords. The keyword extraction method provided by the present application can enable the target transcription engine to perform real audio transcription, and according to the transcribed text and the original text, the target keywords with transcription errors of the target transcription engine can be extracted from the original text, with higher accuracy. The whole process can be carried out without manual participation, which is more time-saving, labor-saving and efficient, and saves labor costs. Description of the Drawings

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0038] Figure 1 A schematic diagram of a system architecture related to this application;

[0039] Figure 2 A schematic diagram of a hardware structure of a terminal provided by an embodiment of this application;

[0040] Figure 3 A schematic diagram of a hardware structure of a server provided by an embodiment of this application;

[0041] Figure 4 A schematic flowchart of a keyword extraction method provided by an embodiment of this application;

[0042] Figure 5 A schematic flowchart of another keyword extraction method provided by an embodiment of this application;

[0043] Figure 6 A schematic diagram of the structure of a keyword extraction device provided by an embodiment of this application. Detailed implementation manners

[0044] The following describes the embodiments of this application in combination with the drawings in the embodiments of this application. The terms used in the implementation part of this application are only used to explain the specific embodiments of this application, rather than aiming to limit this application.

[0045] The following describes the embodiments of this application in combination with the drawings. Those of ordinary skill in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0046] The terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of this application. In addition, the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0047] In a possible implementation, as Figure 1 shown, the system architecture involved in this application may include a terminal 101 and a server 102. The terminal 101 can interact with the server 102 through a network (wired network or wireless network). Among them, the server 102 may include one or more servers ( Figure 1 illustrated by taking one server as an example). The terminal and the server cooperate to implement keyword extraction.

[0048] In another possible implementation, the system architecture involved in this application may include a terminal. The terminal has strong data processing capabilities and can implement keyword extraction.

[0049] Next, the product form of the above terminal will be described.

[0050] The above terminal may be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, a robot, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The embodiments of this application do not make any restrictions on this.

[0051] Figure 2 Shows an optional schematic diagram of the hardware structure of the terminal.

[0052] Referring to Figure 2 shown, the terminal may include a radio frequency unit 210, a memory 220, an input unit 230, a display unit 240, a camera 250 (optional), an audio circuit 260 (optional), a speaker 261 (optional), a microphone 262 (optional), a headphone jack 263 (optional), a processor 270, an external interface 280, a power supply 290, and other components. Those skilled in the art can understand that Figure 2 this is only an example of the terminal and does not constitute a limitation on the terminal. It may include more or fewer components than shown, or combine some components, or different components.

[0053] The input unit 230 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the terminal. Specifically, the input unit 230 can include a touch screen 231 (optional) and / or other input devices 232. The touch screen 231 can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object such as a finger, a joint, a stylus, etc. on or near the touch screen), and drive corresponding connection devices according to a preset program. The touch screen can detect the touch action of the user on the touch screen, convert the touch action into a touch signal and send it to the processor 270, and can receive and execute the commands sent by the processor 270; the touch signal at least includes contact coordinate information. The touch screen 231 can provide an input interface and an output interface between the terminal and the user. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch screen. In addition to the touch screen 231, the input unit 230 can also include other input devices. Specifically, the other input devices 232 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0054] The display unit 240 can be used to display information input by the user or information provided to the user, various menus of the terminal, an interactive interface, file display, and / or the playback of any multimedia file.

[0055] The memory 220 can be used to store instructions and data. The memory 220 mainly includes a storage instruction area and a storage data area. The storage data area can store various data, such as multimedia files, texts, etc.; the storage instruction area can store software units such as an operating system, applications, instructions required for at least one function, or their subsets or extended sets. It can also include a non-volatile random access memory; it provides the processor 270 with management of hardware, software, and data resources in the computing processing device, supports control software and applications. It is also used for the storage of multimedia files, and the storage of running programs and applications.

[0056] The processor 270 is the control center of the terminal, connecting various parts of the entire terminal through various interfaces and lines. By running or executing instructions stored in the memory 220 and invoking data stored in the memory 220, it executes various functions of the terminal and processes data, thereby exercising overall control over the terminal. Optionally, the processor 270 may include one or more processing units; preferably, the processor 270 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 270. In some embodiments, the processor and the memory may be implemented on a single chip, and in some embodiments, they may also be separately implemented on independent chips. The processor 270 can also be used to generate corresponding operation control signals, send them to corresponding components of the computing and processing device, read and process data in the software, especially read and process the data and programs in the memory 220, so that each functional module therein executes corresponding functions, thereby controlling the corresponding components to act according to the requirements of the instructions.

[0057] Among them, the memory 220 can be used to store software codes related to the keyword extraction method. The processor 270 can execute the software codes in the memory 220 or can also schedule other units (such as the above-mentioned input unit 230 and display unit 240) to implement corresponding functions.

[0058] The radio frequency unit 210 (optional) can be used to receive and transmit information or signals during a call. For example, after receiving the downlink information of the base station, it is given to the processor 270 for processing; in addition, it sends the designed uplink data to the base station. Generally, the radio frequency unit 210 includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the radio frequency unit 210 can also communicate with network devices and other devices through wireless communication. This wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0059] Among them, in the embodiments of the present application, the radio frequency unit 210 can send data to other devices and can also receive data sent by other devices. It should be understood that the radio frequency unit 210 is optional and can be replaced by other communication interfaces, such as a network interface.

[0060] The terminal further includes a power supply 290 (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the processor 270 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.

[0061] The terminal further includes an external interface 280, which can be a standard Micro USB interface or a multi-pin connector, and can be used to connect the terminal to other devices for communication and can also be used to connect a charger to charge the terminal.

[0062] Although not shown, the terminal may further include a flashlight, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which will not be elaborated here.

[0063] Next, the product form of the above server will be described.

[0064] Figure 3 A schematic structural diagram of the above server is provided, as Figure 3 shown, the server may include a bus 301, a processor 302, a communication interface 303, and a memory 304. The processor 302, the memory 304, and the communication interface 303 communicate with each other through the bus 301.

[0065] The bus 301 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 3 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0066] The processor 302 can be any one or more of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Micro Processor (MP), or a Digital Signal Processor (DSP), etc.

[0067] The memory 304 may include volatile memory, such as random access memory (RAM). The memory 304 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0068] The memory 304 can be used to store software code related to the keyword extraction method. The processor 302 can call the software code stored in the memory 304 and can also schedule other units to implement corresponding functions.

[0069] The processors in the above terminal and server (such as processor 270 and processor 302) can be hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general-purpose processor, digital signal processor (DSP), microprocessor, or microcontroller, etc.), or a combination of these hardware circuits. For example, the processor can be a hardware system with the function of executing instructions, such as CPU, DSP, etc., or a hardware system without the function of executing instructions, such as ASIC, FPGA, etc., or a combination of the above hardware systems without the function of executing instructions and the hardware systems with the function of executing instructions.

[0070] In order to be able to automatically and accurately extract target keywords that are mis-transcribed by the target transcription engine but are expected to be correctly transcribed from the original text, in the initial stage of thinking, a model can be trained for the original text, and then the model can be generated into a patch package and deployed to the target transcription engine. However, the keyword generation efficiency of the patch package is low, and there are relatively large false trigger situations.

[0071] In order to automatically and accurately extract target keywords from the original text without introducing a patch package, this application provides a keyword extraction method and related device.

[0072] Optionally, the keyword extraction method and related device provided by this application can be applied to a variety of scenarios, including but not limited to the following scenarios.

[0073] Scenario 1: In a meeting scenario, it is usually possible to obtain speech materials related to the meeting content before the meeting. To achieve screen projection during the meeting, target keywords can be extracted from the original speech materials, such as keywords like uncommon technical terms and names. The target keywords are configured into the target transcription engine so that when the audio of the speaker during the meeting is transcribed by the target transcription engine, the target keywords can be accurately transcribed. For example, if the target keyword "Zhang Hong" is not configured into the target transcription engine, when the speaker mentions "Zhang Hong", the target transcription engine may transcribe it as "Zhang Hong", but by configuring the target keyword "Zhang Hong" into the target transcription engine, it can be accurately transcribed and projected onto the screen, improving the transcription recognition accuracy and user experience.

[0074] Scenario 2: In a medical training scenario, the training teacher usually writes training materials before the training. To enable the trainees to review and revise the training content, target keywords can be extracted from the original training materials, such as medical terms. The target keywords are configured into the target transcription engine so that when the audio of the training teacher during the training is transcribed by the target transcription engine, target keywords such as medical terms can be accurately transcribed.

[0075] Of course, in addition to this, there can be other scenarios, which are not specifically limited in this application.

[0076] To enable those skilled in the art to better understand this application, the keyword extraction method and related devices provided by this application are introduced in detail through the following embodiments.

[0077] Please refer to Figure 4 , which shows a schematic flowchart of a keyword extraction method provided by an embodiment of this application. The keyword extraction method may include:

[0078] Step S401, obtain the original text and the audio data of the original text.

[0079] Here, the original text refers to the original electronic text.

[0080] If the original paper materials are obtained in this embodiment, the paper materials can be recognized to obtain the electronic text. Optionally, a scanner or a multifunctional printer or a mobile phone application (such as a photo scanning application) can be used to scan the paper materials into electronic text; optionally, optical character recognition software can also be used to perform OCR (Optical Character Recognition) recognition on the paper materials to obtain the electronic text; optionally, an online scanning service can also be used to scan the paper materials into electronic text.

[0081] Of course, the above methods for converting paper materials into electronic text are only examples, and there can be other methods in addition to this, which are not limited in this application.

[0082] Optionally, to avoid extracting incorrect target keywords in this application due to errors during the conversion of paper materials to electronic texts, this embodiment can perform manual fine calibration on the electronic text to obtain the original text.

[0083] This embodiment can also obtain the audio data of the original text. If the audio data of the original text cannot be directly obtained, the audio data can be generated based on the original text.

[0084] Here, the method for generating audio data based on the original text includes but is not limited to the following methods.

[0085] The first method: The audio data of the original text can be generated by means of manual recording.

[0086] The second method: The audio data of the original text can be generated by using a speech synthesis method. Specifically, the original text can be subjected to audio synthesis to obtain synthesized audio, and the synthesized audio can be used as the audio data of the original text.

[0087] Optionally, to avoid extracting incorrect target keywords in this application due to errors in the synthesized audio, this embodiment can perform manual verification on the synthesized audio. Then, the synthesized audio after manual verification can be used as the audio data of the original text.

[0088] In a possible implementation, if the target transcription engine that needs to optimize the target keywords can implement cross-lingual transcription, this embodiment can generate the audio data of the original text in a preset number of languages, and then respectively execute the following step S402 and step S403 for the audio data in each language to further improve the transcription ability of the target transcription engine for each language. Here, the process of generating the audio data of the original text in a preset number of languages can refer to the foregoing introduction and will not be elaborated here.

[0089] However, considering that most transcription engines in most scenarios are for within-language transcription, to avoid errors in transcription caused by cross-lingual transcription, preferably, the language of the audio data can be the same as the language of the original text.

[0090] Step S402: Process the audio data into a transcribed text through the target transcription engine.

[0091] Step S403: Compare the transcribed text with the original text to obtain the target keywords.

[0092] Here, the target keywords refer to the specific words that have transcription errors when the target transcription engine transcribes the audio data, that is, the specific words with differences between the transcribed text and the original text.

[0093] Optionally, the specific word may refer to the word that is expected to be correctly transcribed by the target transcription engine. For example, the words that are expected to be correctly transcribed by the target transcription engine in the original text can be added to a preset word set in advance. Then, the target keyword needs to be a word within the preset word set.

[0094] Optionally, the specific word can also be a word that meets a preset rule. Here, the preset rule can be set according to the actual scenario. For example, in some scenarios, the preset rule can be that the number of Chinese characters is greater than or equal to 2.

[0095] Of course, the specific word can also be other, and the present application does not make specific limitations.

[0096] It should be noted that if the target keyword includes multiple words (here, in Chinese, the word refers to a Chinese character, and in a foreign language, the word refers to a word), at least one of the multiple words has a transcription error. For example, both the transcribed text and the original text are in Chinese. The original text is "Thank Zhang Hong for his great contributions in the field of artificial intelligence", and the transcribed text is "Thank Zhang Hong in the field of artificial intelligence for his great contributions", then the target keyword can be "Zhang Hong".

[0097] The keyword extraction method provided by the present application first obtains the original text and the audio data of the original text, then processes the audio data into a transcribed text through the target transcription engine, and finally compares the transcribed text with the original text to obtain the target keyword. The keyword extraction method provided by the present application can enable the target transcription engine to perform real audio transcription, and according to the transcribed text and the original text, the target keyword with a transcription error in the target transcription engine can be extracted from the original text, with higher accuracy. The entire process can be carried out without manual participation, which is more time-saving, labor-saving and efficient, and saves labor costs.

[0098] In some embodiments of the present application, the specific process of "Step S403. Compare the transcribed text with the original text to obtain the target keyword" is introduced.

[0099] In an optional embodiment, this embodiment can pre-train a keyword extraction model, and then input the transcribed text and the original text into the keyword extraction model, so that the keyword extraction model can obtain the target keyword by comparing the differences between the transcribed text and the original text.

[0100] Here, the pre-trained keyword extraction model is trained by using the training text marked with keyword tags and the corresponding transcribed text.

[0101] In another optional embodiment, this embodiment can compare the transcribed text with the original text word by word to obtain the difference word information in the original text, and then extract the target keyword from the original text according to the difference word information.

[0102] Here, the differential word information includes at least one of the differential words and the differential positions of the differential words. The differential words refer to the words with transcription errors in the target transcription engine.

[0103] Taking Chinese as an example, optionally, if the differential words include multiple consecutive characters, the position information of the differential words may include the position information of each character in the multiple consecutive characters, or may only be the position information of the first character in the multiple consecutive characters.

[0104] Taking the original text as "amyotrophic lateral sclerosis" and the transcribed text as "amyotrophic toilet sclerosis" as an example, the differential words include "lateral" and "cord"; optionally, the position information of the differential words may be "3; 4", or may be "3". Here, "3" in the position information represents the position of "lateral", and "4" represents the position of "cord".

[0105] Optionally, the process of "extracting the target keyword from the original text according to the differential word information" may include: performing word segmentation on the original text to obtain the word segmentation result, and extracting the word segment where the differential word information is located from the word segmentation result as the target keyword.

[0106] Of course, the process of "extracting the target keyword from the original text according to the differential word information" may also have other implementation manners, which are not limited in this application.

[0107] In another optional embodiment, considering that the original text may contain a large number of words, directly comparing word by word will consume a lot of time and be inefficient. To improve the efficiency, the following embodiment is provided.

[0108] Optionally, this embodiment can compare the transcribed text with the original text sentence by sentence to obtain the differential sentences in the original text and the incorrect sentences in the transcribed text; then compare the differential sentences and the incorrect sentences word by word to obtain the target keyword.

[0109] The above-mentioned differential sentences refer to the sentences in the original text that are different from the transcribed text, and the incorrect sentences refer to the sentences in the transcribed text corresponding to the differential sentences. That is, in this embodiment, the differential sentences and the incorrect sentences correspond one by one.

[0110] For example, the original text is "AI is now changing the way we live and work at an unprecedented speed. With its powerful data processing capabilities and intelligent decision-making support, it has brought unprecedented opportunities and challenges to various fields. In the future, I will continue to pay attention to the development trends of AI, strive to learn and master relevant skills, and make full preparations for the arrival of this intelligent era. Finally, I would like to sincerely thank Teacher Zhang Hong. It is he who led me into the magical and vast field of artificial intelligence, giving me the opportunity to explore and learn in this field full of infinite possibilities. Thank you."

[0111] The transcribed text is "Love is now changing the way we live and work at an unprecedented speed. With its powerful data processing capabilities and intelligent decision-making support, it has brought unprecedented opportunities and challenges to various fields. In the future, I will continue to pay attention to the development trends of love, strive to learn and master relevant skills, and make full preparations for the arrival of this intelligent era. Finally, I would like to sincerely thank Teacher Zhang Hong. It is he who led me into the magical and vast field of artificial intelligence, giving me the opportunity to explore and learn in this field full of infinite possibilities. Thank you."

[0112] Then, the different sentence 1 is "AI is now changing the way we live and work at an unprecedented speed", and the corresponding incorrect sentence 1 is "Love is now changing the way we live and work at an unprecedented speed"; the different sentence 2 is "I will continue to pay attention to the development trends of AI", and the corresponding incorrect sentence 2 is "I will continue to pay attention to the development trends of love"; the different sentence 3 is "I would like to sincerely thank Teacher Zhang Hong", and the corresponding incorrect sentence 3 is "I would like to sincerely thank Teacher Zhang Hong".

[0113] Optionally, the process of "comparing the different sentences and the incorrect sentences word by word to obtain the target keywords" may include: comparing the different sentences and the incorrect sentences word by word to obtain the different word information in the different sentences, where the different word information includes different words and / or the different positions of different words; extracting the target keywords from the different sentences according to the different word information.

[0114] Still taking the original text and the transcribed text shown above as an example, comparing the different sentence 1 with the incorrect sentence 1 word by word, the different word obtained is "AI", and the different position of the different word is "0"; comparing the different sentence 2 with the incorrect sentence 2 word by word, the different word obtained is "AI", and the different position of the different word is "6"; comparing the different sentence 3 with the incorrect sentence 3 word by word, the different word obtained is "Hong", and the different position of the different word is "8".

[0115] Optionally, in order to manage the differential sentences, error sentences, and differential word information, and improve the extraction efficiency of target keywords, in this embodiment, the differential sentences, error sentences, and differential word information can be stored in the recognition difference list in the form of a ternary structure.

[0116] In this embodiment, the process of "extracting target keywords from differential sentences according to differential word information" can be implemented in multiple ways. Here are two implementation methods provided below.

[0117] The first implementation method: Segment the differential sentence to obtain a segmentation sequence. According to the differential word information, extract the target segmentation from the segmentation sequence, and determine the target segmentation as the target keyword.

[0118] It should be noted that if multiple differential sentences are obtained previously, in this embodiment, each differential sentence among the multiple differential sentences is segmented, and then the target segmentation is extracted according to the corresponding differential word information, and the target segmentation is determined as the target keyword. Optionally, the target keywords can be stored in a preset keyword set for easy management.

[0119] Optionally, a segmentation tool can be used to segment the differential sentence to obtain a segmentation sequence; the differential sentence can also be subjected to grammatical analysis to segment it according to the grammatical analysis result to obtain an analysis sequence. Of course, the segmentation process can also be other, and this application does not limit it.

[0120] Optionally, the segmentation tool can be the Jieba segmentation tool or the HanLP tool or the LTP tool. In specific implementation, the segmentation tool can also be other, and this application will not list them one by one.

[0121] The following gives an example to introduce the process of "extracting target segmentation from the segmentation sequence according to differential word information and determining the target segmentation as the target keyword".

[0122] Taking the differential sentence "One-key video opens a new world" and the transscribed text "One thing frequency opens a new vision", with the differential words being "key" and "vision" and the corresponding differential positions being "1" as an example, the segmentation sequence of "One-key video opens a new world" is "One-key / video / open / new world". Then, according to the differential words and the corresponding differential positions, the target segmentation, that is, the target keywords, can be determined as "One-key" and "video" (it can be seen that there are two "visions" in the segmentation sequence of the differential sentence, but since there is only one position in the differential position, in this example, the determined target segmentation is "video", rather than "new world").

[0123] The second implementation method: Search for different words from the different sentences according to the different word information, and perform recognition based on the word boundaries according to the found different words, so as to recognize the forward adjacent words and / or backward adjacent words of the word segmentation formed by the different words in the different sentences, and form word segmentation with the forward adjacent words and / or backward adjacent words of the recognized different words as the target keywords.

[0124] It should be noted that the above two implementation methods are only examples and do not limit this application.

[0125] As introduced above, there may also be a preset word set or preset rules. Optionally, after obtaining the target keywords through the above implementation methods, the following steps may further be included: using the target keywords as candidate keywords, removing the candidate keywords that are not in the preset word set or do not meet the preset rules, and using the remaining candidate keywords as the target keywords configured in the target transcription engine of this application.

[0126] By comparing the original text with the transcribed text, this application can accurately identify the target keywords that the target transcription engine actually transcribes incorrectly from the original text, and configuring the target keywords into the target transcription engine can effectively improve the transcription recognition accuracy of the target transcription engine.

[0127] In a possible implementation, in order to be able to determine the specific transcription recognition accuracy, after comparing the different sentences and the incorrect sentences word by word to obtain the different word information in the different sentences in the previous text, this embodiment can also determine the number of incorrect words and the total number of words included in the original text, and calculate the transcription recognition accuracy according to the number of incorrect words and the total number of words included in the original text.

[0128] For example, the transcription recognition accuracy = 1 - the number of incorrect words / the total number of words included in the original text.

[0129] In a more preferred implementation, considering that after configuring the target keywords (that is, the target keyword set in the previous text) into the target transcription engine, the target transcription engine may still make mistakes when transcribing audio data. To further improve the transcription recognition accuracy of the target transcription engine, this embodiment also provides another keyword extraction method.

[0130] See Figure 5 As shown, it is a flowchart of another keyword extraction method provided by an embodiment of this application, including:

[0131] Step S501: Obtain the original text and the audio data of the original text.

[0132] Step S502: Process the audio data into a transcribed text through the target transcription engine.

[0133] Step S503: Compare the transcribed text with the original text to obtain target keywords.

[0134] Step S501 - Step S503 correspond to Step S401 - Step S403 in the previous text one by one. For details, please refer to the previous introduction and will not be elaborated here.

[0135] Step S504: Configure the target keywords to the target transcription engine.

[0136] Specifically, configure the target keywords to the target transcription engine through the keyword optimization interface for engine keyword optimization.

[0137] Step S505: Determine whether the preset loop end condition is met. If not, return to Step S502; if so, end.

[0138] Optionally, the preset loop end condition can be any one of the following conditions or a combination of multiple conditions: the transcription recognition accuracy rate reaches the preset accuracy rate threshold (e.g., 98%), the transcribed text is exactly the same as the original text, and the preset number of loop times is reached (e.g., 2 times).

[0139] Specifically, the preset loop end condition can be any one of the following: the transcription recognition accuracy rate reaches the preset accuracy rate threshold; the transcription recognition accuracy rate reaches the preset accuracy rate threshold, or the preset number of loop times is reached; the transcribed text is exactly the same as the original text; the transcribed text is exactly the same as the original text, or the preset number of loop times is reached; the preset number of loop times is reached. It should be noted that the above 98% and 2 times are only examples and can be set to others according to needs in actual scenarios. For example, when comparing the transcribed text and the original text, no target keywords can be obtained, etc. This application does not make limitations. However, it is worth noting that the transcription recognition accuracy rate is preferably not set to 100% to avoid over - extraction due to incorrect recognition by the target transcription engine, which may lead to mis - recognition of the target transcription engine.

[0140] The above introduces the keyword extraction method provided by the embodiments of this application. Next, the device corresponding to the above keyword extraction method will be introduced.

[0141] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a keyword extraction device provided by the embodiments of this application. The keyword extraction device may include: a data acquisition module 601, an audio transcription module 602, and a keyword extraction module 603.

[0142] The data acquisition module 601 is used to acquire the original text and the audio data of the original text;

[0143] An audio transcription module 602, configured to process audio data into a transcribed text through a target transcription engine;

[0144] A keyword extraction module 603, configured to compare the transcribed text with the original text to obtain target keywords, where the target keywords refer to specific words that have transcription errors when the target transcription engine transcribes the audio data.

[0145] In a possible implementation, the keyword extraction device provided by the embodiments of the present application may further include: a keyword configuration module.

[0146] The keyword configuration module is configured to configure the target keywords into the target transcription engine.

[0147] The above-mentioned audio transcription module is further configured to, after the keyword configuration module configures the target keywords into the target transcription engine, process the audio data into a transcribed text again through the target transcription engine until a preset loop end condition is met.

[0148] In a possible implementation, when the above-mentioned keyword extraction module compares the transcribed text with the original text to obtain target keywords, it may specifically be used for:

[0149] Compare the transcribed text with the original text sentence by sentence to obtain the different sentences in the original text and the incorrect sentences in the transcribed text, where the different sentences and the incorrect sentences correspond one by one;

[0150] Compare the different sentences and the incorrect sentences word by word to obtain the target keywords.

[0151] In a possible implementation, when the above-mentioned keyword extraction module compares the different sentences and the incorrect sentences word by word to obtain the target keywords, it may specifically be used for:

[0152] Compare the different sentences and the incorrect sentences word by word to obtain the different word information in the different sentences, where the different word information includes different words and / or the different positions of the different words;

[0153] Extract the target keywords from the different sentences according to the different word information.

[0154] In a possible implementation, when the above-mentioned keyword extraction module extracts the target keywords from the different sentences according to the different word information, it may specifically be used for:

[0155] Segment the different sentences to obtain a segmentation sequence;

[0156] Extract the target segments from the segmentation sequence according to the different word information, and determine the target segments as the target keywords.

[0157] In a possible implementation, the keyword extraction device provided by the embodiments of the present application may further include: an accuracy rate calculation module.

[0158] The accuracy rate calculation module is configured to, after comparing the different sentences and the error sentences word by word to obtain the different word information in the different sentences, determine the number of incorrect words and the total number of words included in the original text, and calculate the transcribing recognition accuracy rate according to the number of incorrect words and the total number of words included in the original text.

[0159] In a possible implementation, the above-mentioned preset loop end condition is any one or a combination of the following conditions: the transcribing recognition accuracy rate reaches a preset accuracy rate threshold, the transcribing text is exactly the same as the original text, and a preset number of loops is reached.

[0160] In a possible implementation, when the above-mentioned data acquisition module acquires the audio data of the original text, it may specifically be used for:

[0161] Performing audio synthesis on the original text to obtain synthesized audio, and using the synthesized audio as the audio data of the original text.

[0162] The keyword extraction device provided by the embodiments of the present application first acquires the original text and the audio data of the original text, then processes the audio data into a transcribing text through a target transcribing engine, and finally compares the transcribing text with the original text to obtain target keywords. The keyword extraction device provided by the present application can enable the target transcribing engine to perform real audio transcribing, and according to the transcribing text and the original text, it can extract the target keywords where the target transcribing engine has transcribing errors from the original text, with higher accuracy. The whole process can be carried out without manual participation, which is more time-saving, labor-saving and efficient, and saves labor costs.

[0163] The embodiments of the present application further provide an electronic device, which may include: at least one processor, at least one communication interface, at least one memory, and at least one communication bus.

[0164] In the embodiments of the present application, the number of the processor, the communication interface, the memory, and the communication bus is at least one, and the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0165] The processor may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application, etc.;

[0166] The memory may include high-speed RAM memory and may also include non-volatile memory, etc., such as at least one disk memory;

[0167] Among them, the memory stores a program, and the processor can call the program stored in the memory, and the program is used to implement the steps of the keyword extraction method provided in the above embodiments.

[0168] The embodiments of the present application also provide a computer storage medium, and the storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the steps of the keyword extraction method provided in the above embodiments.

[0169] The embodiments of the present application also provide a computer program product, including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device can implement the steps of the keyword extraction method provided in the above embodiments.

[0170] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.

[0171] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware, including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for the present application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0172] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0173] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)), etc.

Claims

1. A keyword extraction method, characterized in that: include: Acquire original text and audio data of the original text; Processing the audio data into transcribed text by a target transcription engine; The transcribed text is compared with the original text to obtain a target keyword, wherein the target keyword refers to a specific word that is transcribed incorrectly when the target transcription engine transcribes the audio data.

2. The keyword extraction method according to claim 1, characterized in that: After comparing the transcribed text with the original text to obtain the target keyword, the method further includes: The target keyword is configured to the target transcription engine, and the audio data is processed into a transcribed text by the target transcription engine, until a preset loop end condition is met.

3. The keyword extraction method according to claim 2, characterized in that: The step of comparing the transcribed text with the original text to obtain target keywords includes: Comparing the transcribed text with the original text sentence by sentence to obtain difference sentences in the original text and erroneous sentences in the transcribed text, wherein the difference sentences and the erroneous sentences correspond to each other one by one; The difference sentence and the erroneous sentence are compared word by word to obtain the target keyword.

4. The keyword extraction method according to claim 3, characterized in that: The step of comparing the difference sentence and the erroneous sentence word by word to obtain the target keyword comprises: Comparing the difference sentence and the erroneous sentence word by word to obtain difference word information in the difference sentence, wherein the difference word information includes difference words and / or difference positions of the difference words; The target keyword is extracted from the difference sentence according to the difference word information.

5. The keyword extraction method according to claim 4, characterized in that: The step of extracting the target keyword from the difference sentence according to the difference word information includes: Segmenting the difference sentences to obtain a segmentation sequence; According to the difference word information, a target segmented word is extracted from the segmented word sequence, and the target segmented word is determined as the target keyword.

6. The keyword extraction method according to claim 4 or 5, characterized in that: After comparing the difference sentence and the wrong sentence word by word to obtain the difference word information in the difference sentence, the method further includes: Determining the number of the erroneous words and the total number of words contained in the original text; The transcription recognition accuracy is calculated based on the number of erroneous words and the total number of words contained in the original text.

7. The keyword extraction method according to claim 6, characterized in that: The preset loop end condition is any one of the following conditions or a combination of multiple conditions: the transcription recognition accuracy reaches a preset accuracy threshold, the transcribed text is completely consistent with the original text, and a preset number of cycles is reached.

8. The keyword extraction method according to claim 1 or 2, characterized in that: The process of obtaining the audio data of the original text includes: Perform audio synthesis on the original text to obtain synthesized audio, and use the synthesized audio as audio data of the original text.

9. A keyword extraction device, characterized in that: include: Data acquisition module, audio transcription module and keyword extraction module; The data acquisition module is used to acquire the original text and the audio data of the original text; The audio transcription module is used to process the audio data into a transcribed text through a target transcription engine; The keyword extraction module is used to compare the transcribed text with the original text to obtain a target keyword, wherein the target keyword refers to a specific word that is transcribed incorrectly when the target transcription engine transcribes the audio data.

10. An electronic device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the steps of the keyword extraction method as described in any one of claims 1 to 8.

11. A computer storage medium, characterized in that: The storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the steps of the keyword extraction method as described in any one of claims 1 to 8.

12. A computer program product, characterized in that The method comprises computer-readable instructions, which, when executed on an electronic device, enable the electronic device to implement the steps of the keyword extraction method as described in any one of claims 1 to 8.