Data Processing Method, Apparatus, Server, and Storage Medium
By performing text recognition and keyword database query on voice data, highlighting the key information of speech content in the simultaneous interpretation system, the problem that existing systems cannot effectively extract and display key information and improve users' intuitive understanding ability.
Patent Information
- Application Number
- CN201980100284.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-04
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2039-11-04
AI Technical Summary
The existing simultaneous interpretation system cannot effectively extract and highlight key information about the speech content, making it difficult for users to intuitively understand the focus of the speech content.
By text recognition of speech data, we look up target fragments in the keyword database that meet preset conditions, and when presenting the recognition text, we use a presentation format different from other texts to highlight the target fragments.
It realizes the extraction and highlighting of key information of voice data, allowing users to intuitively understand the focus of the speech content and improves the user experience.
Smart Images

Figure CN114402384B_ABST
Abstract
Description
Technical Field
[0001] This application relates to simultaneous interpretation technology, and particularly to a data processing method, apparatus, server, and storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, the concept of artificial intelligence (AI) has gradually moved from the black technology in the laboratory to reality and been applied to all aspects of real life.
[0003] The simultaneous interpretation system is a voice translation product for conference scenarios that has emerged in recent years. It uses AI technology to provide multilingual text translation and text display for the speech content of conference speakers.
[0004] In related simultaneous interpretation systems, the speech content is displayed through text, but for users, they cannot truly and intuitively understand the key information of the speech content through the displayed content. Summary of the Invention
[0005] To solve the related technical problems, embodiments of this application provide a data processing method, apparatus, server, and storage medium.
[0006] Embodiments of this application provide a data processing method applied to a server, including:
[0007] Obtain the voice data to be processed, perform text recognition on the voice data to obtain the recognized text; the recognized text is used for presentation when playing the voice data;
[0008] Search for a keyword library according to the recognized text, and determine a target segment in the recognized text that meets the first preset condition;
[0009] Determine the first presentation format of the target segment, so as to present the target segment in the first presentation format when presenting the recognized text; the first presentation format is different from the second presentation format; the second presentation format is the presentation format of other texts in the recognized text except the target segment.
[0010] In the above solution, determining the target segment in the recognized text that meets the first preset condition includes at least one of the following:
[0011] Determine a target segment in the recognized text that matches any keyword in the keyword library;
[0012] Determine at least two keywords in the recognized text; determine the target segment based on the weights of the keywords in the at least two keywords.
[0013] In the above solution, the keyword library includes at least one keyword table;
[0014] Determining the first presentation format of the target segment includes:
[0015] Determining the target keyword table corresponding to the target segment; the target keyword table includes keywords that match the target segment;
[0016] Taking the format corresponding to the target keyword table as the first presentation format.
[0017] In the above solution, the keyword library includes at least two keyword tables; each keyword table in the at least two keyword tables corresponds to a different format; each keyword table in the at least two keyword tables corresponds to a different priority;
[0018] Determining the target keyword table corresponding to the target segment includes:
[0019] Determining at least two candidate keyword tables corresponding to the target segment;
[0020] Taking the candidate keyword table with the highest priority in the at least two candidate keyword tables as the target keyword table.
[0021] In the above solution, the method further includes:
[0022] Performing word segmentation on the recognized text to obtain at least one word;
[0023] Filtering the at least one word, and taking the words obtained after filtering as the word segmentation result;
[0024] Updating the first keyword table based on the word segmentation result; the first keyword table is a keyword table in the keyword library; the keywords and the weights of the keywords in the first keyword table change with the change of the voice data to be processed.
[0025] In the above solution, updating the first keyword table based on the word segmentation result includes:
[0026] For each word in the word segmentation result, determining the occurrence times and the number of word elements of the corresponding word;
[0027] Determining the weight of the corresponding word based on the occurrence times and the number of word elements; the weight changes with the change of the occurrence times of the corresponding word in the recognized text; the recognized text changes with the change of the voice data to be processed;
[0028] Determining the words in the word segmentation result that meet the second preset condition as keywords;
[0029] Update the first keyword table according to the keywords that meet the second preset condition and the weights corresponding to the keywords; each keyword corresponds to at least one language.
[0030] In the above solution, determining the words in the word segmentation result that meet the second preset condition includes at least one of the following:
[0031] Determine the words in the word segmentation result whose weights exceed the preset weight threshold;
[0032] Determine the words in the word segmentation result whose occurrence times exceed the preset occurrence threshold.
[0033] In the above solution, each keyword in the first keyword table corresponds to a font change factor, and the font change factor is related to the weight;
[0034] Determining the first presentation format of the target segment includes:
[0035] When the target keyword table corresponding to the target segment is the first keyword table, determine the format corresponding to the font change factor as the first presentation format.
[0036] In the above solution, the method further includes:
[0037] Extract terms from the bilingual data of the machine translation model, and generate a second keyword table based on the extracted terms; the second keyword table is one of the keyword tables in the keyword library.
[0038] An embodiment of the present application further provides a data processing device, including:
[0039] An acquisition unit, configured to obtain the speech data to be processed, perform text recognition on the speech data, and obtain the recognition text; the recognition text is used for presentation when the speech data is played;
[0040] A first processing unit, configured to search the keyword library according to the recognition text, and determine the target segment in the recognition text that meets the first preset condition;
[0041] A second processing unit, configured to determine the first presentation format of the target segment, so as to present the target segment in the first presentation format when presenting the recognition text; the first presentation format is different from the second presentation format; the second presentation format is the presentation format of other texts in the recognition text except the target segment.
[0042] An embodiment of the present application further provides a server, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of any of the above data processing methods.
[0043] An embodiment of the present application also provides a storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of any of the above data processing methods are implemented.
[0044] The data processing method, device, server, and storage medium provided by the embodiments of the present application obtain voice data to be processed, perform text recognition on the voice data to obtain recognition text; the recognition text is used for presentation when the voice data is played; a keyword library is searched according to the recognition text to determine a target segment in the recognition text that meets a first preset condition; a first presentation format of the target segment is determined, so that the target segment is presented in the first presentation format when the recognition text is presented; the first presentation format is different from a second presentation format; the second presentation format is the presentation format of other characters in the recognition text except the target segment. In this way, key information can be extracted from the voice data and the key information can be highlighted in the recognition text, so that the user can intuitively understand the key information of the voice data. Description of the Drawings
[0045] Figure 1 It is a schematic diagram of the system architecture for the application of simultaneous interpretation methods in the related art;
[0046] Figure 2 It is a schematic flowchart of a data processing method according to an embodiment of the present application;
[0047] Figure 3 It is another schematic flowchart of a data processing method according to an embodiment of the present application;
[0048] Figure 4 It is a schematic flowchart of a method for determining a first presentation format according to an embodiment of the present application;
[0049] Figure 5 It is a schematic diagram of the composition structure of a data processing device according to an embodiment of the present application;
[0050] Figure 6 It is a schematic diagram of the composition structure of a server according to an embodiment of the present application. Detailed Embodiments
[0051] The present application will be further described in detail below with reference to the drawings and embodiments.
[0052] Before elaborating on the technical solutions of the embodiments of the present application in detail, a system for the application of simultaneous interpretation methods in the related art will be briefly described first.
[0053] Figure 1 It is a schematic diagram of the system architecture for the application of simultaneous interpretation methods in the related art; as Figure 1As shown, the system may include: a machine simultaneous interpretation server, a voice processing server, an audience mobile terminal, a personal computer (PC) client, and a display screen.
[0054] In practical applications, a speaker can give a conference speech through the PC client. During the conference speech, the PC client collects the speaker's voice data and sends the collected voice data to the machine simultaneous interpretation server. The machine simultaneous interpretation server identifies the voice data through the voice processing server to obtain an identification result (the identification result can be an identification text in the same language as the voice data, or a translation text in another language obtained by translating the identification text); the machine simultaneous interpretation server can send the identification result to the PC client, and the PC client can project the identification result onto the display screen; it can also send the identification result to the audience mobile terminal (specifically, according to the language required by the user, the corresponding identification result in the corresponding language is sent) to display the identification result for the user, so as to realize translating the speaker's speech content into the language required by the user and displaying it.
[0055] However, only text recognition and translation of voice data are performed, and the speech content is displayed through text. The key information in the speech content is not extracted, and the key information cannot be prominently displayed to the user. For the user, the key information of the speech content cannot be truly and intuitively understood through the displayed content, which is not convenient enough.
[0056] Based on this, in various embodiments of the present application, the voice data is identified to obtain an identification text, the keyword library is used to query the identification text, and the target segment is determined; when presenting the identification text, the target segment is presented in a format different from the other texts in the identification text except the target segment; thus, the key information (i.e., the target segment) of the voice data can be extracted, and the key information can be prominently displayed, enabling the user to intuitively understand the key information of the voice data.
[0057] An embodiment of the present application provides a data processing method, which is applied to a server. Figure 2 It is a schematic flowchart of a data processing method according to an embodiment of the present application; as Figure 2 shown, the method includes:
[0058] Step 201: Obtain the voice data to be processed, perform text recognition on the voice data, and obtain an identification text;
[0059] Here, the identification text is used for presentation when playing the voice data.
[0060] Step 202: Search the keyword library according to the identification text, and determine the target segment in the identification text that meets the first preset condition;
[0061] Step 203: Determine a first presentation format for the target segment, so as to present the target segment in the first presentation format when presenting the recognized text;
[0062] Here, the first presentation format is different from the second presentation format; the second presentation format is the presentation format of other texts in the recognized text except the target segment.
[0063] Among them, in step 201, in practical applications, the voice data to be processed can be collected by a first terminal and sent to the server. The first terminal can be a mobile terminal such as a personal computer or a tablet computer. The first terminal can be provided with or connected to a voice collection module, such as a microphone, and the voice collection module is used to collect sound to obtain the voice data to be processed.
[0064] In step 201, presenting the recognized text when playing the voice data means presenting the recognized text while playing the voice data, that is, the data data processing method is applied to the scenario of simultaneous interpretation.
[0065] Specifically, in the scenario of simultaneous interpretation, when a speaker is giving a speech, a first terminal (such as Figure 1 the PC shown) uses a voice collection module to collect the speech content in real time, that is, to obtain the voice data to be processed. A communication connection can be established between the first terminal and the server, and the first terminal sends the obtained voice data to the server, and the server can obtain the voice data to be processed in real time. The server performs text recognition on the voice data to be processed, obtains the recognized text and presents it, that is, realizes presenting the recognized text while playing the voice data.
[0066] The simultaneous interpretation scenario can adopt a system architecture such as Figure 1 shown. The method of this application is applied to the server. The server can be a newly added server in the Figure 1 system architecture, used to implement the solution of this application (that is, the method shown in Figure 2 ), or it can be an improvement on the voice processing server in the Figure 1 architecture to implement the solution of this application.
[0067] In practical applications, the recognized text obtained according to the voice data can correspond to one or more languages, and the recognized texts in different languages are used to be presented to users in different languages.
[0068] Here, the recognized text corresponds to at least one language. The recognized text can be the recognized text in the same language as the voice data to be processed (denoted as the first language), or the recognized text in other languages after translating the recognized text in the first language. Specifically, it can be the recognized text in the second language, …, the recognized text in the Nth language, where N is greater than or equal to 1.
[0069] When the recognized text is the text in the same language as the voice data to be processed, the text recognition of the voice data to obtain the recognized text includes:
[0070] Performing automatic speech recognition (ASR) on the voice data to obtain the recognized text in the first language; the first language is the same as the language corresponding to the voice data.
[0071] When the recognized text is the text in a different language from the voice data to be processed, the text recognition of the voice data to obtain the recognized text includes:
[0072] Performing automatic speech recognition on the voice data to obtain the recognized text in the first language; the first language is the same as the language corresponding to the voice data;
[0073] Using a preset translation model to perform machine translation (MT) on the recognized text in the first language to obtain the recognized text in other languages.
[0074] By performing text recognition on the voice data in the above manner, the obtained recognized text corresponds to at least one language, that is, the recognized text in the first language, the recognized text in the second language, …, the recognized text in the Nth language can be obtained according to the voice data, where N is greater than or equal to 1.
[0075] Here, the translation model is used to translate the text in one language into the text in another language.
[0076] In an embodiment, after the server obtains the recognized text, it can send the obtained recognized text to the second terminal held by the user (such as Figure 1 the viewer mobile terminal shown), and the recognized text is presented by the second terminal when playing the voice data, and the user can read the recognized text to understand the content of the voice data. Here, the user holding the second terminal can also select a language through the human-computer interaction interface of the second terminal, and the second terminal sends the selected language to the server, and the server sends the recognized text in the corresponding language according to the language selected by the user.
[0077] In another embodiment, the server may also send the recognized text to the first terminal, and the first terminal presents the recognized text in at least one language through the connected display screen (i.e., using screen mirroring technology for screen mirroring), and the user reads the recognized text in the corresponding language to understand the content of the voice data.
[0078] Among them, in step 202, in practical applications, there may be one or more target segments in the recognized text. The target segment refers to a string of characters in the recognized text, such as terms, keywords, etc.
[0079] In step 202, determining the target segments in the recognized text that meet the first preset condition includes at least one of the following:
[0080] Determining the target segments in the recognized text that match any keyword in the keyword library;
[0081] Determining at least two keywords from the recognized text; determining the target segments based on the weights of the keywords in the at least two keywords.
[0082] Specifically, when a string of characters in the recognized text only matches one keyword in the keyword library, the characters that match one keyword are considered as one target segment.
[0083] When a string of characters in the recognized text can match at least two keywords in the keyword library, determine the weights of the at least two keywords, and determine the target segments based on the keyword with the higher weight.
[0084] For example, the keyword library includes two keywords: translation, machine translation. When the recognized text contains a string of characters: machine translation, the characters "machine translation" can match the above two keywords. At this time, determine the weights of the keywords "translation" and "machine translation". If the weight of the keyword "translation" is higher, then determine the target segment as: translation; otherwise, if the weight of the keyword "machine translation" is higher, then determine the target segment as: machine translation.
[0085] In practical applications, the selection criteria for target segments may be different. For example, it may be for technical terms, repeatedly mentioned content, etc. in the recognized text; in order to determine target segments according to multiple criteria, the keyword library may consist of at least one keyword list.
[0086] Based on this, in one embodiment, the keyword library may include at least one keyword list;
[0087] Determining the first presentation format of the target segments includes:
[0088] Determine the target keyword table corresponding to the target segment; the target keyword table includes keywords that match the target segment;
[0089] Use the format corresponding to the target keyword table as the first presentation format.
[0090] Here, the second presentation format can be a preset presentation format for identifying text. The first presentation format corresponds to the keyword table and is different from the second presentation format.
[0091] In practical applications, the keyword library can include at least two keyword tables; each keyword table in the at least two keyword tables corresponds to a different format; each keyword table in the at least two keyword tables corresponds to a different priority;
[0092] When there are at least two keyword tables corresponding to the target segment (that is, the keywords that match the target segment exist in at least two keyword tables), at this time, determining the first presentation format of the target segment includes:
[0093] Determine at least two candidate keyword tables corresponding to the target segment;
[0094] Use the candidate keyword table with a higher priority in the at least two candidate keyword tables as the target keyword table.
[0095] For example, the keyword library includes: Keyword Table One and Keyword Table Two; the priority of Keyword Table One is higher than that of Keyword Table Two; Keyword Table One corresponds to Presentation Format One, and Keyword Table Two corresponds to Presentation Format Two. Keyword Table One includes keywords A and B; Keyword Table Two includes keywords B and C; the server searches the keyword library according to the recognition text and determines the target segment: keyword B; that is, the keywords that match the target segment exist in two keyword tables. Since the two keyword tables correspond to different presentation formats respectively; at this time, select Presentation Format One corresponding to Keyword Table One with a very high priority as the first presentation format of the target segment.
[0096] Here, in order to enable users to more intuitively understand which keyword table the target segment is determined based on, each keyword table in the at least two keyword tables can correspond to a different presentation format.
[0097] Here, considering that different keyword tables may have the same keywords, in order to determine the first presentation format corresponding to the target segment existing in different keyword tables, here each keyword table in the at least two keyword tables can correspond to a different priority, and select the presentation format corresponding to the keyword table with a higher priority.
[0098] In actual application, the weight of each word can be associated with the repetition degree of each word in the recognized text, and the weight of each word can be updated according to the repetition degree of the word, so that the determined target segment can more accurately reflect the key points of the voice data, thereby helping the user intuitively understand the key content of the voice data.
[0099] Based on this, in one embodiment, the method further includes:
[0100] Segment the recognized text to obtain at least one word;
[0101] Filter the at least one word, and use the words obtained after filtering as the segmentation result;
[0102] Based on the segmentation result, update the first keyword table; the first keyword table is one of the keyword tables in the keyword library; the keywords and the weights of the keywords in the first keyword table change with the change of the voice data to be processed.
[0103] Here, the filtering of the at least one word includes:
[0104] Filter out the words that are the same as each stop word in the preset stop word table from the at least one word, and use the words obtained after filtering as the segmentation result.
[0105] The stop word table can be preset, and the stop word table can include conventional pause words, such as: this, of, etc., and can also include: stop words that the user hopes to filter out and will not become the target segment, such as: country names, etc. that are easily mentioned repeatedly but do not need to be presented specially.
[0106] Specifically, the updating of the first keyword table based on the segmentation result includes:
[0107] For each word in the segmentation result, determine the occurrence times and the number of word forms of the corresponding word;
[0108] Determine the weight of the corresponding word based on the occurrence times and the number of word forms; the weight changes with the change of the occurrence times of the corresponding word in the recognized text; the recognized text changes with the change of the voice data to be processed;
[0109] Determine the words that meet the second preset condition in the segmentation result as keywords;
[0110] Update the first keyword table according to the keywords that meet the second preset condition and the weights corresponding to the keywords; the keywords correspond to at least one language.
[0111] Here, as the speech data to be processed continuously changes, the recognized text continuously changes, and the word segmentation results obtained based on the recognized text also continuously change, so that the occurrence times of corresponding words continuously change; in this embodiment, the weight of a word is related to the occurrence times, so that the weight of the word changes as the speech data to be processed continuously changes.
[0112] The following makes a specific description of the first keyword table.
[0113] The words in the first keyword table are statistically counted in units of n-gram (n represents the number of word elements, with a maximum of 3). For example: the number of word elements of "machine" is 1; "machine translation" is composed of the words "machine" and "translation", and its number of word elements is 2; "machine translation evaluation" is composed of the words "machine", "translation", and "evaluation", and its number of word elements is 3.
[0114] Accumulate the occurrence times of each word in the first keyword table, and store the occurrence times as a global variable in the first keyword table. Each word can correspond to 3 attributes:
[0115] Frequency attribute (i.e., occurrence times), built-in value attribute (the built-in value is related to the number of word elements. In an example, the value of 1-gram can be 1, the value of 2-gram is 3, and the value of 3-gram is 5), weight attribute (weight value = frequency * built-in value).
[0116] The format of the first keyword table can be: n-gram (representing the word), freq (representing the frequency attribute), value (representing the built-in value attribute), weight (representing the weight attribute).
[0117] For example: the first keyword table may include:
[0118] Machine (i.e., n-gram), 20 (i.e., freq), 1 (i.e., value), 20 (i.e., weight); corresponding to at least one language, for example, English: Machine;
[0119] Machine Translation (i.e., n-gram), 12 (i.e., freq), 3 (i.e., value), 36 (i.e., weight); corresponding to at least one language, for example, English: Machine Translation;
[0120] Machine Translation Evaluation (i.e., n-gram), 4 (i.e., freq), 5 (i.e., value), 20 (i.e., weight); corresponding to at least one language, for example, English: Machine Translation Evaluation.
[0121] It should be noted that considering that the frequency of low-order grams must be higher than that of high-order grams, for example, the frequency corresponding to "machine" (a low-order gram) must be higher than the frequencies corresponding to "machine translation" and "machine translation evaluation" (high-order grams). And many terms are high-order grams, although there are also some terms that are low-order grams. Therefore, when both high-order grams and low-order grams match, the target segment can be selected based on the weight, that is, when the target segment matches at least two keywords, the target segment is determined according to the keyword with the higher weight.
[0122] Specifically, determining the words in the word segmentation result that meet the second preset condition includes at least one of the following:
[0123] Determining the words in the word segmentation result whose weight exceeds the preset weight threshold;
[0124] Determining the words in the word segmentation result whose occurrence times exceed the preset occurrence threshold.
[0125] Here, the preset weight threshold and the preset occurrence threshold can be preset and stored in the server.
[0126] Specifically, each keyword in the first keyword table corresponds to a font change factor, and the font change factor is related to the weight;
[0127] Determining the first presentation format of the target segment includes:
[0128] When the target keyword table corresponding to the target segment is the first keyword table, determining the format corresponding to the font change factor as the first presentation format.
[0129] Here, considering that the keywords in the first keyword table and the weights of each keyword are constantly changing, the weight can be mapped into a factor used to change the font of the keyword, that is, the font change factor. The font change factor can be a decimal or an integer (for example: with a step of 0.5, specifically, numbers such as 0.5, 1.0, 1.5, 2.0, etc. can be used); during the continuous change of the voice data, as the weight of the keyword changes, the size of the font also changes correspondingly. Here, the font change factor specifically refers to the font size by which the target segment needs to be enlarged; assuming that the original font size (i.e., the second presentation format) of the recognized text is 2 and the determined font change factor is 1.0; then the first presentation format is: the font size is 3. The size of the font can have a maximum limit, and the font size will no longer change after reaching the maximum limit.
[0130] It should be noted that the data processing method can be applied to the simultaneous interpretation scenario of a meeting. During the meeting, the voice data to be processed is constantly changing. Correspondingly, the recognized text is constantly changing, and thus the word segmentation results obtained based on the recognized text are also constantly changing. By using the method of this embodiment, the first keyword table can be continuously updated based on the word segmentation results. After the meeting ends and the update of the first keyword table is completed, the first keyword table can be deleted from the keyword library to save storage space.
[0131] In practical applications, in order to match the recognized text in at least one language, the keywords in the first keyword table also need to correspond to at least one language, so as to determine the target segments included in the recognized text in different languages and present them in the first presentation format.
[0132] Based on this, in one embodiment, the method further includes:
[0133] After determining the keywords, use a preset translation engine to translate the keywords to obtain keywords in other languages.
[0134] Correspondingly, the updating of the first keyword table according to the keywords that meet the second preset condition and the weights corresponding to the keywords includes:
[0135] Update the first keyword table according to the keywords, the keywords in other languages, and the weights corresponding to the keywords.
[0136] Here, for each keyword, there can be corresponding keywords in the first language, the second language,..., the Nth language; there is a corresponding relationship between the language of the recognized text and the language of the keyword, and the first language is the language corresponding to the voice data.
[0137] It should be noted that in order to determine the target segments in the recognized text in any language, the recognized text in the same language as the voice data (i.e., the first language) can be segmented to obtain at least one keyword, the keyword can be translated to obtain the translation results corresponding to each keyword, and each keyword and the translation result corresponding to the keyword are stored in the keyword table in a corresponding manner; thus, for the recognized text in any language, the target segments can be determined by querying the keyword table. Here, translating the keyword means translating the keyword in the same language as the voice data (i.e., the first language) to obtain keywords in the second language,..., the Nth language.
[0138] First, the recognized text in the first language is tokenized to determine keywords. After determining the keywords, a preset translation engine is used to translate the keywords to obtain keywords in other languages, considering that the translation engine is more accurate in translating short content, thereby improving the accuracy of each keyword in the keyword list.
[0139] Of course, in order to determine the target segment of the recognized text in any language, the recognized text in any language can also be tokenized separately to obtain the tokenization result corresponding to the recognized text in that language, and the keyword list is updated based on the tokenization result; that is, each recognized text in a language corresponds to a keyword list in the corresponding language; there is no limitation here.
[0140] In practical applications, in order to specially display professional terms (a type of keyword), a keyword list containing professional terms can be preset in advance to determine the professional terms that need to be specially displayed in the recognized text.
[0141] Based on this, in one embodiment, the method further includes:
[0142] Extract terms from the bilingual data of the machine translation model, and generate a second keyword list based on the extracted terms; the second keyword list is used as a keyword list in the keyword library.
[0143] Here, methods such as text-reranking, Bootstrapping, and deep learning can be combined for term extraction, and there is no limitation on the method of term extraction.
[0144] The format of the second keyword list is: keyword, weight; the keyword corresponds to at least one language. Taking keywords in two languages as an example, the second keyword list includes:
[0145] Machine translation (i.e., the word in the first language), machine translation (i.e., the corresponding word in the second language), 0.03 (i.e., the weight);
[0146] Speech recognition (i.e., the word in the first language), automatic speech recognition (i.e., the corresponding word in the second language), 0.02 (i.e., the weight).
[0147] In another embodiment, the method further includes: receiving the keywords manually set and the weights corresponding to the keywords, and generating a third keyword list based on the keywords manually set and the weights corresponding to the keywords.
[0148] The format of the third keyword list can be: keyword, weight; the keyword corresponds to at least one language. Taking keywords in two languages as an example, the second keyword list may include:
[0149] Penicillin (i.e., the word in the first language), Penicillin (i.e., the corresponding word in the second language), 0.5 (weight).
[0150] Here, the third keyword list is different from the first keyword list and the second keyword list. The third keyword list is set by professional and technical personnel in the corresponding field according to their experience. This is because in each field, there are certain professional terms. For example, in the fields of medicine, aerospace, real estate, etc., the keywords set by professional and technical personnel in the field are more authoritative and accurate. The priority of the third keyword list can be higher than that of the first keyword list, and the priority of the first keyword list can be higher than that of the second keyword list.
[0151] It should be noted that during simultaneous interpretation, the keywords in the second keyword list and the third keyword list will not change, but the keywords in the first keyword list will change continuously with the change of voice data. After the simultaneous interpretation is completed, the second keyword list and the third keyword list are still stored in the keyword library. The first keyword list can be deleted from the keyword library to save storage space; of course, the first keyword list can also be saved corresponding to the recognized text to facilitate users to organize files, which is not limited here.
[0152] In addition, in order to determine the target segment in the recognized text of any language, it should be understood that for the second keyword list and the third keyword, the keywords included therein can be translated into texts in other languages to obtain translation results, and the keywords and the corresponding translation results are stored in the keyword list correspondingly. Thus, for the recognized text of any language, the target segment can be determined by querying the keyword list.
[0153] The data processing method provided by the embodiments of the present application can be specifically applied to the simultaneous interpretation scenario, such as the simultaneous interpretation of a meeting. In this scenario, the speaker gives a speech, the server obtains the voice data of the speaker, performs text recognition on the voice data to obtain the recognized text; uses the keyword library to determine the target segment in the recognized text, and highlights the target segment (i.e., presents it in the first presentation format) to help users more directly determine the key points of the speech and the professional terms mentioned in the speech; thus helping users better accept the speech content.
[0154] It should be understood that the order of describing each step (such as generating the first keyword list, generating the second keyword list, generating the third keyword list, etc.) in the above embodiments does not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0155] The data processing method provided by the embodiment of this application obtains the voice data to be processed, performs text recognition on the voice data to obtain the recognized text; the recognized text is used for presentation when the voice data is played; searches for a keyword library according to the recognized text, and determines a target segment in the recognized text that meets the first preset condition; determines the first presentation format of the target segment, so as to present the target segment in the first presentation format when presenting the recognized text; the first presentation format is different from the second presentation format; the second presentation format is the presentation format of other texts in the recognized text except the target segment. In this way, key information can be extracted from the content of the voice data and highlighted, so that users can intuitively understand the key information of the voice content, help users better accept the speech content, and improve the user experience.
[0156] Figure 3 Another schematic flowchart of the data processing method according to the embodiment of this application; as Figure 3 shown, the method includes:
[0157] Step 301: Pre-generate a keyword library;
[0158] Here, the method can be applied to servers, mobile terminals, cloud devices, etc.
[0159] In practical applications, in order to determine the target segment according to multiple criteria (for example, it can be technical terms in the recognized text, repeatedly mentioned content, etc.), the keyword library can be composed of at least one keyword table.
[0160] Based on this, in one embodiment, the keyword library may include a term list T1;
[0161] The step 301 includes: extracting terms from the large-scale bilingual data of the machine translation model, and generating a term list T1 according to the extracted terms; the term list T1 is used as a keyword table in the keyword library.
[0162] The term list T1 is equivalent to Figure 2 the second keyword table in the method shown, and each term has a weight. The format of the term list T1 can be as shown in Table 1 below:
[0163] Word in the first language Word in the second language Weight Machine translation machine translation 0.03 Speech recognition automatic speech recognition 0.02
[0164] Table 1
[0165] Here, the keyword library may further include a term list T2;
[0166] The step 301 includes: obtaining a manually maintained term list T2 as a keyword table in the keyword library.
[0167] Here, considering that each field has certain professional terms (including abbreviations of terms, etc.), such as in the fields of medicine, aerospace, real estate, etc., manual maintenance of the terms in the corresponding fields has higher accuracy. Therefore, a term list T2 is provided.
[0168] The term list T2 is equivalent to Figure 2 the third keyword table in the method shown. Each term has a weight, and its format can be as shown in Table 2 below:
[0169] Word in the first language Word in the second language Weight Penicillin Penicillin 0.5
[0170] Table 2
[0171] Step 302: Determine the speech data during simultaneous interpretation, perform text recognition on the speech data, and obtain the recognized text.
[0172] Here, the step 302 includes: obtaining the speech data of the speaker (denoted as S); performing text recognition on the speech data to obtain the recognized text.
[0173] The recognized text includes: text in the same language as the speech data (denoted as text T), and translation texts in other languages obtained after translating text T (denoted as text R). There can be multiple translation texts, that is, multiple translation texts in various languages are obtained after translating text T.
[0174] Step 303: Search the keyword library according to the recognized text to obtain the target segment, determine the first presentation format of the target segment, so as to present the target segment in the first presentation format when presenting the recognized text;
[0175] Here, the first presentation format is different from the second presentation format; the second presentation format is the presentation format of the other text in the recognized text except the target segment.
[0176] Here, for text T, the step 303 includes:
[0177] Step 3031: Search the term list T1, the term list T2, and the term list D according to text T;
[0178] Here, the term list D is updated according to the speech data to be processed, and the keywords and the weights of the keywords change with the change of the speech data to be processed; the priority of the term list T2 is higher than that of the term list D, and the priority of the term list D is higher than that of the term list T1;
[0179] Step 3032: When some segments in text T exist in the term list T2, present the font of the included segments in the first presentation format corresponding to the term list T2;
[0180] Here, the first presentation format can be F + 4 (i.e., the font size plus 4, where F is the initial font size of the text), and it is marked in red;
[0181] Step 3033: When a partial segment in text T exists in the glossary D, present the font of the included segment in accordance with the first presentation format corresponding to the glossary D;
[0182] Here, the first presentation format can be F + 3 (i.e., the font size plus 3); the first presentation format can also include color setting for the text, such as marking the color of the text as green to highlight the included segment;
[0183] It should be noted that if the segment existing in the glossary D also exists in the glossary T2, then this segment is presented in accordance with the first presentation format corresponding to the glossary T2.
[0184] Step 3034: When a partial segment in text T exists in the glossary T1, present the font of the included segment in accordance with the first presentation format corresponding to the glossary T1;
[0185] Here, the first presentation format can be F + 2 (i.e., the font size plus 2); the first presentation format can also be color setting for the text, such as marking the color of the text as blue to highlight the included segment.
[0186] It should be noted that if the segment existing in the glossary T1 also exists in the glossary T2, then this segment is presented in accordance with the first presentation format corresponding to the glossary T2; it should be noted that if the segment existing in the glossary T1 also exists in the glossary D but does not exist in the glossary T2, then this segment is presented in accordance with the first presentation format corresponding to the glossary T2.
[0187] The operation for text R is the same as the above operation for text T, and reference can be made to steps 3031 - 3034, which will not be elaborated here.
[0188] Here, updating the glossary D according to the to - be - processed voice data may include:
[0189] Segment text T to obtain at least one word; filter out the words that are the same as each stop word in the preset stop - word list from the at least one word, and use the words obtained after filtering as the segmentation result; update the glossary D based on the segmentation result.
[0190] Here, a stop word list is used to filter at least one word obtained by word segmentation. This is because during simultaneous interpretation, the content of the speaker is relatively scarce. Judging keywords directly based on the repetition rate of the text provides too little information, and there is a lot of noise in the extracted keywords. Filtering at least one word obtained by word segmentation through the stop word list can reduce keyword noise.
[0191] Here, T and R can be separated, and only T is subjected to word segmentation operations to obtain the term list D; then, a translation engine is used to translate each word in the term list D. This is because the translation engine is more accurate in translating short content.
[0192] The words in the term list D are statistically analyzed in units of n-gram (n is at most 3). The description of n-gram has been specifically described in the Figure 2 method shown, and will not be elaborated here.
[0193] The term list D is equivalent to Figure 2 the first keyword list in the method shown. The method for updating the term list D can refer to the Figure 2 method for updating the first keyword list in, and will not be elaborated here.
[0194] As the simultaneous interpretation process progresses, the keywords in the term list T1 and the term list T2 will not change, but the keywords in the term list D are constantly changing, that is, the attributes of the words (specifically referring to the frequency attribute and the weight attribute) are also changing. These changes in attributes can also be reflected by a method. Specifically, the weight can be mapped into a font change factor, which is used as the factor for magnifying the keyword; the font change factor can be a decimal or an integer (assuming a step of 0.5, the font change factor can be 0.5, 1.0, 1.5, 2.0, etc.). During the simultaneous interpretation process, according to the font change factor, the keywords in the text will be gradually magnified. Of course, there is a maximum limit for the font size, and it will no longer change if it exceeds the maximum limit.
[0195] Through the above solution, the simultaneous interpretation display screen in front of the exhibition booth receives and presents the speech recognition result (such as text T) and the machine translation result (text R) of the speaker. In presenting the above results, some texts will be displayed in different colors and different font sizes (different colors and different font sizes can represent target segments determined based on different term lists. For example, the term list T2 is a manually maintained keyword list with the highest credibility, and the font size of the target segment determined based on the term list T2 can also be the largest), so as to prominently remind the audience.
[0196] The data processing method provided by the present application can determine the key information (such as the above-mentioned terms) in the recognized text in a simultaneous interpretation scenario, and display the key information in the speaker's speech by changing its font size and color, so as to prominently remind the user and enable the user to capture the main content of the speaker in a short time; in this way, the user can have a general understanding of the content of the speech without having to watch all the content on the entire screen, which is particularly suitable for scenarios where the speaker speaks faster.
[0197] Figure 4 FIG. 1 is a flow chart of a method for determining a first presentation format according to an embodiment of the present application; Figure 4 As shown, the method includes:
[0198] Step 401: when determining a target segment in the recognized text that meets a first preset condition, determining a candidate keyword table corresponding to the target segment;
[0199] Here, the candidate keyword table includes keywords that match the target segment;
[0200] Step 402: Determine the number of the candidate keyword tables. When the number of the candidate keyword tables is one, execute step 403; when the number of the candidate keyword tables is at least two, execute step 404;
[0201] Step 403: taking the candidate keyword table as the target keyword table, and taking the format corresponding to the candidate keyword table as the first presentation format.
[0202] Step 404: Determine the priority corresponding to each of the at least two candidate keyword tables, sort the at least two candidate keyword tables according to the priority, and determine the candidate keyword table with the highest priority; and use the format corresponding to the candidate keyword table with the highest priority as the first presentation format.
[0203] It should be noted that when each of the at least two keyword tables corresponds to a different format, and each of the at least two keyword tables corresponds to a different priority, the method described in step 404 can be used. When the keyword library includes at least two keyword tables, the format corresponding to the candidate keyword table with the highest priority is used as the first presentation format. If the keyword tables in the keyword library correspond to the same format, there is no need to use the operation of step 404, but directly select the format corresponding to any one of the candidate keyword tables as the first presentation format.
[0204] In order to implement the data processing method of the embodiment of the present application, the embodiment of the present application also provides a data processing device. Figure 5 Schematic diagram of the structure of the data processing device according to the embodiment of the present application;Figure 5 As shown, the data processing device includes:
[0205] An acquisition unit 51, configured to obtain voice data to be processed, perform text recognition on the voice data, and obtain recognition text; the recognition text is used for presentation when the voice data is played;
[0206] A first processing unit 52, configured to search a keyword library according to the recognition text and determine a target segment in the recognition text that meets a first preset condition;
[0207] A second processing unit 53, configured to determine a first presentation format of the target segment, so as to present the target segment in the first presentation format when presenting the recognition text; the first presentation format is different from a second presentation format; the second presentation format is the presentation format of other texts in the recognition text except the target segment.
[0208] In one embodiment, the first processing unit 52 is configured to use at least one of the following methods to determine a target segment in the recognition text that meets a first preset condition:
[0209] Determine a target segment in the recognition text that matches any keyword in the keyword library;
[0210] Determine at least two keywords from the recognition text; determine the target segment based on the weights of the keywords in the at least two keywords.
[0211] In one embodiment, the second processing unit 53 is configured to determine a target keyword table corresponding to the target segment; the target keyword table includes keywords that match the target segment; use the format corresponding to the target keyword table as the first presentation format.
[0212] Here, the keyword library includes at least one keyword table.
[0213] Here, the keyword library may include at least two keyword tables; each keyword table in the at least two keyword tables corresponds to a different format; each keyword table in the at least two keyword tables corresponds to a different priority.
[0214] In one embodiment, the second processing unit 53 is configured to determine at least two candidate keyword tables corresponding to the target segment;
[0215] Use the candidate keyword table with a higher priority in the at least two candidate keyword tables as the target keyword table.
[0216] In one embodiment, the device further includes a third processing unit, configured to perform word segmentation on the recognition text to obtain at least one word;
[0217] Filter the at least one word, and use the word obtained after filtering as the word segmentation result;
[0218] Update the first keyword table based on the word segmentation result; the first keyword table is one of the keyword tables in the keyword library; the keywords and keyword weights in the first keyword table change with the change of the speech data to be processed.
[0219] Here, the third processing unit is specifically configured to determine the occurrence times and word elements of each word in the word segmentation result;
[0220] Determine the weight of the corresponding word based on the occurrence times and the word elements; the weight changes with the change of the occurrence times of the corresponding word in the recognition text; the recognition text changes with the change of the speech data to be processed;
[0221] Determine the words in the word segmentation result that meet the second preset condition as keywords;
[0222] Update the first keyword table according to the keywords that meet the second preset condition and the weights corresponding to the keywords; the keywords correspond to at least one language.
[0223] Here, determining the words in the word segmentation result that meet the second preset condition includes at least one of the following:
[0224] Determine the words in the word segmentation result whose weights exceed the preset weight threshold;
[0225] Determine the words in the word segmentation result whose occurrence times exceed the preset occurrence threshold.
[0226] In one embodiment, each keyword in the first keyword table corresponds to a font change factor, and the font change factor is related to the weight.
[0227] The second processing unit 53 is configured to determine the format corresponding to the font change factor as the first presentation format when the target keyword table corresponding to the target segment is the first keyword table.
[0228] In one embodiment, the device further includes a fourth processing unit configured to extract terms from the bilingual data of the machine translation model and generate a second keyword table based on the extracted terms; the second keyword table is one of the keyword tables in the keyword library.
[0229] In practical applications, the obtaining unit 51 can be implemented through a communication interface; the first processing unit 52, the second processing unit 53, the third processing unit, and the fourth processing unit can all be implemented by a processor in the server, such as a central processing unit (CPU), a digital signal processor (DSP), a microcontroller unit (MCU), or a field-programmable gate array (FPGA), etc.
[0230] It should be noted that: when the device provided in the above embodiment performs data processing, only the division of the above program modules is used for illustration. In practical applications, the above processing can be allocated to different program modules according to needs, that is, the internal structure of the terminal is divided into different program modules to complete all or part of the above-described processing. In addition, the device provided in the above embodiment and the data processing method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be elaborated here.
[0231] Based on the hardware implementation of the above device, an embodiment of the present application further provides a server. Figure 6 It is a schematic diagram of the hardware composition structure of the server according to the embodiment of the present application. As Figure 6 shown, the server 60 includes a memory 63, a processor 62, and a computer program stored on the memory 63 and executable on the processor 62; when the processor 62 located in the server executes the program, it implements the method provided by one or more technical solutions on the server side.
[0232] Specifically, when the processor 62 located in the server 60 executes the program, it implements: obtaining voice data to be processed, performing text recognition on the voice data to obtain recognition text; the recognition text is used for presentation when playing the voice data; searching a keyword library according to the recognition text to determine a target segment in the recognition text that meets a first preset condition; determining a first presentation format of the target segment, so as to present the target segment in the first presentation format when presenting the recognition text; the first presentation format is different from a second presentation format; the second presentation format is the presentation format of other texts in the recognition text except the target segment.
[0233] It should be noted that the specific steps implemented when the processor 62 located in the server 60 executes the program have been detailed above and will not be elaborated here.
[0234] It can be understood that the server further includes a communication interface 61; each component in the server is coupled together through a bus system 64. It can be understood that the bus system 64 is configured to implement connection and communication between these components. In addition to including a data bus, the bus system 64 further includes a power bus, a control bus, a status signal bus, etc.
[0235] It can be understood that the memory 63 in this embodiment may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM, Read-Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory may be a disk memory or a tape memory. The volatile memory may be a random access memory (RAM, RandomAccess Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM, Static Random Access Memory), synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), dynamic random access memory (DRAM, Dynamic Random Access Memory), synchronous dynamic random access memory (SDRAM, SynchronousDynamic Random Access Memory), double data rate synchronous dynamic random access memory (DDRSDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memories described in the embodiments of the present application are intended to include, but are not limited to, these and any other suitable types of memories.
[0236] The method disclosed in the embodiments of the present application above can be applied to, or implemented by, the processor 62. The processor 62 may be an integrated circuit chip with the ability to process signals. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 62 or by instructions in the form of software. The above-mentioned processor 62 may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 62 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present application, it can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, and this storage medium is located in the memory. The processor 62 reads the information in the memory and combines its hardware to complete the steps of the foregoing method.
[0237] The embodiments of the present application also provide a storage medium, specifically a computer storage medium, and more specifically a computer-readable storage medium. A computer instruction, that is, a computer program, is stored thereon, and when the computer instruction is executed by a processor, the method provided by one or more of the above technical solutions on the server side is implemented.
[0238] In several embodiments provided by the present application, it should be understood that the disclosed methods and intelligent devices can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed with each other may be through some interfaces, and the indirect coupling or communication connection of the devices or units may be electrical, mechanical, or other forms.
[0239] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0240] In addition, in each embodiment of the present application, each functional unit can be entirely integrated into a second processing unit, or each unit can be separately regarded as a unit alone, or two or more units can be integrated into one unit; the above-mentioned integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0241] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as removable storage devices, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0242] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. And the foregoing storage medium includes: various media such as removable storage devices, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0243] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence.
[0244] In addition, among the technical solutions described in the embodiments of the present application, they can be combined arbitrarily without conflict.
[0245] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, and all of them should be covered by the protection scope of the present application.
Claims
1. A data processing method, applied to a server, includes: Obtain the speech data to be processed, perform text recognition on the speech data, and obtain the recognized text; The recognized text is used for presentation when the speech data is played; Search the keyword library according to the recognized text, and determine the target segment in the recognized text that meets the first preset condition; Determine the first presentation format of the target segment, so as to present the target segment in the first presentation format when presenting the recognized text; The first presentation format is different from the second presentation format; The second presentation format is the presentation format of other texts in the recognized text except the target segment; Wherein, the keyword library includes at least two keyword tables; each keyword table in the at least two keyword tables corresponds to a different format; Each keyword table in the at least two keyword tables corresponds to a different priority.
2. The method according to claim 1, wherein, The determining the target segment in the recognized text that meets the first preset condition includes at least one of the following: Determine the target segment in the recognized text that matches any keyword in the keyword library; Determine at least two keywords from the recognized text; determine the target segment based on the weights of the keywords in the at least two keywords.
3. The method according to claim 1, wherein, The determining the first presentation format of the target segment includes: Determine the target keyword table corresponding to the target segment; the target keyword table includes keywords that match the target segment; Take the format corresponding to the target keyword table as the first presentation format.
4. The method according to claim 3, wherein, The determining the target keyword table corresponding to the target segment includes: Determine at least two candidate keyword tables corresponding to the target segment; Take the candidate keyword table with the highest priority in the at least two candidate keyword tables as the target keyword table.
5. The method according to claim 1, wherein, The method further includes: Perform word segmentation on the recognized text to obtain at least one word; Filter the at least one word, and take the words obtained after filtering as the word segmentation result; Update the first keyword table based on the word segmentation result; the first keyword table is a keyword table in the keyword library; the keywords and the weights of the keywords in the first keyword table change with the change of the speech data to be processed.
6. The method according to claim 5, wherein, The updating the first keyword table based on the word segmentation result includes: For each word in the word segmentation result, determine the occurrence times and the number of word elements of the corresponding word; Determine the weight of the corresponding word based on the occurrence times and the number of word elements; the weight changes with the change of the occurrence times of the corresponding word in the recognized text; the recognized text changes with the change of the speech data to be processed; Determine the words in the word segmentation result that meet the second preset condition as keywords; Update the first keyword table according to the keywords that meet the second preset condition and the weights corresponding to the keywords; the keywords correspond to at least one language.
7. The method according to claim 6, wherein, The determining the words in the word segmentation result that meet the second preset condition includes at least one of the following: Determine the words in the word segmentation result whose weights exceed the preset weight threshold; Determine the words in the word segmentation result whose occurrence times exceed the preset occurrence times threshold.
8. The method according to claim 5, wherein, Each keyword in the first keyword table corresponds to a font change factor, and the font change factor is related to the weight; Determining the first presentation format of the target segment includes: When the target keyword table corresponding to the target segment is the first keyword table, determining the format corresponding to the font change factor as the first presentation format.
9. The method according to claim 1, wherein, The method further includes: Performing term extraction on the bilingual data of the machine translation model, and generating a second keyword table based on the extracted terms; the second keyword table is one of the keyword tables in the keyword library.
10. A data processing device, comprising: An acquisition unit, configured to obtain the speech data to be processed, perform text recognition on the speech data, and obtain the recognized text; The recognized text is used for presentation when the speech data is played; A first processing unit, configured to search the keyword library according to the recognized text, and determine a target segment in the recognized text that meets the first preset condition; wherein, the keyword library includes at least two keyword tables; each keyword table in the at least two keyword tables corresponds to a different format; each keyword table in the at least two keyword tables corresponds to a different priority; A second processing unit, configured to determine the first presentation format of the target segment, so as to present the target segment in the first presentation format when presenting the recognized text; the first presentation format is different from the second presentation format; the second presentation format is the presentation format of other texts in the recognized text except the target segment.
11. A server, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the method according to any one of claims 1 to 9 are implemented.
12. A storage medium, having computer instructions stored thereon, wherein when the instructions are executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Voice interaction method and device, computer equipment and storage medium
CN109658931A
Text display method and device
CN110263149A