A computer-aided input method, device, computer equipment and storage medium
By obtaining tone data in real time and generating real-time subtitles using keyword subtitles library, it solves the problem that participants find it difficult to clearly obtain speech content in online meetings, and improves the quality and efficiency of the meeting.
Patent Information
- Application Number
- CN202210562992.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2042-05-23
AI Technical Summary
During the online meeting, it is difficult for participants to clearly obtain the speech content due to their fast speech speed or accent, which affects the quality of the meeting.
By obtaining the tone data in the voice content of the video conference in real time, using the preset keyword subtitle library for matching query, and combining with the conference user ID filtering, real-time subtitles are generated.
It realizes real-time subtitles generated in online meetings, improves the experience of participants, and can record meeting minutes to improve the real-time and accuracy of subtitles.
Smart Images

Figure CN114974225B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer technology, and in particular to a computer-assisted input method, device, computer equipment and storage medium. Background Art
[0002] Currently, due to geographical restrictions, companies or meetings cannot be held in person, and online meetings need to be conducted through video or voice.
[0003] Before holding an online meeting, the organizer of the meeting creates an online conference room and notifies the participants to log in to the conference room on time to attend the meeting. During the online meeting, the content and images of the participants' speeches are obtained by connecting to the computer or the computer's own audio and video input device, and shared on each participant's terminal.
[0004] The above-mentioned prior art solutions have the following defects:
[0005] During an online meeting, if the user speaks too fast or with an accent, other participants may not be able to clearly understand what the user is saying, affecting the quality of the meeting. Summary of the Invention
[0006] In order to generate subtitles in real time during online meetings, the present application provides a computer-assisted input method, apparatus, computer device, and storage medium.
[0007] The above-mentioned invention objective of this application is achieved through the following technical solutions:
[0008] A computer-assisted input method, the computer-assisted input method being applied to a device connected to a computer and transmitting data, comprising:
[0009] Acquire the video conference voice content in real time, and acquire the tone data in the video conference voice content;
[0010] Performing a matching query from a preset keyword and subtitle library based on the tone data to obtain keyword data to be screened and common voice tone data;
[0011] Obtaining a corresponding conference user identifier from the video conference voice content, and filtering the keyword data to be filtered according to the conference user identifier to obtain keyword subtitles;
[0012] Based on the common voice tone data and the keyword subtitles, real-time conference subtitles corresponding to the tone data are generated and displayed on the user terminal.
[0013] By adopting the above technical solution, the inventor found that it is possible to add real-time subtitles during online meetings to solve the problem that participants cannot clearly understand what the user who is speaking is saying. However, in the process of recognizing text through voice, it takes a long time and the real-time nature of the recognized subtitles cannot be achieved. That is, when the text content is recognized, the subtitles displayed are already the content that the user who is speaking has said before. Therefore, the keyword subtitle library is pre-built, and when the conference voice content is obtained in real time, the tone data is obtained from the conference voice content, and the tone data is used to perform a matching query in the keyword subtitle library, so that the matching In this way, proper nouns or nouns frequently said by users are matched from the keyword subtitle library, which reduces the degree of text recognition from voice. At the same time, the matched keywords to be screened are screened through the corresponding conference user ID, and the keyword subtitles can be determined. After being pieced together with the commonly used voice tone data to obtain the real-time subtitles of the conference, the efficiency of obtaining real-time subtitles can be improved, thereby realizing the generation of corresponding real-time subtitles during online meetings and improving the experience of participants in online meetings. At the same time, when obtaining real-time subtitles of the meeting, it can not only improve the experience of participants in the meeting, but also record the real-time subtitles of the meeting, which is helpful to generate corresponding meeting minutes.
[0014] In a preferred example, the present application may be further configured as follows: in performing a matching query from a preset keyword subtitle library according to the tone data, constructing the keyword subtitle library includes:
[0015] Acquire group domain data within a preset group of people, and acquire domain keyword information based on the group domain data, wherein the domain keyword information includes a plurality of keyword words;
[0016] Obtain historical meeting text data of the group of people, perform a matching query in the historical meeting text data based on the keyword words in the field keyword information, and obtain group keyword words and keyword frequency corresponding to each group keyword word;
[0017] After splitting each of the group keyword words into individual Chinese words, obtaining the keyword tone corresponding to each of the Chinese words, and generating a tone key value according to the keyword tone;
[0018] After associating the tone key value with the corresponding group keyword word, the keyword subtitle library is obtained after sorting the group keyword words according to the keyword word frequency.
[0019] By adopting the above technical solution, by determining the attributes of the preset personnel group, that is, the group field data, the corresponding field keyword information is obtained, and the keyword subtitle library can be specifically constructed based on the attributes of the preset personnel group, thereby streamlining the number of samples in the keyword subtitle library to improve the efficiency of matching the keywords to be screened in the actual use of the keyword subtitle library. At the same time, since the data structure of the tone characteristics of the speech is relatively complex, converting the keyword tone into a tone key value can simplify the data structure in the keyword subtitle library and improve the efficiency of recognition and matching.
[0020] In a preferred example, the present application can be further configured as follows: performing a matching query from a preset keyword subtitle library based on the tone data to obtain keyword data to be screened, specifically including:
[0021] Calculating a key value to be matched for the tone data;
[0022] The key value to be matched is input into the keyword subtitle library, and a matching query is performed in sequence. If the corresponding tone key value can be successfully matched, the associated group keyword is used as the keyword data to be screened.
[0023] By adopting the above technical solution, since the data structure of the characteristic points of the tone is relatively complex, the tone data is converted into the key value to be matched, so that in the process of matching and identifying in the keyword subtitle library, the feature matching can be converted into the comparison of the inherited character string, thereby improving the matching efficiency.
[0024] In a preferred example, the present application may be further configured as follows: calculating the key value to be matched of the tone data specifically includes:
[0025] Splitting the tone data to obtain word tone data;
[0026] The key values to be matched are calculated one by one according to the order of the tone data.
[0027] By adopting the above technical solution, since users usually speak a sentence coherently, by first splitting the acquired tone data, the sentence spoken by the user can be recognized, thereby improving the real-time performance of subtitles.
[0028] In a preferred example, the present application may be further configured as follows: generating and displaying real-time conference subtitles corresponding to the tone data on the user side based on the common voice tone data and the keyword subtitles, specifically including:
[0029] splicing the common voice tone data and the keyword subtitles according to the order of the tone data to obtain subtitle data to be recognized;
[0030] After semantic recognition is performed on the subtitle data to be recognized, the real-time subtitles of the conference are obtained.
[0031] By adopting the above technical solution and performing semantic recognition on the subtitle data to be recognized, it is possible to verify whether the text subtitles corresponding to the recognized sentences are fluent, and the degree of fit between the obtained real-time subtitles of the meeting and the content spoken by the user can be improved.
[0032] The second object of the present invention is achieved through the following technical solutions:
[0033] A computer-assisted input device, comprising:
[0034] A tone extraction module is used to obtain the video conference voice content in real time and obtain the tone data in the video conference voice content;
[0035] A subtitle matching module is used to perform a matching query from a preset keyword subtitle library based on the tone data to obtain keyword data to be screened and common voice tone data;
[0036] A subtitle screening module is used to obtain a corresponding conference user identifier from the video conference voice content, and screen the keyword data to be screened according to the conference user identifier to obtain keyword subtitles;
[0037] The subtitle generation module is used to generate and display the conference real-time subtitles corresponding to the tone data on the user end based on the common voice tone data and the keyword subtitles.
[0038] By adopting the above technical solution, the inventor found that it is possible to add real-time subtitles during online meetings to solve the problem that participants cannot clearly understand what the user who is speaking is saying. However, in the process of recognizing text through voice, it takes a long time and the real-time nature of the recognized subtitles cannot be achieved. That is, when the text content is recognized, the subtitles displayed are already the content that the user who is speaking has said before. Therefore, the keyword subtitle library is pre-built, and when the conference voice content is obtained in real time, the tone data is obtained from the conference voice content, and the tone data is used to perform a matching query in the keyword subtitle library, so that the matching In this way, proper nouns or nouns frequently said by users are matched from the keyword subtitle library, which reduces the degree of text recognition from voice. At the same time, the matched keywords to be screened are screened through the corresponding conference user ID, and the keyword subtitles can be determined. After being pieced together with the commonly used voice tone data to obtain the real-time subtitles of the conference, the efficiency of obtaining real-time subtitles can be improved, thereby realizing the generation of corresponding real-time subtitles during online meetings and improving the experience of participants in online meetings. At the same time, when obtaining real-time subtitles of the meeting, it can not only improve the experience of participants in the meeting, but also record the real-time subtitles of the meeting, which is helpful to generate corresponding meeting minutes.
[0039] The third objective of this application is achieved through the following technical solutions:
[0040] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned computer-assisted input method when executing the computer program.
[0041] The fourth objective of this application is achieved through the following technical solutions:
[0042] A computer-readable storage medium stores a computer program, which implements the steps of the computer-assisted input method when executed by a processor.
[0043] In summary, this application includes at least one of the following beneficial technical effects:
[0044] 1. The keyword subtitle library is pre-built, and when the conference voice content is obtained in real time, the tone data is obtained from the conference voice content, and the tone data is used to perform a matching query in the keyword subtitle library. In this way, proper nouns or nouns frequently spoken by users can be matched from the keyword subtitle library in a matching manner, which reduces the degree of text recognition from voice. At the same time, the matched keywords to be screened are screened by the corresponding conference user ID to determine the keyword subtitles. After the keyword subtitles are pieced together with the commonly used voice tone data to obtain the real-time subtitles of the conference, the efficiency of obtaining the real-time subtitles can be improved, thereby realizing the generation of corresponding real-time subtitles during online meetings and improving the experience of the participants of the online meetings;
[0045] 2. When receiving real-time subtitles of a meeting, it not only improves the meeting experience of participants, but also enables them to record the real-time subtitles of the meeting, which helps to generate corresponding meeting minutes;
[0046] 3. By determining the attributes of the preset personnel group, that is, the group domain data, the corresponding domain keyword information is obtained, and a keyword subtitle library can be specifically constructed based on the attributes of the preset personnel group, thereby streamlining the number of samples in the keyword subtitle library to improve the efficiency of matching the keywords to be screened during the actual use of the keyword subtitle library. At the same time, since the data structure of the tone feature of the voice is relatively complex, converting the keyword tone into the tone key value can simplify the data structure in the keyword subtitle library and improve the efficiency of recognition and matching;
[0047] 4. By performing semantic recognition on the subtitle data to be recognized, it is possible to verify whether the text subtitles corresponding to the recognized sentences are fluent, which can improve the degree of fit between the obtained real-time conference subtitles and the content spoken by the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a flow chart of a computer-assisted input method according to an embodiment of the present application;
[0049] Figure 2 This is another implementation flow chart of the computer-assisted input method in one embodiment of the present application;
[0050] Figure 3 This is a flowchart for implementing step S21 in the computer-assisted input method in one embodiment of the present application;
[0051] Figure 4 This is a flowchart for implementing step S211 in the computer-assisted input method in one embodiment of the present application;
[0052] Figure 5 is a flowchart for implementing step S40 in the computer-assisted input method in one embodiment of the present application;
[0053] Figure 6 This is a principle block diagram of a computer-assisted input device in one embodiment of the present application;
[0054] Figure 7 It is a schematic diagram of a device in one embodiment of the present application. DETAILED DESCRIPTION
[0055] The present application is further described in detail below with reference to the accompanying drawings.
[0056] In one embodiment, if Figure 1 As shown, the present application discloses a computer-assisted input method, which specifically includes the following steps:
[0057] S10: Acquire the video conference voice content in real time, and acquire the tone data in the video conference voice content.
[0058] In this embodiment, the video conference voice content refers to the voice data triggered by the user during the video conference, and the tone data refers to the tone of the voice data triggered by the user in the video conference voice content.
[0059] Specifically, when a user conducts an online video conference or voice conference, if he needs to speak in the conference, he must have a voice receiving device connected to the computer or other mobile terminal. The device can be a connected microphone or headphones with a microphone, or it can be a built-in microphone in the device to obtain the voice data input by the user into the computer or other mobile terminal.
[0060] Since there may be multiple users speaking at the same time in an online meeting, in order to better identify the content of each user's speech and generate subtitles that match the speech content, the user can trigger the microphone-turning command in the video conferencing software and trigger the corresponding voice according to the topic of the current meeting. The video conference voice content can be obtained in real time through the user's microphone, and the corresponding tone data can be extracted from the video conference voice content.
[0061] S20: performing a matching query from a preset keyword subtitle library according to the tone data to obtain keyword data to be filtered and common voice tone data.
[0062] In this embodiment, the keyword subtitle database refers to a database storing Chinese words used for directionality recognition of the user's voice data in a video conference.
[0063] Specifically, the keyword subtitle library is pre-built, combined with Figure 2 , building the keyword subtitle library includes:
[0064] S201: Acquire group domain data within a preset group of people, and acquire domain keyword information based on the group domain data, wherein the domain keyword information includes a number of keyword words.
[0065] Specifically, when organizing an online meeting, participants may have the same or similar user attributes, such as belonging to the same unit, enterprise, organization, or a group of people established to do a certain thing. Therefore, the group domain data in this embodiment refers to the common attributes of the group of people, such as the main business direction of a unit or enterprise.
[0066] Furthermore, after obtaining the group field data, based on the characteristics of the group field data, such as the main business direction of the enterprise, the nouns associated with the characteristics are crawled to form the keyword information of the field. For example, if the main business direction is the production of motherboards of electronic equipment, then the keyword information of the field is professional nouns related to the motherboards, components, production equipment and production processes of electronic equipment, and keyword words are obtained to form the keyword information.
[0067] S202: Obtain historical meeting text data of a group of people, perform a matching query in the historical meeting text data based on the keyword words in the domain keyword information, and obtain group keyword words and the keyword frequency corresponding to each group keyword word.
[0068] In this embodiment, the group keyword words refer to nouns that have been hit in the domain keyword information of the group of people during the video conference.
[0069] Specifically, the historical meeting text data is obtained by collecting meeting minutes and recordings of the group of people and performing text recognition, as well as other methods. Furthermore, after obtaining the historical meeting text data, each keyword in the domain keyword information is used to perform a matching query within the historical meeting text data. Successfully matched keyword information is used as the group keyword word, and the number of successful matches for each group keyword word is counted as the corresponding keyword frequency.
[0070] S203: After splitting each group keyword word into individual Chinese words, obtain the keyword tone corresponding to each Chinese word, and generate a tone key value according to the keyword tone.
[0071] Specifically, when a user is speaking, they speak word by word. Therefore, in order to be able to directionally identify the corresponding professional terms from the pitch data using the keyword subtitle library, the group keyword words are split into the corresponding single Chinese words. For example, "computer motherboard" is split into "ji", "suan", "ji", "zhu", and "ban", and the pitch of each Chinese word corresponding to each group keyword word is obtained one by one as the keyword pitch. Further, by identifying the feature points of each keyword pitch, the corresponding string is generated as the pitch key value to improve the efficiency of identifying and matching from the keyword subtitle library.
[0072] S204: After associating the pitch key value with the corresponding group keyword word, and sorting according to the keyword frequency of the group keyword word, the keyword subtitle library is obtained.
[0073] Specifically, through the key-value storage method, the pitch key value is associated with the corresponding group keyword one by one, so that during the identification and matching, the matching can be carried out one by one for each pitch key value to continuously refine the result of the identification and matching. For example, if the first Chinese word spoken by the user is "ji", then the group keyword word starting with "ji" at the beginning is identified and matched in the keyword subtitle library through the key value of "ji" as the first keyword word group. Further, if the second Chinese word spoken by the user is "suan", then the group keyword word starting with the first two characters "ji suan" is further screened out from the first keyword word group, thereby identifying the unique noun.
[0074] Further, after associating the pitch key value with the group keyword word and sorting according to the keyword frequency, the keyword subtitle library is obtained.
[0075] Combined with Figure 3 and Figure 4 , after obtaining the keyword subtitle library, when performing a matching query from the preset keyword subtitle library through the pitch data to obtain the keyword data to be screened, it specifically includes:
[0076] S21: Calculate the key value to be matched for the pitch data, including:
[0077] S211: Split the pitch data to obtain the word pitch data.
[0078] S212: Calculate the key value to be matched one by one in the order of the pitch data.
[0079] In this embodiment, the key value to be matched refers to the string used for matching in the keyword subtitle library.
[0080] Specifically, after obtaining the tone data, the tone data is split into each Chinese word to obtain word tone data. The tone data can be split based on pitch or pauses, etc. Therefore, the split Chinese word can be the tone data corresponding to a single Chinese character or a phrase consisting of multiple Chinese characters, as the word tone data.
[0081] Furthermore, after obtaining the word tone data, the feature points of the word tone data are obtained one by one according to the order of the tone data, that is, the order in which the user speaks, and the key values corresponding to the feature points of the word tone data are calculated one by one through a hash algorithm or other corresponding algorithm as the key values to be matched.
[0082] S22: Input the key value to be matched into the keyword subtitle library, and perform matching query in sequence. If the corresponding tone key value can be successfully matched, the associated group keyword will be used as the keyword data to be screened.
[0083] Specifically, the key value to be matched is input into the keyword subtitle library, and matching queries are performed one by one with the corresponding tone key value in the order in which each key value to be matched is input into the keyword subtitle library. If the matching query is successful, it means that the content spoken by the user includes keywords in the keyword subtitle library; at the same time, in order to improve the matching efficiency, the degree of fit between the feature points extracted from the tone data of the keyword and the actual tone data of the user's speech and the actual tone is reduced. Therefore, when matching the corresponding tone key value through the key value to be matched, there will be multiple tone key values matched with the same key value to be matched. Therefore, the corresponding tone key value matched by the same key value to be matched is used as the keyword data to be screened.
[0084] S30: Obtain a corresponding conference user ID from the video conference voice content, filter the keyword data to be filtered according to the conference user ID, and obtain keyword subtitles.
[0085] Specifically, when constructing the keyword subtitle database, the keywords corresponding to each conference user ID are classified into a category, and the word frequency of each keyword is counted, that is, the keywords frequently spoken by the user corresponding to the conference user ID.
[0086] Furthermore, if there are multiple keyword data to be screened associated with the same key value to be matched, the keyword with the highest frequency is screened out as the keyword subtitle according to the corresponding word frequency of the keyword to be screened.
[0087] S40: Based on the commonly used voice tone data and keyword subtitles, generate and display the conference real-time subtitles corresponding to the tone data on the user end.
[0088] Combine Figure 1 and Figure 5 , specifically including:
[0089] S41: splicing the common voice tone data and the keyword subtitles in the order of the tone data to obtain the subtitle data to be recognized;
[0090] S42: After semantic recognition is performed on the subtitle data to be recognized, real-time subtitles for the conference are obtained.
[0091] Specifically, the commonly used voice tone data and the keyword subtitles are spliced to obtain the subtitle data to be identified, and the corresponding text is identified from the commonly used voice tone data in the subtitle data to be identified to obtain a complete sentence; further, the sentence is semantically recognized to check whether the sentence is fluent, and if so, the sentence is used as the real-time subtitle of the meeting; if not, the recognized text of the commonly used voice tone data is replaced in descending order according to the word frequency of the keyword data to be screened until a fluent sentence is obtained. Optionally, in order to improve the efficiency of obtaining real-time subtitles of the meeting, the number of adjustments can be limited, for example, twice, three times, etc. If the real-time subtitles of the meeting obtained by the user still have defects, the real-time subtitles of the meeting can be marked to modify the subsequent strategy for adjusting the sentences.
[0092] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0093] In one embodiment, a computer-assisted input device is provided, which corresponds one-to-one with the computer-assisted input method in the above embodiment. Figure 6 As shown, the computer-assisted input device includes a tone extraction module, a subtitle matching module, a subtitle screening module, and a subtitle generation module. The functional modules are described in detail as follows:
[0094] A tone extraction module is used to obtain the voice content of the video conference in real time and obtain the tone data in the voice content of the video conference;
[0095] The subtitle matching module is used to perform matching queries from a preset keyword subtitle library based on the tone data to obtain the keyword data to be screened and the common voice tone data;
[0096] The subtitle screening module is used to obtain the corresponding conference user ID from the video conference voice content, and screen the keyword data to be screened according to the conference user ID to obtain the keyword subtitles;
[0097] The subtitle generation module is used to generate and display real-time conference subtitles corresponding to the tone data on the user side based on commonly used voice tone data and keyword subtitles.
[0098] Optionally, the computer-assisted input device further includes:
[0099] A group attribute acquisition module is used to obtain group domain data within a preset group of people, and obtain domain keyword information based on the group domain data, wherein the domain keyword information includes a number of keyword words;
[0100] The keyword crawling module is used to obtain historical meeting text data of a group of people, perform matching queries in the historical meeting text data based on the keyword words in the domain keyword information, and obtain group keyword words and the keyword frequency corresponding to each group keyword word;
[0101] A key value generation module is used to split each group keyword word into individual Chinese words, obtain the keyword tone corresponding to each Chinese word, and generate a tone key value based on the keyword tone;
[0102] The subtitle library construction module is used to associate the tone key value with the corresponding group keyword word, sort the group keyword words according to the keyword frequency, and obtain the keyword subtitle library.
[0103] Optionally, the subtitle filtering module includes:
[0104] Key value calculation submodule, used to calculate the key value to be matched for the tone data;
[0105] The matching and screening submodule is used to input the key value to be matched into the keyword subtitle library and perform matching queries in sequence. If the corresponding tone key value can be successfully matched, the associated group keyword will be used as the keyword data to be screened.
[0106] Optionally, the Key value calculation submodule includes:
[0107] A tone splitting unit, used for splitting the tone data to obtain word tone data;
[0108] The key value calculation unit is used to calculate the key values to be matched one by one according to the order of the tone data.
[0109] Optionally, the subtitle generation module includes:
[0110] The subtitle splicing submodule is used to splice the common voice tone data and keyword subtitles in the order of the tone data to obtain the subtitle data to be recognized;
[0111] The semantic recognition verification submodule is used to obtain real-time conference subtitles after performing semantic recognition on the subtitle data to be recognized.
[0112] The specific definition of a computer-assisted input device can be found in the definition of a computer-assisted input method described above and will not be further elaborated here. Each module in the aforementioned computer-assisted input device may be implemented in whole or in part through software, hardware, or a combination thereof. Each of the aforementioned modules may be embedded in or independent of a processor in a computer device in hardware form, or may be stored in a memory in a computer device in software form, so that the processor can call and execute the corresponding operations of each of the aforementioned modules.
[0113] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store a keyword subtitle library corresponding to each group of people. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a computer-assisted input method is implemented.
[0114] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0115] Acquire the audio content of the video conference in real time and obtain the tone data in the audio content of the video conference;
[0116] According to the tone data, a matching query is performed from the preset keyword subtitle library to obtain the keyword data to be filtered and the common voice tone data;
[0117] Obtain the corresponding conference user ID from the video conference voice content, filter the keyword data to be filtered according to the conference user ID, and obtain the keyword subtitles;
[0118] Based on commonly used voice tone data and keyword subtitles, real-time conference subtitles corresponding to the tone data are generated and displayed on the user end.
[0119] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0120] Acquire the audio content of the video conference in real time and obtain the tone data in the audio content of the video conference;
[0121] According to the tone data, a matching query is performed from the preset keyword subtitle library to obtain the keyword data to be filtered and the common voice tone data;
[0122] Obtain the corresponding conference user ID from the video conference voice content, filter the keyword data to be filtered according to the conference user ID, and obtain the keyword subtitles;
[0123] Based on commonly used voice tone data and keyword subtitles, real-time conference subtitles corresponding to the tone data are generated and displayed on the user end.
[0124] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0125] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0126] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A computer-assisted input method, characterized in that: The computer-assisted input method is applied to a device connected to a computer and transmitting data, including: Acquire the video conference voice content in real time, and acquire the tone data in the video conference voice content; Performing a matching query from a preset keyword and subtitle library based on the tone data to obtain keyword data to be screened and common voice tone data, wherein performing a matching query from a preset keyword and subtitle library based on the tone data, constructing the keyword and subtitle library includes: Acquire group domain data within a preset group of people, and acquire domain keyword information based on the group domain data, wherein the domain keyword information includes a plurality of keyword words; Obtain historical meeting text data of the group of people, perform a matching query in the historical meeting text data based on the keyword words in the field keyword information, and obtain group keyword words and keyword frequency corresponding to each group keyword word; After splitting each of the group keyword words into individual Chinese words, obtaining the keyword tone corresponding to each of the Chinese words, and generating a tone key value according to the keyword tone; After associating the tone key value with the corresponding group keyword word, the keyword subtitle library is obtained after sorting the group keyword words according to the keyword word frequency; Obtaining a corresponding conference user identifier from the video conference voice content, and filtering the keyword data to be filtered according to the conference user identifier to obtain keyword subtitles; Based on the common voice tone data and the keyword subtitles, real-time conference subtitles corresponding to the tone data are generated and displayed on the user terminal.
2. The computer-assisted input method according to claim 1, wherein: The method of performing a matching query from a preset keyword subtitle library based on the tone data to obtain keyword data to be screened specifically includes: Calculating a key value to be matched for the tone data; The key value to be matched is input into the keyword subtitle library, and a matching query is performed in sequence. If the corresponding tone key value can be successfully matched, the associated group keyword is used as the keyword data to be screened.
3. The computer-assisted input method according to claim 2, wherein: The calculating the key value to be matched of the tone data specifically includes: Splitting the tone data to obtain word tone data; The key values to be matched are calculated one by one according to the order of the tone data.
4. The computer-assisted input method according to any one of claims 1 to 3, wherein: The generating and displaying, on the user side, conference real-time subtitles corresponding to the tone data based on the common voice tone data and the keyword subtitles specifically includes: splicing the common voice tone data and the keyword subtitles according to the order of the tone data to obtain subtitle data to be recognized; After semantic recognition is performed on the subtitle data to be recognized, the real-time subtitles of the conference are obtained.
5. A computer-assisted input device, characterized in that: The computer-assisted input device comprises: A tone extraction module is used to obtain the video conference voice content in real time and obtain the tone data in the video conference voice content; A subtitle matching module is used to perform a matching query from a preset keyword subtitle library based on the tone data to obtain keyword data to be screened and common voice tone data; A group attribute acquisition module is used to acquire group domain data within a preset group of people, and acquire domain keyword information based on the group domain data, wherein the domain keyword information includes a number of keyword words; A keyword crawling module is used to obtain historical meeting text data of the personnel group, perform matching queries in the historical meeting text data according to the keyword words in the field keyword information, and obtain group keyword words and keyword frequency corresponding to each group keyword word; A key value generating module is used to split each of the group keyword words into individual Chinese words, obtain the keyword tone corresponding to each of the Chinese words, and generate a tone key value according to the keyword tone; A subtitle library construction module is used to associate the tone key value with the corresponding group keyword word, and then sort the group keyword words according to the keyword word frequency to obtain the keyword subtitle library; A subtitle screening module is used to obtain a corresponding conference user identifier from the video conference voice content, and screen the keyword data to be screened according to the conference user identifier to obtain keyword subtitles; The subtitle generation module is used to generate and display the conference real-time subtitles corresponding to the tone data on the user end based on the common voice tone data and the keyword subtitles.
6. The computer-assisted input device according to claim 5, wherein: The subtitle screening module includes: A key value calculation submodule, configured to calculate a key value to be matched for the tone data; The matching and screening submodule is used to input the key value to be matched into the keyword subtitle library, perform matching query in sequence, and if the corresponding tone key value can be successfully matched, the associated group keyword will be used as the keyword data to be screened.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the computer-assisted input method according to any one of claims 1 to 4 are implemented.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the computer-assisted input method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Video data processing method and device, computer equipment and storage medium
CN109714608A