A speech recognition system based on semantic understanding of computer application scenarios

By designing a speech recognition system based on computer application scenarios, using voice decomposition, filtering, conversion, extraction and display modules to extract and display key information in the voice, the problem that the existing system cannot effectively extract voice focus and improve user experience and information acquisition efficiency.

CN115273821BActive Publication Date: 2025-05-23PINGDINGSHAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210818227.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2025-05-23
Estimated Expiration
2042-07-13

AI Technical Summary

Technical Problem

The existing voice recognition system cannot effectively extract key information in the voice when receiving voice, resulting in users wasting time when obtaining information and poor user experience.

Method used

Design a speech recognition system based on semantic understanding of computer application scenarios. Through speech decomposition, filtering, conversion, extraction and display modules, speech is converted into text, and important fields are extracted through text meaning extraction technology and displayed to users.

Benefits of technology

It realizes the rapid extraction and display of voice information, saves users time to obtain information, improves user experience, and highlights important content through tone judgment and color display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115273821B_ABST
    Figure CN115273821B_ABST
Patent Text Reader

Abstract

The present invention discloses a speech recognition system based on semantic understanding of computer application scenarios, including a speech decomposition module for receiving a user's speech and decomposing the user's speech into a plurality of sub-speech with equal time; a speech filtering module respectively analyzes each of the sub-speech, deletes the sub-speech with blank sound, and splices the remaining sub-speech to obtain a new speech; a speech conversion module converts the new speech into a text segment using a speech-to-text conversion technology, and stores the text segment in a buffer; a text extraction module extracts important fields from the text segment in the buffer using a text extraction technology; and a text display module displays the important fields on a user interface. The present invention extracts the converted text to obtain key information in the text, and displays the extracted key points, thereby saving the user's time in obtaining information and improving the user's experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer speech, and in particular to a speech recognition system based on semantic understanding of computer application scenarios. Background Art

[0002] In the current instant messaging software, most of them support voice sending, that is, recording the user's voice and sending it to the other party. Since this function meets the user's convenience needs when used, it has been widely used nowadays. However, when it is used, it often has certain disadvantages. For example, when there is no way to receive voice, receiving voice is really overwhelming. For example, when receiving a long voice, it often takes a lot of time to extract the semantic focus, which makes people bored. Therefore, the disadvantages of receiving voice are some problems that need to be solved.

[0003] The existing solution to the above problem is to use speech-to-text technology to convert speech into text and display it to the user. However, this function requires the user to trigger it manually. At the same time, when the speech is converted into text, the key points are not extracted and displayed, but all the information is displayed. It is obviously a waste of time for people to watch such long passages, and there is no good user experience. Summary of the invention

[0004] The purpose of the present invention is to overcome the problems existing in the above-mentioned prior art and provide a speech recognition system based on semantic understanding of computer application scenarios, which can extract the converted text, obtain the key information in the text, and display the extracted key points, thereby saving the user's time in obtaining information and improving the user experience.

[0005] To this end, the present invention provides a speech recognition system based on semantic understanding of computer application scenarios, comprising:

[0006] A speech decomposition module is used to receive the user's speech and decompose the user's speech into a plurality of sub-speech with equal time, and mark each sub-speech with a timestamp, wherein the timestamp is the starting time of the corresponding sub-speech in the speech;

[0007] A speech filtering module analyzes each of the sub-speech respectively, deletes the sub-speech with blank sound, sorts the remaining sub-speech according to the order of the timestamps on each of the sub-speech, and splices the sorted sub-speech to obtain a new speech;

[0008] A speech conversion module converts the new speech into text segments using speech-to-text conversion technology, and stores the text segments in a buffer area;

[0009] A text extraction module extracts important fields from the text segments in the buffer using a text extraction technique;

[0010] The text display module displays the important fields on the user interface.

[0011] Furthermore, the text extraction technology comprises the following steps:

[0012] Decomposing the text segment into a plurality of sentences, and numbering and marking each sentence according to its position in the text segment;

[0013] Extracting keywords from each of the sentences in sequence, and arranging the keywords from each sentence to form a keyword sequence;

[0014] When the first keyword and the last keyword in the keyword sequence are the same, output the sentence corresponding to the last keyword;

[0015] When at least three identical keywords appear in the keyword sequence, obtain the keyword with the highest frequency of occurrence, and compare the obtained keyword with the first keyword and the last keyword respectively. When the obtained keyword is the same as the first keyword or the last keyword, output the sentence corresponding to the first keyword or the last keyword, otherwise output all the sentences.

[0016] Furthermore, when the acquired keyword is the same as the first keyword or the last keyword, the corresponding sentence is output, including the following steps:

[0017] The obtained keyword is compared with the first keyword and the last keyword respectively, and the results are true when they are the same, and false when they are different;

[0018] When the results of the comparison between the first keyword and the last keyword are both true, output the sentence corresponding to the last keyword;

[0019] When the comparison results of the first keyword and the last keyword are different, the sentence corresponding to the keyword with the true comparison result.

[0020] Furthermore, when the text display module is displayed on the user interface, the output sentence is displayed below the voice, and a full-text button is set at the same time. When the user clicks the full-text button, the text segment is displayed.

[0021] Furthermore, the speech conversion module removes repeated sentences in the text segment before storing the text segment in the buffer area.

[0022] Furthermore, when removing repeated sentences in a text segment, the following steps are included:

[0023] Decomposing the text segment into a plurality of sentences, and numbering and marking each sentence according to its position in the text segment;

[0024] Compare the similarities of the two statements respectively to obtain similarity values ​​of the two statements, and when the similarity value is greater than a set value, delete one of the statements and traverse all the statements;

[0025] Arrange the remaining sentences in sequence according to their numbers to obtain a concise paragraph, and output the concise paragraph.

[0026] Furthermore, it also includes:

[0027] A tone judgment module, used for detecting the speech speed and pitch of the speech, obtaining the tone of the speech according to the weighted speech speed and pitch, and judging the tone degree of the speech;

[0028] The display prediction module is used to control the text display module to display the important fields in corresponding colors according to the tone of the voice.

[0029] The present invention provides a speech recognition system based on semantic understanding of computer application scenarios, which has the following beneficial effects:

[0030] The present invention refines the information displayed after voice conversion into text through text processing technology, filters out unimportant information content, and the remaining displayed information is the key information of the converted text, so that when users read information, they can save unnecessary time, quickly extract the key content of the information, and quickly review the information, thereby improving the user experience;

[0031] When extracting the key content of the information, the present invention first removes the modal particles in the voice information according to the time interval of each word, and then obtains the speech with practical meaning in the voice information through semantic technology, and finally converts the speech into words for preservation;

[0032] The present invention also obtains the importance of the speech by detecting the tone of voice in the speech, and highlights the content with high importance by using color differentiation. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic block diagram of the system of the present invention;

[0034] Figure 2 It is a schematic flow chart of the text extraction technology of the present invention;

[0035] Figure 3 It is a schematic block diagram of the process of outputting corresponding sentences when the keywords are the same in the present invention;

[0036] Figure 4 The present invention is a schematic block diagram of the process of removing repeated sentences in a text segment. DETAILED DESCRIPTION

[0037] A specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings, but it should be understood that the protection scope of the present invention is not limited by the specific implementation.

[0038] In the present application document, component models and structures that are not clearly defined are all prior arts known to those skilled in the art, and those skilled in the art may make settings according to actual needs, and are not specifically limited in the embodiments of the present application document.

[0039] Specifically, Figure 1-4 As shown, the embodiment of the present invention provides a speech recognition system based on semantic understanding of computer application scenarios, including: a speech decomposition module, a speech filtering module, a speech conversion module, a text extraction module and a text display module. The following is a detailed description of the working of each module.

[0040] A speech decomposition module is used to receive the user's speech and decompose the user's speech into several sub-speech of equal time. Each sub-speech is marked with a timestamp, and the timestamp is the starting time of the corresponding sub-speech in the speech; the user speech of the present invention refers to the received speech used for decomposition, which is decomposed into several sub-speech of equal time, each with a timestamp. First, the playback time of the user's speech is obtained, for example, a 40-second speech is divided into 40 sub-speech at intervals of 1 second, and each sub-speech has a corresponding timestamp. The timestamp of the first sub-speech is 0 seconds, the timestamp of the second sub-speech is 1 second, and so on.

[0041] The speech filtering module analyzes each of the sub-speech respectively, deletes the sub-speech with blank sound, sorts the remaining sub-speech according to the order of timestamps on each sub-speech, and splices the sorted sub-speech to obtain a new speech; the blank speech in the present invention is a speech with only noise and no useful sound after the sub-speech is played. After filtering the sub-speech without sound, the obtained sub-speech is arranged in the order of timestamps to obtain a new speech. This new speech is the subsequent research object and has saved some unnecessary time for users.

[0042] The speech conversion module converts the new speech into text segments using speech-to-text conversion technology, and stores the text segments in a buffer area; this module is implemented using currently mature speech-to-text conversion technology.

[0043] The text extraction module extracts the important fields from the text segments in the buffer using the text extraction technology; this module further refines the information in the above text segments to obtain the key information therein, namely the important fields.

[0044] The text display module displays the important fields on the user interface for users to view.

[0045] In the above technical solution, the speech decomposition module, the speech filtering module, the speech conversion module, the text extraction module and the text display module interact with each other, so that after the speech received by the user is initially filtered, the text segments are extracted, and the extracted text segments are subjected to a second key extraction, so that when the user uses it, he or she can not only obtain the key content quickly and accurately without listening to the speech, which easily saves the time of obtaining information for the user, and at the same time helps the user improve the accuracy of obtaining information.

[0046] In the present invention, the text extraction technology comprises the following steps:

[0047] (1) decomposing the text segment into a plurality of sentences, and numbering and marking each sentence according to its position in the text segment;

[0048] (2) extracting the keywords of each of the sentences in sequence, and arranging the keywords of each sentence to form a keyword sequence;

[0049] (3) when the first keyword and the last keyword in the keyword sequence are the same, outputting the sentence corresponding to the last keyword;

[0050] (iv) when at least three identical keywords appear in the keyword sequence, obtain the keyword with the highest frequency of occurrence, and compare the obtained keyword with the first keyword and the last keyword respectively; when the obtained keyword is the same as the first keyword or the last keyword, output the sentence corresponding to the first keyword or the last keyword; otherwise, output all the sentences.

[0051] In the above technical solution, steps (i) to (iv) are performed in a logical order. The present invention obtains the language structure of the text segment by means of keywords, which is in the form of general-specific-general, general-specific, specific-general or specific-general-specific. By finding the context of the text segment, the key information of the text segment is obtained. The meaning of each sentence in the text segment is determined by means of keywords. The final output in step (iv) is the central sentence of the text segment, which is also the sentence that can express the meaning of the entire text segment. The user can directly read the sentence to grasp the entire text segment.

[0052] Meanwhile, in the present invention, when the acquired keyword is the same as the first keyword or the last keyword, the corresponding sentence is output, including the following steps:

[0053] (1) comparing the obtained keyword with the first keyword and the last keyword respectively, where the results are true when they are the same and false when they are different;

[0054] (2) When the results of the comparison between the first keyword and the last keyword are both true, output the sentence corresponding to the last keyword;

[0055] (3) When the comparison results of the first keyword and the last keyword are different, the sentence corresponding to the keyword with the true comparison result.

[0056] In the above technical solution, steps (1) to (3) are performed in sequence according to a logical order. The steps (1) to (3) are suitable for the writing context of a general-specific, specific-general or specific-general-specific structure. In step (1), the writing context of a general-specific, specific-general or specific-general-specific structure is determined. Step (2) indicates that the writing context is a specific-general structure, and the output is the part of the sentence that represents the general. Step (3) indicates that the writing context is a general-specific or specific-general-specific structure, and the output is also the part of the sentence that represents the general.

[0057] When extracting the key points of a text segment, the present invention combines the key points with the context of the text, so that the core meaning of the extracted text segment is clearer and more specific.

[0058] At the same time, in the present invention, when the text display module is displayed on the user interface, the output sentence is displayed below the voice, and a full-text button is set. When the user clicks the full-text button, the text segment is displayed.

[0059] In the present invention, the speech conversion module removes repeated sentences in the text segment before storing the text segment in the buffer area, so that when extracting key points, the processing of sufficient sentences can be reduced, thereby improving the calculation speed of the system.

[0060] Meanwhile, in the present invention, when removing repeated sentences in a text segment, the following steps are included:

[0061] <1> Decomposing the text segment into a plurality of sentences, and numbering and marking each sentence according to its position in the text segment;

[0062] <2> Compare the similarities of the two statements respectively to obtain similarity values ​​of the two statements, and when the similarity value is greater than a set value, delete one of the statements and traverse all the statements;

[0063] <3> Arrange the remaining sentences in sequence according to their numbers to obtain a concise paragraph, and output the concise paragraph.

[0064] In the above technical solution, the steps <1> to <3> The present invention compares the sentences in the text segment in a logical order, finds the repetitions between the sentences, deletes the repetitive sentences, and obtains concise and clear sentences, which are the concise text segments described in the present invention.

[0065] The present invention also includes: a tone judgment module and a display prediction module. The tone judgment module is used to detect the speech speed and pitch of the speech, and obtain the tone of the speech according to the weighted speech speed and pitch, and judge the tone of the speech; the display prediction module is used to control the text display module to display the important fields in a corresponding color according to the tone of the speech.

[0066] The present invention also obtains the importance of the speech by detecting the tone in the speech, and highlights the content with high importance by using color differentiation. The present invention uses percentage to express the tone, and the larger the percentage value, the heavier the tone.

[0067] The above disclosures are only several specific embodiments of the present invention. However, the embodiments of the present invention are not limited thereto. Any changes that can be conceived by those skilled in the art should fall within the protection scope of the present invention.

Claims

1. A speech recognition system based on semantic understanding of computer application scenarios, It is characterized in that include: A speech decomposition module is used to receive the user's speech and decompose the user's speech into a plurality of sub-speech with equal time, and mark each sub-speech with a timestamp, wherein the timestamp is the starting time of the corresponding sub-speech in the speech; A speech filtering module analyzes each of the sub-speech respectively, deletes the sub-speech with blank sound, sorts the remaining sub-speech according to the order of the timestamps on each of the sub-speech, and splices the sorted sub-speech to obtain a new speech; A speech conversion module converts the new speech into text segments using speech-to-text conversion technology, and stores the text segments in a buffer area; A text extraction module extracts important fields from the text segments in the buffer using a text extraction technique; A text display module displays the important fields on the user interface; The text extraction technology comprises the following steps: Decomposing the text segment into a plurality of sentences, and numbering and marking each sentence according to its position in the text segment; Extracting the keywords of each of the sentences in turn, and arranging the keywords of each sentence to form a keyword sequence; When the first keyword and the last keyword in the keyword sequence are the same, output the sentence corresponding to the last keyword; When at least three identical keywords appear in the keyword sequence, obtain the keyword with the highest frequency of occurrence, and compare the obtained keyword with the first keyword and the last keyword respectively. When the obtained keyword is the same as the first keyword or the last keyword, output the sentence corresponding to the first keyword or the last keyword, otherwise output all the sentences.

2. A speech recognition system based on computer application scenario semantic understanding as claimed in claim 1, It is characterized in that When the acquired keyword is the same as the first keyword or the last keyword, the corresponding sentence is output, including the following steps: The obtained keyword is compared with the first keyword and the last keyword respectively, and the results are true when they are the same, and false when they are different; When the results of the comparison between the first keyword and the last keyword are both true, output the sentence corresponding to the last keyword; When the comparison results of the first keyword and the last keyword are different, the sentence corresponding to the keyword with the true comparison result.

3. A speech recognition system based on computer application scenario semantic understanding as claimed in claim 2, It is characterized in that When the text display module is displayed on the user interface, the output sentence is displayed below the voice, and a full-text button is provided. When the user clicks the full-text button, the text segment is displayed.

4. A speech recognition system based on computer application scenario semantic understanding as claimed in claim 1, It is characterized in that The speech conversion module removes repeated sentences in the text segment before storing the text segment in a buffer area.

5. A speech recognition system based on computer application scenario semantic understanding as claimed in claim 4, It is characterized in that When removing repeated statements in a text segment, the following steps are included: Decomposing the text segment into a plurality of sentences, and numbering and marking each sentence according to its position in the text segment; Compare the similarities of the two statements respectively to obtain similarity values ​​of the two statements, and when the similarity value is greater than a set value, delete one of the statements and traverse all the statements; Arrange the remaining sentences in sequence according to their numbers to obtain a concise paragraph, and output the concise paragraph.

6. A speech recognition system based on computer application scenario semantic understanding as claimed in claim 1, It is characterized in that Also includes: A tone judgment module, used for detecting the speech speed and pitch of the speech, obtaining the tone of the speech according to the weighted speech speed and pitch, and judging the tone degree of the speech; The display prediction module is used to control the text display module to display the important fields in corresponding colors according to the tone of the voice.

Citation Information

Patent Citations

  • Speech recognition method and related product

    CN111161739A

  • Conference summary recording method and device and electronic equipment

    CN112687272A