Method, System and Electronic Device for Identifying and Displaying Key Subsequences of Speech Sequences

By identifying pause points in the phonological sequence and dividing them into subsequences, using multiple processes to identify and display pictures of chemical substances and reaction relationships, the problem of key subsequence recognition in phonological communication in the field of chemistry is solved, and communication efficiency is improved.

CN114783422BActive Publication Date: 2025-07-08IOL WUHAN INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202210403490.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-21
Publication Date
2025-07-08
Estimated Expiration
2042-04-21

AI Technical Summary

Technical Problem

In the vocabulary communication in the field of chemistry and molecular fields, the prior art cannot effectively highlight the key subsequences in the vocabulary sequence, resulting in the inability of both parties to the voice interaction to quickly and accurately absorb professional content, affecting communication efficiency.

Method used

By identifying pauses in the phonological sequence, splitting them into multiple subsequences, and using multiple processes to identify key subsequences, displaying pictures of the relationship between chemical substances and chemical reactions in real time, such as chemical formulas and reaction formulas.

Benefits of technology

It realizes the rapid and accurate identification and display of key subsequences in phonological sequences in the field of chemistry and other occasions, and improves the efficiency of speech communication, especially professional communication in the field of chemistry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114783422B_ABST
    Figure CN114783422B_ABST
Patent Text Reader

Abstract

The present invention provides a method, system and electronic device for identifying and displaying key subsequences of a speech sequence, belonging to the technical field of speech recognition. The method includes step S100 of obtaining a speech sequence; S200 of identifying whether there is a key subsequence in the speech sequence; if there is, the key subsequence is displayed in a predetermined format while the speech sequence is being played; if not, the speech sequence is directly played; the key subsequence includes the reaction relationship of at least one chemical substance and / or a combination of chemical substances. The system includes a speech receiving end, a speech recognition end, a speech display end and a speech broadcast end; the speech broadcast end is used to broadcast the key subsequence while the key subsequence is displayed in a predetermined format on the display interface by the speech display end. The present invention also provides an electronic device for implementing the above method. The present invention can realize the identification and display of key subsequences of a speech sequence, which is helpful for focusing on key points during online speech teaching in the chemical field and saving system resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of speech recognition, and in particular relates to a method, system and electronic device for identifying and displaying a key subsequence of a speech sequence, and a storage medium for implementing the method. Background Art

[0002] With the development of artificial intelligence technology and the popularization of smart terminals, the needs of recording and converting text have been greatly met by providing functions such as conference voice record recognition, conference recording conversion to text, and key conference archive indexing through terminal modules such as mobile phones. Nowadays, voice navigation, voice wake-up, voice dialing, voice-to-text and other functions have become popular in various terminals. Intelligent voice control has evolved from a teasing application when users are bored to a functional application that can truly help users solve practical problems. Intelligent voice applications are maturing, and the terminal industry is ushering in a new revolution featuring intelligent voice control.

[0003] For example, the Chinese invention patent CN202111600180.3 under review proposes a voice information processing method that can realize voice-to-text conversion without sending voice messages to customers or protecting the personal privacy of voice messages, thereby ensuring the security of information interaction in real-time chat programs.

[0004] However, the inventors found that although most of the real-time speech translation engines in the prior art can realize real-time recognition of continuous audio streams, real-time recognition and translation of speech input content, convert it into text information and return the corresponding text stream, in some special occasions, especially occasions containing specific vocabulary and specific format vocabulary, such as speech communication in the chemical field, molecular field and other occasions, the speech sequence usually contains a large number of professional chemical vocabulary. The prior art's word-by-word and sentence-by-sentence original speech-to-text translation method cannot highlight the key subsequences in these speech sequences, making it impossible for both parties in the speech interaction to quickly and accurately absorb the key content, affecting the efficiency of speech communication. Summary of the invention

[0005] In order to solve the above technical problems, the present invention proposes a method, system and electronic device for identifying and displaying a key subsequence of a speech sequence, and a storage medium for implementing the method.

[0006] In a first aspect of the present invention, a method for identifying and displaying a key subsequence of a speech sequence is provided, the method comprising the following steps:

[0007] S100: Acquire a speech sequence;

[0008] S200: Identify whether there is a key subsequence in the speech sequence;

[0009] If it exists, while playing the voice sequence, the key subsequence is displayed in a predetermined format;

[0010] If it does not exist, the voice sequence is directly played;

[0011] The key subsequence includes the reaction relationship of at least one chemical substance and / or chemical substance combination;

[0012] The displaying of the key subsequence in a predetermined format includes: displaying at least one chemical formula corresponding to the at least one chemical substance and / or a chemical reaction formula, a chemical molecular formula, an electronic structural formula of a chemical substance, etc. corresponding to the chemical substance combination in the form of a picture.

[0013] As the implementation basis of the solution of the present invention, the voice sequence obtained in step S100 includes multiple pause points;

[0014] After step S100 and before step S200, the method further includes the following steps:

[0015] Taking the pause points as units, the voice sequence is segmented into multiple voice subsequences.

[0016] Meanwhile, to ensure real-time performance and accuracy, while saving system process resources and ensuring the adaptability of the process and the processing process, step S100 further includes:

[0017] After obtaining the voice sequence, determining the first number of pause points included in the voice sequence;

[0018] Activating the recognition processes of the second number of key subsequences;

[0019] As a specific implementation means, step S200 includes the following sub-steps:

[0020] S201: Identifying the current pause point in the obtained voice sequence;

[0021] S202: Translating the voice subsequence before the current pause point into a text subsequence;

[0022] S203: Judging whether the text subsequence contains a predetermined key subsequence;

[0023] If it contains, while playing the voice subsequence, the key subsequence is displayed in a predetermined format;

[0024] If it does not contain, directly play the text subsequence, obtain the next pause point in the voice sequence, take the next pause point as the current pause point, and return to step S202.

[0025] Step S200 further includes:

[0026] Identify whether there is a key subsequence in the speech sequence through the identification process of the second quantity of key subsequences.

[0027] As a more specific method solution, the method for identifying and displaying key subsequences of a speech sequence provided by the second aspect of the present invention includes the following steps:

[0028] S510: Obtain the current speech sequence;

[0029] S520: Identify the first quantity of pause points included in the current speech sequence, and activate the identification process of the second quantity of key subsequences;

[0030] S530: Segment the current speech sequence into a third quantity of speech subsequences with the first quantity of pause points as units;

[0031] S540: Use the identification process of the second quantity of key subsequences to identify in parallel whether each of the speech subsequences contains a predetermined key subsequence;

[0032] S550: If two adjacent speech subsequences contain the same predetermined key subsequence, then merge the two adjacent speech subsequences into one speech subsequence;

[0033] S560: Play the speech subsequences in sequence, and if the currently played speech subsequence contains a predetermined key subsequence, then display the predetermined key subsequence in a predetermined format at the same time;

[0034] The key subsequence includes the reaction relationship of at least one chemical substance and / or chemical substance combination.

[0035] Compared with the method described in the first aspect, the further improvement of the method described in the second aspect lies in steps S540 - S560, especially step S550 "If two adjacent speech subsequences contain the same predetermined key subsequence, then merge the two adjacent speech subsequences into one speech subsequence". In actual operation, due to this improvement, the processing efficiency of key subsequences can be further improved, and at the same time, the number of processes can be saved.

[0036] Specifically, the second quantity is greater than the third quantity;

[0037] The displaying the key subsequence in a predetermined format includes: displaying at least one chemical formula corresponding to the at least one chemical substance and / or a chemical reaction formula corresponding to a chemical substance combination as a picture.

[0038] To implement the method described in the first aspect or the second aspect, in the third aspect of the present invention, a system for identifying and displaying key subsequences of a speech sequence is provided. The system includes a speech receiving end, a speech recognition end, a speech display end, and a speech broadcast end;

[0039] Among them, the specific functions of each sub-end module are implemented as follows:

[0040] The speech receiving end is used to receive the speech sequence;

[0041] The speech recognition end is used to identify the pause points included in the speech sequence and whether the speech sequence includes key subsequences;

[0042] The speech display end is used to display the key subsequences on the display interface in a predetermined format;

[0043] Among them, the speech recognition end identifies whether the speech sequence includes key subsequences with the pause points included in the speech sequence as nodes;

[0044] The key subsequences include the reaction relationships of at least one chemical substance and / or chemical substance combination;

[0045] The displaying the key subsequences on the display interface in a predetermined format includes: displaying the at least one chemical formula corresponding to the at least one chemical substance and / or the chemical reaction formula corresponding to the chemical substance combination as pictures.

[0046] The speech broadcast end is used to broadcast the key subsequences while the speech display end displays the key subsequences on the display interface in a predetermined format.

[0047] More specifically, the speech recognition end includes a pause point recognition unit, a subsequence segmentation unit, a subsequence recognition unit, and a subsequence merging unit;

[0048] The pause point recognition unit is used to identify the pause points included in the speech sequence;

[0049] The subsequence segmentation unit is used to segment the speech sequence into multiple speech subsequences with the pause points as units;

[0050] The subsequence recognition unit is used to identify whether each speech subsequence includes a predetermined key subsequence;

[0051] The subsequence merging unit is used to merge two adjacent speech subsequences into one speech subsequence if the two adjacent speech subsequences include the same predetermined key subsequence.

[0052] In a fourth aspect of the present invention, there is provided a terminal device, which may be, for example, a data interaction device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program may be a data interaction program. When the processor executes the computer program, the steps of the method described in the first aspect or the second aspect are implemented.

[0053] In a fifth aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, which when executed by a processor, implements the steps of the method described in the first aspect or the second aspect.

[0054] In a sixth aspect of the present invention, there is provided an electronic device, comprising: a processor for executing a method for identifying and displaying key subsequences of a voice sequence described in the first aspect or the second aspect; and a memory coupled to the processor for storing instructions executed by the processor.

[0055] The technical solution of the present invention aims at the problem that in some special occasions, especially in the occasions of voice communication in fields such as the chemical field and the molecular field, which contain specific vocabulary and vocabulary in specific formats, the voice sequence usually contains a large number of professional chemical terms. It can highlight the key subsequences in these voice sequences, enabling both parties of the voice interaction to quickly and accurately absorb the key content and improving the efficiency of voice communication.

[0056] In subsequent further embodiments, it can also be seen that the present invention can also real-time recognize the audio stream as text, which is applicable to long sentence voice input, audio-visual subtitles, meetings, speech subtitles on the same screen, intelligent language processing. It can also perform intelligent error correction on the intermediate results of recognition, display the intermediate text results in real time, quickly recognize the key audio stream, and quickly display the recognition results in the form of a picture text summary. For example, at least one chemical formula corresponding to a chemical substance and / or a chemical reaction formula, a chemical molecular formula, an electronic structural formula of a chemical substance, etc. are displayed in a picture.

[0057] The further advantages of the present invention will be further detailed in the specific embodiment part in conjunction with the accompanying drawings of the specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0059] Figure 1It is the main flowchart of a method for identifying and displaying key subsequences of a speech sequence according to an embodiment of the present invention;

[0060] Figure 2 It is to implement Figure 1 The flowchart of computer program instructions for the multi-modal speech translation method described above;

[0061] Figure 3 It is the main flowchart of a method for identifying and displaying key subsequences of a speech sequence according to another preferred embodiment of the present invention;

[0062] Figure 4 It is to implement Figure 1 Or Figure 3 The schematic diagram of the unit module of the key subsequence recognition and display system of the speech sequence for the method for identifying and displaying key subsequences of a speech sequence described above;

[0063] Figure 5 It is Figure 4 The schematic diagram of a further preferred embodiment of the key subsequence recognition and display system of the speech sequence described above;

[0064] Figure 6 It is Figure 4 The internal structure schematic diagram of the speech recognition end in the key subsequence recognition and display system of the speech sequence described above;

[0065] Figure 7 It is the schematic diagram of the display effect after the implementation of the technical solution of the present invention;

[0066] Figure 8 It is to implement Figure 1 Or Figure 2 Or Figure 3 The structure schematic diagram of the terminal device for the method described above. Detailed implementation manners

[0067] Next, in combination with the accompanying drawings and specific implementation manners, the invention will be further described.

[0068] Figure 1 It is the main flowchart of a method for identifying and displaying key subsequences of a speech sequence according to an embodiment of the present invention.

[0069] In Figure 1 it, the method can be generally summarized into two main steps S100 and S200, and each step is specifically summarized as follows:

[0070] S100: Obtain a speech sequence;

[0071] S200: Identify whether there is a key subsequence in the speech sequence;

[0072] If it exists, the key subsequence is displayed in a predetermined format while playing the voice sequence;

[0073] If it does not exist, the voice sequence is played directly;

[0074] In particular, it should be emphasized that the present invention is directed to occasions including specific vocabulary and specific format vocabulary, such as in the fields of chemistry, molecules, etc. In voice communication in such occasions, the voice sequence usually contains a large number of professional chemical vocabulary, which is the improvement motivation and main concept of the present invention.

[0075] Therefore, the key subsequence includes at least one reaction relationship of chemical substances and / or combinations of chemical substances;

[0076] The displaying of the key subsequence in a predetermined format includes: displaying at least one chemical formula corresponding to the at least one chemical substance and / or a chemical reaction formula corresponding to a combination of chemical substances in the form of a picture.

[0077] Specifically, at least one chemical formula corresponding to a chemical substance and / or a chemical reaction formula, chemical molecular formula, electronic structural formula of a chemical substance, etc. corresponding to a combination of chemical substances are displayed in the form of a picture.

[0078] Figure 2 is to implement Figure 1 A schematic flowchart of computer program instructions for the multimodal voice translation method.

[0079] Figure 2 In it, the step S200 includes the following sub-steps:

[0080] S201: Identify the current pause point in the acquired voice sequence;

[0081] S202: Translate the voice subsequence before the current pause point into a text subsequence;

[0082] S203: Determine whether the text subsequence contains a predetermined key subsequence;

[0083] If it contains, the key subsequence is displayed in a predetermined format while playing the voice subsequence, and the next pause point in the voice sequence is acquired, and the next pause point is used as the current pause point, and return to step S202;

[0084] If it does not contain, the text subsequence is played directly, and the next pause point in the voice sequence is acquired, and the next pause point is used as the current pause point, and return to step S202.

[0085] More specifically, the step S100 further includes:

[0086] After obtaining the speech sequence, determine the first quantity of pause points included in the speech sequence;

[0087] Activate the recognition processes of the second quantity of key subsequences;

[0088] The speech sequence obtained in step S100 includes multiple pause points;

[0089] After step S100 and before step S200, the method further includes the following steps:

[0090] Taking the pause points as units, segment the speech sequence into multiple speech subsequences.

[0091] Step S200 further includes:

[0092] Identify whether there are key subsequences in the speech sequence through the recognition processes of the second quantity of key subsequences.

[0093] On the basis of Figure 1 - Figure 2 See further Figure 3 . Figure 3 It is the main flowchart of a method for identifying and displaying key subsequences of a speech sequence according to another preferred embodiment of the present invention.

[0094] Figure 3 On the basis of Figure 1 - Figure 2 Further improvements are made, and the method flow includes steps S510 - S560. The specific implementation of each step is as follows:

[0095] S510: Obtain the current speech sequence;

[0096] S520: Identify the first quantity of pause points included in the current speech sequence, and activate the recognition processes of the second quantity of key subsequences;

[0097] S530: Segment the current speech sequence into the third quantity of speech subsequences with the first quantity of pause points as units;

[0098] S540: Use the recognition processes of the second quantity of key subsequences to parallelly identify whether each speech subsequence contains a predetermined key subsequence;

[0099] S550: If two adjacent speech subsequences contain the same predetermined key subsequence, then merge the two adjacent speech subsequences into one speech subsequence;

[0100] S560: Sequentially play the speech subsequences, and if the currently played speech subsequence contains a predetermined key subsequence, then simultaneously display the predetermined key subsequence in a predetermined format;

[0101] The key subsequence includes the reaction relationships of at least one chemical substance and / or chemical substance combination.

[0102] The second quantity is greater than the third quantity;

[0103] The displaying of the key subsequence in a predetermined format includes: displaying at least one chemical formula corresponding to the at least one chemical substance and / or a chemical reaction formula corresponding to the chemical substance combination as a picture.

[0104] Figure 3 A further improvement of the method lies in steps S540 - S560, especially step S550 "If two adjacent voice subsequences contain the same predetermined key subsequence, then merge the two adjacent voice subsequences into one voice subsequence". In actual operation, due to this improvement, the processing efficiency of the key subsequence can be further improved, and at the same time, the number of processes can be saved.

[0105] Figure 4 is to implement Figure 1 or Figure 3 The schematic diagram of the unit module of the key subsequence recognition and display system for the voice sequence of the method for recognizing and displaying the key subsequence of a voice sequence.

[0106] In Figure 4 it shows that a key subsequence recognition and display system for a voice sequence includes a voice receiving end, a voice recognition end, and a voice display end.

[0107] Among them, in this embodiment, the voice receiving end is used to receive a voice sequence;

[0108] The voice recognition end is used to recognize the pause points included in the voice sequence and whether the voice sequence includes a key subsequence;

[0109] The voice display end is used to display the key subsequence on the display interface in a predetermined format;

[0110] Among them, the voice recognition end recognizes whether the voice sequence includes a key subsequence with the pause points included in the voice sequence as nodes;

[0111] The key subsequence includes the reaction relationships of at least one chemical substance and / or chemical substance combination;

[0112] The displaying of the key subsequence on the display interface in a predetermined format includes: displaying at least one chemical formula corresponding to the at least one chemical substance and / or a chemical reaction formula corresponding to the chemical substance combination as a picture.

[0113] Specifically, the key subsequences are displayed on the display interface in a predetermined format, including: displaying at least one chemical formula corresponding to a chemical substance and / or a chemical reaction formula, chemical molecular formula, electronic structural formula of a chemical substance, etc. corresponding to a chemical substance combination in the form of a picture.

[0114] Based on this, further refer to Figure 4 . The system further includes a voice broadcast terminal; Figure 5 .

[0115] The voice broadcast terminal is used to broadcast the key subsequences while displaying the key subsequences on the display interface in a predetermined format at the voice display terminal.

[0116] Figure 6 Then it shows Figure 4 The internal structure schematic diagram of the voice recognition terminal in the key subsequence recognition and display system of the voice sequence.

[0117] Among them, the voice recognition terminal includes a pause point recognition unit, a subsequence segmentation unit, a subsequence recognition unit, and a subsequence merging unit;

[0118] The pause point recognition unit is used to recognize the pause points included in the voice sequence;

[0119] The subsequence segmentation unit is used to segment the voice sequence into multiple voice subsequences with the pause points as units;

[0120] The subsequence recognition unit is used to recognize whether each voice subsequence contains a predetermined key subsequence;

[0121] The subsequence merging unit is used to merge two adjacent voice subsequences into one voice subsequence if the two adjacent voice subsequences contain the same predetermined key subsequence.

[0122] Figure 7 It is the schematic diagram of the display effect after the implementation of the technical solution of the present invention.

[0123] Figure 7 The current voice sequence obtained in

[0124] "There are two cases in the reaction between sodium hydroxide and carbon dioxide: First: When a small amount of carbon dioxide reacts with sodium hydroxide, sodium carbonate and water are formed; Second: When an excessive amount of carbon dioxide reacts with sodium hydroxide, sodium bicarbonate is formed";

[0125] Among them, multiple voice pause points can be recognized, which are specifically manifested at the positions of corresponding punctuation marks, such as "colon :, semicolon ;, period." etc. The recognition and extraction of voice pause points belong to the prior art, and this embodiment will not expand on this.

[0126] For the voice subsequence before the first voice pause point (there are two cases in the reaction of sodium hydroxide and carbon dioxide), identify the predetermined key subsequence contained therein, that is, "chemical substance and / or chemical substance combination" (sodium hydroxide and carbon dioxide);

[0127] Therefore, while playing the voice sequence, display the key subsequence (sodium hydroxide and carbon dioxide) in a predetermined format, such as Figure 7 the electronic (molecular) structural formula of sodium hydroxide and carbon dioxide described in

[0128] Similarly, continue to identify the predetermined key subsequence, that is, "chemical substance and / or chemical substance combination" (a small amount of carbon dioxide reacts with sodium hydroxide to form sodium carbonate and water). While playing the voice sequence, display the key subsequence (a small amount of carbon dioxide reacts with sodium hydroxide to form sodium carbonate and water) in a predetermined format, such as Figure 7 the first reaction combination formula of sodium hydroxide and carbon dioxide described in

[0129] Similarly, continue to identify the predetermined key subsequence, that is, "chemical substance and / or chemical substance combination" (excess carbon dioxide reacts with sodium hydroxide to form sodium bicarbonate). While playing the voice sequence, display the key subsequence (excess carbon dioxide reacts with sodium hydroxide to form sodium bicarbonate) in a predetermined format, such as Figure 7 the second reaction combination formula of sodium hydroxide and carbon dioxide described in

[0130] It can be seen that the technical solution of the present invention is essentially to generate and display a summary picture containing the key subsequence when it is recognized that there is a key subsequence in the original voice sequence.

[0131] Therefore, the technical solution of the present invention can also achieve the following:

[0132] S2011: Identify the current pause point in the acquired voice sequence;

[0133] S2021: Translate the voice subsequence before the current pause point into a text subsequence;

[0134] S2031: Determine whether the text subsequence contains a predetermined key subsequence;

[0135] If it contains, generate a picture summary corresponding to the voice subsequence based on the predetermined key subsequence, and display the picture summary corresponding to the voice subsequence in a predetermined format while playing the voice subsequence.

[0136] If not included, directly play the text subsequence, obtain the next pause point in the voice sequence, use the next pause point as the current pause point, and return to step S2021.

[0137] It can be seen that in view of the problem that in some special occasions, especially in occasions containing specific vocabulary or specific format vocabulary, such as in the chemical field, molecular field, etc., in voice communication, the voice sequence usually contains a large number of professional chemical vocabulary, the present invention can highlight the key subsequences in these voice sequences, enabling both parties of the voice interaction to quickly and accurately absorb the key content and improving the voice communication efficiency.

[0138] The present invention can also real-time recognize the audio stream as text, which is applicable to long sentence voice input, audio-visual subtitles, meetings, speech subtitles on the same screen, etc., intelligent language processing. It can also perform intelligent error correction on the intermediate results of recognition, display the intermediate text results in real-time, quickly recognize the key audio stream, and quickly display the recognition results in the form of a picture text summary. For example, at least one chemical formula corresponding to a chemical substance and / or a chemical reaction formula, chemical molecular formula, electronic structural formula of a chemical substance combination corresponding to the chemical substance are displayed in a picture.

[0139] It should be noted that Figure 1 - Figure 3 the above steps or the above method and process can all be automatically implemented through computer program instructions. Therefore, referring to Figure 8 , a terminal device is provided. The terminal device can be a data interaction device, including a bus, a processor, and a memory. The memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium.

[0140] Specifically, the terminal device can be an electronic device, including: a processor for executing Figure 1 or Figure 2 or Figure 3 a method for identifying and displaying key subsequences of a voice sequence of the above method; and a memory coupled to the processor for storing instructions executed by the processor.

[0141] These computer program instructions can also be loaded onto a computer or other programmable power secondary equipment, so that a series of operation steps are executed on the computer or other programmable power secondary equipment to generate computer-implemented processing. Thus, the instructions executed on the computer or other programmable power secondary equipment provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or Figure 1 one block or multiple blocks.

[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent substitutions, and any modification or equivalent substitution that does not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

[0143] For the part of the module structure not specifically defined in the present invention, the content recorded in the prior art shall prevail. The prior art mentioned in the foregoing background art section of the present invention can be used as a part of the present invention to understand the meaning of some technical features or parameters. The protection scope of the present invention shall be subject to the content actually recorded in the claims.

Claims

1. A method for identifying and displaying key subsequences of a voice sequence, characterized in that, The method includes the following steps: S1: Obtain a speech sequence; the speech sequence includes professional chemical terms; S2: Identify the current pause point in the obtained speech sequence; S3: Translate the speech subsequence before the current pause point into a text subsequence; S4: Determine whether the text subsequence contains a predetermined key subsequence; the key subsequence includes the reaction relationship of at least one chemical substance and / or chemical substance combination; If it contains, generate a picture summary corresponding to the speech subsequence based on the predetermined key subsequence, display the picture summary corresponding to the speech subsequence in a predetermined format while playing the speech subsequence; and obtain the next pause point in the speech sequence, use the next pause point as the current pause point, and return to step S2; If it does not contain, directly play the speech subsequence, obtain the next pause point in the speech sequence, use the next pause point as the current pause point, and return to step S2; Step S1 further includes: After obtaining the speech sequence, determine the first number of pause points included in the speech sequence; Segment the speech sequence into a third number of speech subsequences; Activate the recognition processes of a second number of key subsequences; the second number is greater than the third number; Step S4 includes: Identify whether there is a predetermined key subsequence in the speech sequence through the recognition processes of the second number of key subsequences.

2. A method for identifying and displaying key subsequences of a speech sequence according to claim 1, wherein: The speech sequence obtained in step S1 includes multiple pause points; After step S1 and before step S2, the method further includes the following steps: Taking the pause points as units, segment the speech sequence into multiple speech subsequences.

3. A method for identifying and displaying key subsequences of a speech sequence according to claim 1, wherein: Step S4 further includes: If two adjacent speech subsequences contain the same predetermined key subsequence, merge the two adjacent speech subsequences into one speech subsequence.

4. A system for identifying and displaying key subsequences of a speech sequence, which is used to implement the method for identifying and displaying key subsequences of a speech sequence according to any one of claims 1-3; the system includes a speech receiving end, a speech recognition end, and a speech display end, wherein: The speech receiving end is used to receive a speech sequence; The speech recognition end is used to identify the pause points included in the speech sequence and whether the speech sequence contains key subsequences; The speech display end is used to display the key subsequences in a predetermined format on a display interface; Among them, the speech recognition end identifies whether the speech sequence contains key subsequences with the pause points included in the speech sequence as nodes; The key subsequence includes the reaction relationship of at least one chemical substance and / or chemical substance combination; The displaying the key subsequences in a predetermined format on the display interface includes: displaying at least one chemical formula corresponding to the at least one chemical substance and / or a chemical reaction formula corresponding to a chemical substance combination in a picture; 5. The key subsequence recognition and display system for a speech sequence according to claim 4, characterized in that The system further includes a speech broadcast end; The voice broadcast terminal is configured to broadcast the key subsequence while the voice display terminal displays the key subsequence on the display interface in a predetermined format.

6. The key subsequence recognition and display system for a speech sequence according to claim 4, characterized in that The voice recognition terminal includes a pause point recognition unit, a subsequence segmentation unit, a subsequence recognition unit, and a subsequence merging unit; The pause point recognition unit is configured to recognize the pause points included in the voice sequence; The subsequence segmentation unit is configured to segment the voice sequence into a plurality of voice subsequences with the pause points as units; The subsequence recognition unit is configured to recognize whether each of the voice subsequences includes a predetermined key subsequence; The subsequence merging unit is configured to merge two adjacent voice subsequences into one voice subsequence if the two adjacent voice subsequences include the same predetermined key subsequence.

7. An electronic device, characterized in that, Comprising: a processor configured to execute the method for recognizing and displaying the key subsequence of a voice sequence according to any one of claims 1 to 3; and a memory coupled to the processor for storing instructions executed by the processor.

Citation Information

Patent Citations

  • A voice information processing method, electronic device, and computer storage medium

    CN114267352B

  • Conference record generation method based on voice recognition, device and storage medium

    CN110335612A

  • Voice recognition method, voice recognition device, electronic equipment and storage medium

    CN110473551A

  • Audio label intelligent labeling method and device based on teaching video

    CN111510765A

  • Voice-based input method and device

    CN111651961A