Method and apparatus for determining target voice information, electronic device, and storage medium

By updating the first preset keyword in intelligent speech recognition to determine the start and end of the target speech text, and using the second preset keyword to identify the type, the problem of quickly determining the target speech information under limited resources is solved, and efficient speech information recognition is achieved.

CN119889326BActive Publication Date: 2025-11-21CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510052944.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-11-21
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

Existing intelligent speech recognition technologies, when faced with limited data processing resources and low hardware capabilities, are unable to quickly determine the target speech information required by users, resulting in a waste of resources and time.

Method used

The target start and end keywords in the initial speech text are determined by the updated first preset keywords, the target speech text is extracted, and the keyword type is identified by the second preset keywords, so as to achieve rapid recognition of target speech information.

Benefits of technology

It reduces the need for data processing resources, adapts to changes in spoken expression, quickly identifies key voice information, and ensures that no critical information is missed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119889326B_ABST
    Figure CN119889326B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of speech recognition, and discloses a target speech information determination method and device, electronic equipment and a storage medium. The method comprises the following steps: in response to a first determination instruction, determining a target starting keyword and a target ending keyword in an initial speech text based on an updated first preset keyword; determining a target speech text from the initial speech text according to the target starting keyword and the target ending keyword; determining a first target keyword in the target speech text and a type to which the first target keyword belongs according to a second preset keyword, so as to determine target speech information in the target speech text. According to the application, the first target keyword of the corresponding type is determined from the target speech text according to the second preset keyword, key speech information can be quickly recognized, and the target speech information in the target speech text can be quickly determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, specifically to a method, apparatus, electronic device, and storage medium for determining target speech information. Background Technology

[0002] With the continuous development of intelligent technologies, intelligent speech recognition technology has become a hot field. Intelligent speech recognition technology mainly relies on complex recognition models, which can recognize the entire text of the initial speech, that is, directly translate the entire initial speech into speech text. This obviously requires huge data processing resources and data processing time. Especially when data processing resources are limited, hardware level is low, and recognition requirements are not high, it is impossible to quickly determine the target speech information required by the user. Summary of the Invention

[0003] In view of the above problems, this application provides a method, apparatus, electronic device and storage medium for determining target speech information, which can quickly determine the target speech information required by the user.

[0004] According to one aspect of this application, a method for determining target speech information is provided. The method includes: responding to a first determining instruction, determining a target starting keyword and a target ending keyword in an initial speech text based on an updated first preset keyword; wherein the target starting keyword and the target ending keyword are respectively the first occurrence of a starting keyword and the last occurrence of a ending keyword in the initial speech text, and the updated first preset keyword is a word obtained by updating the original first preset keyword; determining a target speech text from the initial speech text based on the target starting keyword and the target ending keyword; and determining a first target keyword in the target speech text and the type to which the first target keyword belongs based on a second preset keyword, thereby determining target speech information in the target speech text.

[0005] In an optional embodiment, the determining method further includes: in response to a second determining instruction, determining a second target keyword in the initial speech text and the type to which the second target keyword belongs, based on the second preset keyword, so as to determine the target speech information in the initial speech text; wherein the number of the second target keywords is greater than or equal to the number of the first target keywords.

[0006] In an optional manner, the determination method further includes: if the specified keyword in the target speech text fails to match all second preset keywords, then the specified keyword is escaped based on a preset character escape function; if the escaped characters include characters representing time, then the specified keyword is determined to be a first target keyword of type time.

[0007] In one optional approach, determining the first target keyword in the target speech text and the type to which the first target keyword belongs, based on the second preset keyword, to determine the target speech information in the target speech text, includes: taking words in the target speech text that match the second preset keyword as the first target keyword, and taking the preset type corresponding to the successfully matched second preset keyword as the type to which the first target keyword belongs; wherein, each second preset keyword corresponds to its own preset type; and taking the type to which the first target keyword belongs as the information header of the first target keyword to determine the target speech information in the target speech text.

[0008] In one optional approach, before using the words in the target speech text that match the second preset keyword as the first target keyword, the determination method further includes: traversing each preset type and using the second preset keyword corresponding to the traversed preset type as the target second preset keyword; wherein each preset type corresponds to at least one second preset keyword, and the second preset keywords corresponding to each preset type include a separator character; matching each word in the target speech text with the target second preset keyword to obtain the words in the target speech text that match the target second preset keyword, thereby obtaining the words in the target speech text that match the second preset keyword.

[0009] In one optional approach, determining the target speech text from the initial speech text based on the target start keyword and the target end keyword includes: taking the speech text following the target start keyword in the initial speech text as the first speech text, and taking the speech text preceding the target end keyword in the initial speech text as the second speech text; and determining the speech text overlapping between the first speech text and the second speech text as the target speech text.

[0010] In one optional approach, the first preset keyword includes a preset starting keyword and a preset ending keyword; determining the target starting keyword and target ending keyword in the initial speech text based on the updated first preset keyword includes: if the target word matches the preset starting keyword, then the target word is determined as the starting keyword in the initial speech text, thereby determining all starting keywords in the initial speech text; wherein, the target word is any word in the initial speech text; if the target word matches the preset ending keyword, then the target word is determined as the ending keyword in the initial speech text, thereby determining all ending keywords in the initial speech text; the first occurrence of the starting keyword in the initial speech text is taken as the target starting keyword, and the last occurrence of the ending keyword in the initial speech text is taken as the target ending keyword.

[0011] According to another aspect of this application, a device for determining target speech information is provided. The device includes: a first response module, responding to a first determination instruction, determining a target start keyword and a target end keyword in an initial speech text based on an updated first preset keyword; wherein the target start keyword and the target end keyword are respectively the first occurrence of the start keyword and the last occurrence of the end keyword in the initial speech text, and the updated first preset keyword is a word obtained by updating the original first preset keyword; a text determination module, determining a target speech text from the initial speech text according to the target start keyword and the target end keyword; and an information determination module, determining a first target keyword in the target speech text and the type to which the first target keyword belongs, based on a second preset keyword, to determine the target speech information in the target speech text.

[0012] According to one aspect of this application, an electronic device is provided, comprising: a controller; and a memory for storing one or more programs, which, when executed by the controller, perform the determination method described above.

[0013] According to one aspect of this application, a computer-readable storage medium is also provided, on which computer-readable instructions are stored, which, when executed by a computer's processor, cause the computer to perform the determination method described above.

[0014] According to one aspect of this application, a computer program product or computer program is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the determination method described above.

[0015] This application uses updated first preset keywords to determine the target start and end keywords in the initial speech text, enabling the extraction of the target speech text and focusing on recognizing a portion of the initial speech text, thereby reducing the need for data processing resources. Furthermore, this application can update the first preset keywords in real time to adapt to constantly changing spoken expressions, thus better identifying the target speech text containing the target speech information. Based on second preset keywords, this application determines the corresponding type of first target keyword from the target speech text, enabling rapid identification of key speech information and quickly pinpointing the target speech information within the target speech text.

[0016] The above description is merely an overview of the technical solutions of the embodiments of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0018] Figure 1 This is a flowchart illustrating a method for determining target speech information according to an exemplary embodiment of this application.

[0019] Figure 2 This is a schematic diagram of a function switch setting interface shown in an exemplary embodiment of this application.

[0020] Figure 3 Based on Figure 1 The exemplary embodiment shown illustrates a flowchart of another method for determining target speech information.

[0021] Figure 4 Based on Figure 1 The exemplary embodiment shown illustrates a flowchart of another method for determining target speech information.

[0022] Figure 5 Based on Figure 1 , Figure 3 , Figure 4 The exemplary embodiment shown in any of the examples illustrates a flowchart of another method for determining target speech information.

[0023] Figure 6 This is a flowchart illustrating the execution scenario of the method for determining the target speech information in this application.

[0024] Figure 7 This is a schematic diagram of the structure of a target voice information determination device shown in an exemplary embodiment of this application.

[0025] Figure 8 This is a schematic diagram of the structure of a computer system for an electronic device, as illustrated in an exemplary embodiment of this application. Detailed Implementation

[0026] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0027] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0028] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0029] In this application, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0030] Intelligent speech recognition technology mainly relies on complex recognition models that can recognize the entire text of the initial speech, that is, directly translate the entire initial speech into speech text. This obviously requires huge data processing resources and data processing time. Especially when data processing resources are scarce, hardware level is low, and recognition requirements are not high, it is impossible to quickly determine the target speech information required by the user.

[0031] Therefore, one aspect of this application provides a method for determining target speech information. Please refer to [link / reference needed]. Figure 1 , Figure 1 This is a flowchart illustrating a method for determining target speech information according to an exemplary embodiment of this application. The method includes at least steps S110 to S130, which are described in detail below:

[0032] S110: In response to the first determination instruction, based on the updated first preset keyword, determine the target start keyword and target end keyword in the initial speech text; wherein, the target start keyword and target end keyword are the first occurrence of the start keyword and the last occurrence of the end keyword in the initial speech text, respectively, and the updated first preset keyword is the word obtained by updating the original first preset keyword.

[0033] The first confirmation instruction indicates that upon recognizing the starting keyword, the function to confirm the target voice information is activated, and upon recognizing the ending keyword, the confirmation process ends. The starting keyword is a word that precedes important information, such as "I'm recording this," "Need to confirm this," "I'm confirming this," "Order information," "Detailed information," "Transfer information," "Information confirmation," "Content recorded," etc. The ending keyword is a word that concludes important information, such as "Thank you, goodbye," "All information confirmed," "Content recorded," "Summary complete," "Everything is handled," etc.

[0034] The first preset keyword update permission is open to users, meaning users can update the first preset keyword in real time through voice input, text input, and other methods, including but not limited to modification, addition, and deletion, to adapt to different usage scenarios. The updated first preset keyword can be persistently stored, for example, in a configuration file named globals.json, ensuring keyword persistence and consistency for quick retrieval when needed.

[0035] For example, the first preset keyword includes a preset starting keyword and a preset ending keyword; S110 includes: if the target word matches the preset starting keyword successfully, then the target word is determined as the starting keyword in the initial speech text, so as to determine all the starting keywords in the initial speech text; wherein, the target word is any word in the initial speech text; if the target word matches the preset ending keyword successfully, then the target word is determined as the ending keyword in the initial speech text, so as to determine all the ending keywords in the initial speech text; the first occurrence of the starting keyword in the initial speech text is taken as the target starting keyword, and the last occurrence of the ending keyword in the initial speech text is taken as the target ending keyword.

[0036] For example, by simultaneously loading preset start and end keywords using two different threads, and then matching each word in the initial speech text with these preset start and end keywords, all start and end keywords in the initial speech text can be identified, thereby determining the target start and end keywords. Parallel thread matching allows for rapid analysis of the initial speech text, quickly determining the target start and end keywords for rapid extraction of the target speech text. The target words matched by the two parallel threads at the same time can be the same or different.

[0037] In some embodiments, in order to determine the target starting keyword and the target ending keyword more quickly, the thread that determines the starting keyword matches the words in the initial text in ascending order (i.e., matches from the beginning words of the initial text) and takes the first matched starting keyword as the target starting keyword; the thread that determines the ending keyword matches the words in the initial text in descending order (i.e., matches from the end words of the initial text) and takes the first matched ending keyword as the target ending keyword.

[0038] In some scenarios, the initial voice text originates from a phone call and contains at least one start keyword and at least one end keyword. Related technologies extract the voice text using a start keyword and an adjacent end keyword, which can lead to the omission of important information. For example, the initial voice text might be: "Hello!... Let me confirm... Everything is handled well. We'll meet at XX at 9:00. Summary complete." Here, "Let me confirm" is the start keyword, and "Everything is handled well" and "Summary complete" are the end keywords. If the extraction method used in these technologies is followed, the recognition of subsequent key information will stop upon recognizing "Everything is handled well," thus missing the crucial information "We'll meet at XX at 9:00."

[0039] Therefore, this embodiment uses the first occurrence of the starting keyword and the last occurrence of the ending keyword as the extraction keywords for the target speech, thus ensuring that no key information is omitted. In some embodiments, the initial speech includes a starting keyword but not an ending keyword. In this case, the speech text between the adjacent words following the first starting keyword (i.e., the target starting keyword) and the end word of the text is taken as the target speech text. That is, by default, the last word in the initial text is taken as the target ending keyword.

[0040] S120: Determine the target speech text from the initial speech text based on the target start keyword and the target end keyword.

[0041] When the target starting keyword is identified, all words in the speech information following that word are analyzed; when the target ending keyword is identified, speech analysis stops, meaning there is no need to analyze the speech text following that keyword. Thus, the target starting keyword and the target ending keyword are used as two end-value keywords to quickly extract the target speech text from the initial speech text, avoiding unnecessary analysis of other parts of the speech text.

[0042] For example, the speech text following the target starting keyword in the initial speech text is taken as the first speech text, and the speech text before the target ending keyword in the initial speech text is taken as the second speech text; the speech text overlapping between the first speech text and the second speech text is determined as the target speech text.

[0043] For example, the thread determining the starting keyword matches words in the initial text in ascending order, taking the first matched starting keyword as the target starting keyword, and the speech text following that word as the first speech text. The thread determining the ending keyword matches words in the initial text in reverse order, taking the first matched ending keyword as the target ending keyword, and the speech text following that word as the second speech text. The intersection of the first and second texts is then processed, and the overlapping text is the target speech text.

[0044] S130: Based on the second preset keyword, determine the first target keyword in the target speech text and the type to which the first target keyword belongs, so as to determine the target speech information in the target speech text.

[0045] The second preset keywords are words used to identify substantive information in the target speech text. These include, but are not limited to, words representing various types of meanings, such as vehicle type, location, transportation status, cargo type, and time type. Each type includes at least one second preset keyword. For example, the second preset keywords for vehicle types include: medium-sized van, truck, moving truck, enclosed truck, etc.

[0046] In some embodiments, the number of first target keywords is multiple, and each first target keyword and its type can be combined, and target speech information can be constructed based on multiple combinations.

[0047] This embodiment uses updated first preset keywords to determine the target start and end keywords in the initial speech text, enabling the extraction of the target speech text and focusing on identifying a portion of the initial speech text, thereby reducing the demand for data processing resources. This embodiment can update the first preset keywords in real time to adapt to constantly changing spoken expressions, thus better identifying the target speech text containing the target speech information. This embodiment also determines the corresponding type of first target keyword from the target speech text based on second preset keywords, enabling rapid identification of key speech information and quickly determining the target speech information within the target speech text.

[0048] This embodiment triggers the target speech information determination method only when a relevant first preset keyword is identified. To provide more flexible target speech information recognition, an alternative method that always triggers the method is also provided, where target speech information determination is not performed only after the first preset keyword is identified. For example... Figure 2 As shown, the function switch has three different activation modes: "Always On", "First Preset Keyword Trigger On", and "Always Off". Users can also modify the preset start keyword and preset end keyword accordingly.

[0049] Please refer to details. Figure 3 , Figure 3 Based on Figure 1 The exemplary embodiment shown illustrates a flowchart of another method for determining target speech information. This method, in the example described, [describes a method for determining target speech information]. Figure 1 Based on S110 to S130 shown, at least S310 is also included, which is described in detail below:

[0050] S310: In response to the second determination instruction, determine the second target keyword in the initial speech text and the type to which the second target keyword belongs based on the second preset keyword, so as to determine the target speech information in the initial speech text; wherein the number of the second target keyword is greater than or equal to the number of the first target keyword.

[0051] The second determination instruction is an instruction that unconditionally triggers the target speech information determination function, that is, without the need to perform a matching process of the first preset keyword on the initial speech text, it directly matches the second preset keyword starting from the first word of the initial speech text.

[0052] For example, the initial voice text is: "Hello! ... Let me confirm... Everything is done. We'll meet at XX at 9:00. That's all." In response to the second determination instruction, each word in the initial voice text is matched one by one with the second preset keyword, starting from "Hello," to directly determine the second target keyword in the initial voice text and the type of the second target keyword, thus constructing the target voice information.

[0053] This method of determination, because it directly recognizes the speech information of the initial speech text, can identify all the key information in the initial speech text without omission, thus making the identified target speech information more complete.

[0054] When identifying time-related keywords, the second preset keyword cannot cover all time-related keywords due to the diversity of colloquial expressions of these keywords, resulting in omissions in the identification of time information.

[0055] Therefore, in another exemplary embodiment of this application, a method for determining time-type keywords is described, for details please refer to [link / reference needed]. Figure 4 , Figure 4 Based on Figure 1 The exemplary embodiment shown illustrates a flowchart of another method for determining target speech information. This method, in the example described, [describes a method for determining target speech information]. Figure 1 Based on S110 to S130 shown, at least S410 to S420 are also included, detailed below:

[0056] S410: If the specified keyword in the target speech text fails to match all the second preset keywords, then the specified keyword is escaped based on the preset character escape function.

[0057] In one embodiment, the preset character escape function is a reversible function, which can escape a specified keyword into a specified character, or escape a specified character into a specified keyword. That is, the preset character escape function is a mapping relationship between a specified keyword and a specified character.

[0058] In another embodiment, the preset character escape function is a decryption function. By inputting the specified keyword and the preset key into the decryption function, the corresponding character can be decrypted.

[0059] S420: If the escaped characters include characters representing time, then the specified keyword is determined to be the first target keyword of type time.

[0060] For example, if the specified keyword "9:01" does not match any of the second preset keywords, it will be escaped as 9:01 or 09-01 based on the preset character escape function. Here, "point" and "-" are characters that represent time. Therefore, "9:01" is determined as the first target keyword of the time type. When outputting the target voice information, it can directly output "time is 9:01", or output "time is 9:01", or output "09-01".

[0061] In some embodiments, if a first target keyword of a time type is determined, it will be converted into a more accurate expression based on the current time. For example, if the current time is 9:00 AM on December 17, 2024, and the first target keyword is "2 PM", the two will be combined, that is, "2 PM" will be converted into "2:00 PM on December 17, 2024". This unifies the time type keywords in the entire target speech text, making it easier for users to quickly extract the corresponding time parameters based on the same time expression standard.

[0062] In another exemplary embodiment of this application, it is described in detail how to determine the first target keyword in the target speech text and the type to which the first target keyword belongs based on the second preset keyword, so as to determine the target speech information in the target speech text. Please refer to [link to relevant documentation] for details. Figure 5 , Figure 5 Based on Figure 1 , Figure 3 , Figure 4 The exemplary embodiment shown in any of the examples illustrates a flowchart of another method for determining target speech information. This method further includes at least S510 to S520 in the above-described S130, which are described in detail below:

[0063] S510: Take the words in the target speech text that match the second preset keyword as the first target keyword, and take the preset type corresponding to the successfully matched second preset keyword as the type to which the first target keyword belongs; wherein, each second preset keyword has its own preset type.

[0064] Different second preset keywords can correspond to the same or different preset types, depending on their meaning. For example, the preset type corresponding to the second preset keyword "in transit" is the transportation status type, the preset type corresponding to the second preset keyword "arrived" is the transportation status type, and the preset type corresponding to the second preset keyword "empty vehicle" is the vehicle status type.

[0065] The target speech text contains multiple words, and all of these words can be matched with various second preset keywords. The matching process also needs to consider the preset type to which each second preset keyword belongs. The following is an example of how to determine the words that match the second preset keywords:

[0066] Iterate through each preset type and take the second preset keyword corresponding to the preset type as the target second preset keyword; wherein, each preset type has at least one second preset keyword, and the second preset keywords corresponding to each preset type include a separator character; match each word in the target speech text with the target second preset keyword to obtain the words in the target speech text that match the target second preset keyword, so as to obtain the words in the target speech text that match the second preset keyword.

[0067] By traversing through the text, each word in the target speech text can be matched without omission. Then, the second preset keyword of the preset type that they traversed is matched one by one, so as to accurately determine all words in the target speech text that successfully match the second preset keyword. This process is repeated to determine the words in the target speech text that match the second preset keyword.

[0068] S520: Use the type to which the first target keyword belongs as the information header of the first target keyword to determine the target speech information in the target speech text.

[0069] The category to which the primary target keyword belongs is derived from a broader overview of its meaning, representing a directory or set of categories to which it belongs. For example, the primary target keyword "morning" belongs to the time category, meaning that "morning" is classified as a time-related word; the primary target keywords "rainy day," "cloudy day," and "sunny day" belong to the weather category.

[0070] Placing the type of the first target keyword before the first target keyword allows for a concise and summary description of the key information of the adjacent first target keyword. For example, if the target speech information is "Time: 9:10", it can intuitively and concisely express the key information that the target speech text wants to convey.

[0071] In another exemplary embodiment of this application, the execution scenarios of the above-mentioned multiple determination methods are illustrated by way of example. Please refer to the following for details. Figure 6 , Figure 6This is a flowchart illustrating the execution scenario of the method for determining the target speech information in this application. The execution entity in this scenario is a speech recognition program. After starting, the speech recognition program determines the file to load based on different determination instructions. These include a first determination instruction (enabled upon matching a first preset keyword), a second determination instruction (always enabled), and a third determination instruction (always disabled). If the speech recognition program detects the first or second determination instruction, it calls the `updateEnginer` function, which is responsible for updating the global rule engine configuration. Inside the function, the previous coroutine job is first canceled (if it exists), and then a new coroutine is started, using IO streaming technology to read the contents of the corresponding files (filter.json and / or global.json).

[0072] If it is the first confirmed instruction, two threads in the speech recognition program simultaneously load the filter.json and global.json files to quickly load the first preset keyword from global.json and the second preset keyword from filter.json into the speech recognition program. For example, the FileUtil.openAssets() method is used to read the globals.json and filter.json files in the corresponding child thread. After reading the two lists from globals.json, they are converted into set collections in the code, as shown in the following example code:

[0073] val startKeywords = setOf("I'll record this", "I'll confirm this", "Need to confirm this", "Order Information", "Detailed Information", "Transfer Information", "Information Confirmation", "Content Record")

[0074] val endKeywords = setOf("Thank you and goodbye","All information confirmed","Content recorded","Summary complete","Task completed","All information confirmed","Everything handled","Handover complete","No other issues","Confirmed","End recording")

[0075] In the `fifter.json` file, after reading each list, the code converts it into a regular expression object, for example:

[0076] val cargo type =

[0077] Regex("[^。]*?(furniture|home appliances|books|clothing|kitchenware|office equipment|personal items|cardboard boxes|home furnishings|electronics|clothing boxes|luggage|moving items|heavy equipment|moving tools|television|refrigerator|sofa|mattress|dining table|office desk|computer|bookshelf|audio equipment|kitchen appliances|small appliances|fitness equipment)[^。]*[。]")

[0078] val time =

[0079] Regex("[^。]*?(morning|noon|afternoon|evening|about to depart|ready to depart|on time|delayed|overtime|urgent|time-sensitive|long distance|short distance|about to arrive|estimated arrival|scheduled time|cargo loading time|urgent transport|rest stop|peak hours|off-peak hours|scheduled delivery time|vehicle dispatch time)[^。]*[.]")

[0080] The system matches words in the initial spoken text against a first preset keyword, for example, using the `matchGlobals(result:String, globals:GlobalDescript?=null)` method. Here, `result` represents the initial spoken text. `Globals` represents the data structure (`startKeywords`, `endKeywords`) in the start and end rule engine, containing a set of different commands and keywords. An example matching code is as follows:

[0081] val startKeywords = setOf("I'll record this", "I'll confirm this", "Need to confirm this", "Order Information", "Detailed Information", "Transfer Information", "Information Confirmation", "Content Record")

[0082] val endKeywords = setOf("Thank you and goodbye","All information confirmed","Content recorded","Summary complete","Task completed","All information confirmed","Everything handled","Handover complete","No other issues","Confirmed","End recording")

[0083] fun extractSubTextWithRegex(text:String,startKeywords:Set <string>,endKeywords:Set <string>):String? {

[0084] / / Construct a regular expression to match paragraphs that begin with any keyword in startKeywords and end with any keyword in endKeywords.

[0085] val startPattern=startKeywords.joinToString("|"){Regex.escape(it)}

[0086] val endPattern=endKeywords.joinToString("|"){Regex.escape(it)}

[0087] val regexPattern="($startPattern).*?($endPattern)"

[0088] val regex=Regex(regexPattern,RegexOption.DOT_MATCHES_ALL)

[0089] / / Find matching paragraphs

[0090] val matchResult=regex.find(text)

[0091] return matchResult?.value

[0092] }

[0093] The `startKeywords` and `endKeywords` sets contain preset start keywords and preset end keywords, respectively. For example, `startKeywords` contains phrases like "I'm recording" and "I'm confirming," indicating that these can be used as start markers; `endKeywords` contains phrases like "Thank you, goodbye" and "Information confirmed," indicating that these can be used as end markers.

[0094] The extractSubTextWithRegex function extracts a portion of the text that meets the following conditions: it starts with a keyword in startKeywords and ends with a keyword in endKeywords.

[0095] The construction of regular expressions: `startPattern` and `endPattern` are regular expression patterns dynamically constructed based on `startKeywords` and `endKeywords`. `joinToString("|")` concatenates all elements in `startKeywords` and `endKeywords` into a single string separated by the `|` symbol, allowing you to match any keyword in the regular expression. `Regex.escape(it)` is used to escape any special characters (such as `.`, `*`, etc.) that may be included, ensuring the correctness of the regular expression.

[0096] Regular expression matching: The regular expression ($startPattern).*? ($endPattern) will match strings that begin with a keyword in startPattern and end with a keyword in endPattern. Here, .*? is a non-greedy match; it tries to match as few characters as possible in the middle, ensuring that the first matching substring is found.

[0097] regex.find(text): Use regex.find() to find the first match in text that meets the given criteria.

[0098] Return value: If a match is found, return that content. If no match is found, return "Nomatching subText found."

[0099] If the match is successful, the target speech text is determined and matched with the second preset keyword to determine the first target keyword in the target speech text and the type of the first target keyword, so as to determine the target speech information in the target speech text.

[0100] If it is the second determination instruction, the speech recognition program loads the filter.json file, that is, loads the second preset keyword in it into the speech recognition program, matches the words in the initial speech text with the second preset keyword, determines the second target keyword in the initial speech text, and the type to which the second target keyword belongs, so as to determine the target speech information in the initial speech text.

[0101] If it is a third determination instruction, then the file is not loaded and the relevant determination steps for the target voice information are not executed.

[0102] Another aspect of this application provides a device for determining target speech information, such as... Figure 7 As shown, Figure 7 This is a schematic diagram illustrating the structure of a target speech information determination device according to an exemplary embodiment of this application. The determination device 700 includes:

[0103] The first response module 710, in response to the first determination instruction, determines the target start keyword and target end keyword in the initial speech text based on the updated first preset keyword; wherein, the target start keyword and target end keyword are the first occurrence of the start keyword and the last occurrence of the end keyword in the initial speech text, respectively, and the updated first preset keyword is the word obtained by updating the original first preset keyword.

[0104] The text determination module 730 determines the target speech text from the initial speech text based on the target start keyword and the target end keyword;

[0105] The information determination module 750 determines the first target keyword in the target speech text and the type to which the first target keyword belongs based on the second preset keyword, so as to determine the target speech information in the target speech text.

[0106] In another exemplary embodiment, the determining device 700 further includes:

[0107] The second response module, in response to the second determination instruction, determines the second target keyword in the initial speech text and the type to which the second target keyword belongs, based on the second preset keyword, so as to determine the target speech information in the initial speech text; wherein, the number of the second target keyword is greater than or equal to the number of the first target keyword.

[0108] In another exemplary embodiment, the determining device 700 further includes:

[0109] In the matching failure module, if the specified keyword in the target speech text fails to match all the second preset keywords, the specified keyword is escaped based on the preset character escape function.

[0110] The first target keyword determination module determines that if the escaped characters include characters representing time, then the specified keyword is determined to be the first target keyword of type time.

[0111] In another exemplary embodiment, the information determination module 750 includes:

[0112] The matching unit takes the words in the target speech text that match the second preset keyword as the first target keyword, and takes the preset type corresponding to the successfully matched second preset keyword as the type to which the first target keyword belongs; wherein, each second preset keyword has its own preset type;

[0113] The unit is determined by using the type of the first target keyword as the information header of the first target keyword to identify the target speech information in the target speech text.

[0114] In another exemplary embodiment, the determining device 700 further includes:

[0115] The traversal module iterates through each preset type and takes the second preset keyword corresponding to the traversed preset type as the target second preset keyword; wherein, each preset type has at least one second preset keyword, and the second preset keywords corresponding to each preset type include separator characters;

[0116] The word matching module matches each word in the target speech text with the target second preset keyword to obtain the words in the target speech text that match the target second preset keyword.

[0117] In another exemplary embodiment, the text determination module 730 includes:

[0118] The text determination unit takes the speech text after the target start keyword in the initial speech text as the first speech text and the speech text before the target end keyword in the initial speech text as the second speech text.

[0119] The target speech text determination unit determines the speech text that overlaps between the first speech text and the second speech text as the target speech text.

[0120] In another exemplary embodiment, the first preset keyword includes a preset start keyword and a preset end keyword; the first response module 710 includes:

[0121] If the target word matches the preset starting keyword, the unit that successfully matches the starting keyword will determine the target word as the starting keyword in the initial speech text, thereby identifying all starting keywords in the initial speech text; where the target word is any word in the initial speech text.

[0122] The unit where the keyword matching is successful ends: If the target word matches the preset end keyword successfully, the target word is determined as the end keyword in the initial speech text, so as to determine all the end keywords in the initial speech text;

[0123] The target keyword determination unit takes the first occurrence of the starting keyword in the initial audio text as the target starting keyword and the last occurrence of the ending keyword in the initial audio text as the target ending keyword.

[0124] This application's determining device identifies the target start keyword and target end keyword in the initial speech text by using updated first preset keywords. This allows for the extraction of the target speech text, focusing on recognizing specific portions of the initial speech text and reducing the need for data processing resources. Furthermore, the device can update the first preset keywords in real time to adapt to changing spoken language, thus improving the identification of the target speech information within the target speech text. Based on second preset keywords, the device identifies the corresponding type of first target keyword from the target speech text, enabling rapid identification of key speech information and quickly pinpointing the target speech information within the target speech text.

[0125] It should be noted that the determining device provided in the above embodiments and the determining method provided in the foregoing embodiments belong to the same concept. The specific ways in which each module and unit performs operations have been described in detail in the method embodiments, and will not be repeated here.

[0126] Another aspect of this application provides an electronic device, including: a controller; and a memory for storing one or more programs, which, when executed by the controller, perform the determination method described above.

[0127] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a computer system for an electronic device, illustrating an exemplary embodiment of this application. It shows a schematic diagram of the structure of a computer system suitable for implementing the embodiments of this application.

[0128] It should be noted that, Figure 8 The computer system 800 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0129] like Figure 8 As shown, the computer system 800 includes a Central Processing Unit (CPU) 801, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on programs stored in Read-Only Memory (ROM) 802 or programs loaded from storage portion 808 into Random Access Memory (RAM) 803. The RAM 803 also stores various programs and data required for system operation. The CPU 801, ROM 802, and RAM 803 are interconnected via a bus 804. An Input / Output (I / O) interface 805 is also connected to the bus 804.

[0130] The following components are connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to I / O interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 810 as needed so that computer programs read from it can be installed into storage section 808 as needed.

[0131] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 809, and / or installed from removable medium 811. When the computer program is executed by central processing unit (CPU) 801, it performs various functions defined in the system of this application.

[0132] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0134] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0135] Another aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the determination method as described above. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.

[0136] Another aspect of this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the determination methods provided in the various embodiments described above.

[0137] According to one aspect of the embodiments of this application, a computer system is also provided, including a Central Processing Unit (CPU), which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) or a program loaded from storage into random access memory (RAM), such as performing the methods described above. Various programs and data required for system operation are also stored in the RAM. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0138] The following components are connected to the I / O interface: input components including keyboards, mice, etc.; output components including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage components including hard drives; and communication components including network interface cards such as LAN (Local Area Network) cards and modems. The communication components perform communication processing via networks such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical discs, magneto-optical discs, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage components as required.

[0139] The above description is merely a preferred exemplary embodiment of this application and is not intended to limit the implementation of this application. Those skilled in the art can easily make corresponding modifications or alterations based on the main concept and spirit of this application. Therefore, the scope of protection of this application should be determined by the scope of protection claimed in the claims.< / string> < / string>

Claims

1. A method for determining target speech information, characterized in that, The determination method includes: In response to the first determination instruction, based on the updated first preset keywords, the target start keyword and the target end keyword in the initial speech text are determined; wherein, the target start keyword and the target end keyword are the first occurrence of the start keyword and the last occurrence of the end keyword in the initial speech text, respectively, and the updated first preset keyword is a word obtained by updating the original first preset keyword; Based on the target start keyword and the target end keyword, the target speech text is determined from the initial speech text; Based on the second preset keyword, the first target keyword in the target speech text and the type to which the first target keyword belongs are determined to determine the target speech information in the target speech text. This includes: taking words in the target speech text that match the second preset keyword as the first target keyword, and taking the preset type corresponding to the successfully matched second preset keyword as the type to which the first target keyword belongs; wherein, each second preset keyword corresponds to its own preset type; and taking the type to which the first target keyword belongs as the information header of the first target keyword to determine the target speech information in the target speech text.

2. The determination method according to claim 1, characterized in that, The determination method further includes: In response to the second determination instruction, based on the second preset keyword, the second target keyword in the initial speech text and the type to which the second target keyword belongs are determined, so as to determine the target speech information in the initial speech text; wherein, the number of the second target keyword is greater than or equal to the number of the first target keyword; the second determination instruction is an instruction that represents the unconditional triggering of the target speech information determination function.

3. The determination method according to claim 1, characterized in that, The determination method further includes: If the specified keyword in the target speech text fails to match all the second preset keywords, then the specified keyword is escaped based on the preset character escape function; If the escaped characters include characters representing time, then the specified keyword is determined to be the first target keyword of type time.

4. The determination method according to claim 1, characterized in that, Before determining the words in the target speech text that match the second preset keyword as the first target keyword, the determination method further includes: Iterate through each preset type and take the second preset keyword corresponding to the preset type as the target second preset keyword; wherein, each preset type has at least one second preset keyword, and the second preset keywords corresponding to each preset type include separator characters; Each word in the target speech text is matched with the target second preset keyword to obtain the words in the target speech text that match the target second preset keyword, thus obtaining the words in the target speech text that match the second preset keyword.

5. The determination method according to any one of claims 1 to 4, characterized in that, The step of determining the target speech text from the initial speech text based on the target start keyword and the target end keyword includes: The speech text following the target starting keyword in the initial speech text is taken as the first speech text, and the speech text before the target ending keyword in the initial speech text is taken as the second speech text. The overlapping speech text between the first speech text and the second speech text is identified as the target speech text.

6. The determining method according to any one of claims 1 to 4, characterized in that, The first preset keyword includes a preset start keyword and a preset end keyword; The step of determining the target start keyword and target end keyword in the initial speech text based on the updated first preset keyword includes: If the target word successfully matches the preset starting keyword, then the target word is determined as the starting keyword in the initial speech text, thereby identifying all starting keywords in the initial speech text; wherein, the target word is any word in the initial speech text; If the target word successfully matches the preset end keyword, then the target word is determined as the end keyword in the initial speech text, thereby identifying all end keywords in the initial speech text; The first occurrence of the starting keyword in the initial speech text is taken as the target starting keyword, and the last occurrence of the ending keyword in the initial speech text is taken as the target ending keyword.

7. A device for determining target speech information, characterized in that, The determining device includes: The first response module, in response to the first determination instruction, determines the target start keyword and the target end keyword in the initial speech text based on the updated first preset keyword; wherein, the target start keyword and the target end keyword are respectively the first occurrence of the start keyword and the last occurrence of the end keyword in the initial speech text, and the updated first preset keyword is a word obtained by updating the original first preset keyword; The text determination module determines the target speech text from the initial speech text based on the target start keyword and the target end keyword; The information determination module determines the first target keyword and the type of the first target keyword in the target speech text based on the second preset keyword, thereby determining the target speech information in the target speech text. This includes: taking words in the target speech text that match the second preset keyword as the first target keyword, and taking the preset type corresponding to the successfully matched second preset keyword as the type of the first target keyword; wherein each second preset keyword corresponds to its own preset type; and taking the type of the first target keyword as the information header of the first target keyword to determine the target speech information in the target speech text.

8. An electronic device, characterized in that, include: Controller; A memory for storing one or more programs, which, when executed by a controller, cause the controller to perform the determination method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, It stores computer-readable instructions that, when executed by the computer's processor, cause the computer to perform the determination method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Program endpoint time detection apparatus and method and program information retrieval system

    CN102073635A

  • Voice keyword screening method and device, travel terminal, equipment and medium

    CN111009240A