Semantic processing terminals, methods, storage media, and vehicles

By classifying and parsing user voice through a semantic processing terminal, and rationally arranging parallel and serial semantic execution, the crosstalk problem in voice command processing in the smart cockpit is solved, thus improving the user experience.

CN119360838BActive Publication Date: 2025-11-14CHINA FAW CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411277510.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2025-11-14
Estimated Expiration
2044-09-12

AI Technical Summary

Technical Problem

In smart cockpits, the processing of multiple voice commands under streaming semantics suffers from inconsistent processing times and unclear logic, leading to crosstalk between different business units during execution and a poor user experience.

Method used

The semantic processing terminal classifies user speech by using the semantic recognition module and intent analysis module, identifies major semantic categories and parses user intent, and controls the business execution to execute in parallel and serial semantic order to reduce crosstalk between semantics.

Benefits of technology

By rationally arranging parallel and serial semantic execution logic under streaming semantics, crosstalk between semantics is reduced, thus improving the user's smart cockpit experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360838B_ABST
    Figure CN119360838B_ABST
Patent Text Reader

Abstract

This application discloses a semantic processing terminal, method, storage medium, and vehicle, belonging to the field of semantic processing technology. The semantic processing terminal includes at least a semantic recognition module, an intent analysis module, and a processing control module. The semantic recognition module is at least used to respond to a user's voice input operation and determine the semantic category of the user's input voice. The intent analysis module is at least used to classify and parse multiple user intents contained in the user's input voice when the semantic category of the user's input voice is streaming semantics, to obtain the semantic execution type of each user intent. The processing control module is at least used to control at least one business execution party to successfully execute parallel semantics before executing serial semantics when the semantic execution type includes parallel semantics and serial semantics, thereby reducing crosstalk between semantics. This application can achieve accurate processing of user intents while reducing crosstalk between semantics, which is beneficial to improving the user's smart cockpit experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of semantic processing technology, and in particular to a semantic processing terminal, method, storage medium and vehicle. Background Technology

[0002] Currently, when users in the vehicle's smart cockpit control the vehicle via voice, there is a one-to-one correspondence between their voice input and the actions of the business parties. Even if the business parties need some time to process the voice commands, it will not interfere with the workflow of other business parties.

[0003] However, with the introduction of streaming semantic support in smart cockpits, users can input multiple voice commands simultaneously and send them to multiple business units for execution. As a result, due to the inconsistent processing time of voice commands by different business units and the unclear internal logic of user voice commands, different business units are prone to crosstalk during the execution of user voice commands, resulting in a poor user experience. Summary of the Invention

[0004] This invention provides a semantic processing terminal, method, storage medium, and vehicle to at least reduce semantic crosstalk and improve user experience.

[0005] In a first aspect, embodiments of the present invention provide a semantic processing terminal, which establishes a connection with at least one business execution party;

[0006] The semantic processing terminal includes at least a semantic recognition module, an intent analysis module, and a processing control module;

[0007] The semantic recognition module is at least used to respond to the user's voice input operation and determine the semantic category of the user's input voice; wherein, the semantic category includes at least ordinary semantics and streaming semantics;

[0008] The intent analysis module, connected between the semantic recognition module and the processing control module, is at least used to classify and parse multiple user intents contained in the user input speech when the semantic category of the user input speech is streaming semantics, so as to obtain the semantic execution type of each user intent; wherein, the semantic execution type includes at least parallel semantics and / or serial semantics.

[0009] The processing control module is connected to at least one of the business execution parties and is used at least to control at least one of the business execution parties to successfully execute the parallel semantics before executing the serial semantics when the semantic execution type includes the parallel semantics and the serial semantics, so as to reduce crosstalk between semantics.

[0010] Optionally, the processing control module is further configured to generate broadcast information based on the semantic execution results of all business executors after all the semantics corresponding to each of the business executors have been executed.

[0011] Optionally, when the semantic execution type only includes the parallel semantics, if the number of sub-semantics in the parallel semantics is not greater than a preset threshold, the processing control module is specifically used to generate fine broadcast information based on the semantic execution result of each sub-semantics in the parallel semantics;

[0012] The detailed broadcast information refers at least to a specific business carrier.

[0013] Optionally, when the semantic execution type only includes the parallel semantics, if the number of sub-semantics in the parallel semantics is greater than a preset threshold, the processing control module is specifically used to determine the size relationship between the number of sub-semantics that failed to execute and the number of sub-semantics that succeeded to execute based on the semantic execution result of each sub-semantics in the parallel semantics, and then generate the broadcast information based on the size relationship;

[0014] If the number of sub-semantics that failed to execute is less than the number of sub-semantics that succeeded to execute, then the broadcast information shall at least provide detailed broadcasts of the sub-semantics that failed to execute and provide summary broadcasts of the sub-semantics that succeeded to execute.

[0015] Optionally, when the semantic execution type includes multiple serial semantics, the processing control module is further configured to control at least one of the business execution parties to execute sequentially according to the parsing time order of the multiple serial semantics until all the multiple serial semantics are executed.

[0016] Optionally, the semantic execution type may further include at least termination semantics;

[0017] The processing control module is further configured to, when the semantic execution type includes the termination semantic, the parallel semantic, and / or the serial semantic, control at least one of the business execution parties to successfully execute the parallel semantic and / or the serial semantic before executing the termination semantic, so as to ensure that the execution of the user intent is maximized.

[0018] Optionally, the parallel semantics at least refers to semantics that can be directly executed by the business executor without secondary interaction with the user;

[0019] The serial semantics refer at least to semantics that require secondary interaction with the user before the business can be executed;

[0020] The term "termination semantics" at least refers to the semantics that prevents subsequent interactions after the business executor performs the semantics.

[0021] Secondly, embodiments of the present invention also provide a semantic processing method, wherein the method is executed using the semantic processing terminal described in the first aspect, the method comprising:

[0022] In response to the user's voice input, the semantic recognition module determines the semantic category of the user's voice input.

[0023] When the semantic category of the user input speech is streaming semantics, the intent analysis module categorizes and parses the multiple user intents contained in the user input speech to obtain the semantic execution type of each user intent.

[0024] When the semantic execution type includes parallel semantics and serial semantics, the serial semantics are executed only after at least one business execution party is successfully executed by the processing control module, so as to reduce interference between semantics.

[0025] The semantic categories include at least ordinary semantics and streaming semantics; the semantic execution types include at least parallel semantics and / or serial semantics.

[0026] Thirdly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the semantic processing method described in the second aspect.

[0027] Fourthly, embodiments of the present invention also provide a vehicle that integrates a semantic processing terminal as described in any of the first aspects.

[0028] This invention provides a semantic processing terminal, method, storage medium, and vehicle. First, in response to a user's voice input, a semantic recognition module determines the semantic category of the user's input voice. Then, when the semantic category of the user's input voice is streaming semantics, an intent analysis module categorizes and parses multiple user intents contained in the user's input voice to obtain the semantic execution type of each user intent. Finally, when the semantic execution type includes parallel semantics and serial semantics, a processing control module controls at least one business execution party to successfully execute parallel semantics before executing serial semantics, thereby reducing semantic interference. Thus, by introducing streaming semantics into the vehicle's smart cockpit, this invention can categorize and parse multiple user intents generated by streaming semantics in a short time. By decomposing the semantic execution type of each user intent and rationally arranging the execution logic between parallel and serial semantics, it achieves accurate processing of user intents while reducing crosstalk between semantics, thus improving the user's smart cockpit experience. Attached Figure Description

[0029] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0030] Figure 1 This is a schematic diagram of the structure of a semantic processing terminal provided in an embodiment of the present invention;

[0031] Figure 2 This is a flowchart of a semantic processing method provided in an embodiment of the present invention. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0033] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms “a,” “said,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms, and “multiple” generally includes at least two unless the context clearly indicates otherwise.

[0034] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0035] It should be understood that although the terms first, second, third, etc., may be used in the embodiments of this application, these descriptions should not be limited to these terms. These terms are only used to distinguish the descriptions. For example, first may also be referred to as second without departing from the scope of the embodiments of this application, and similarly, second may also be referred to as first.

[0036] Depending on the context, the words “if” or “suppose” as used here can be interpreted as “when” or “in response to determination” or “in response to detection.” Similarly, depending on the context, the phrases “if determination” or “if detection (of the stated condition or event)” can be interpreted as “when determination” or “in response to determination” or “when detection (of the stated condition or event)” or “in response to detection (of the stated condition or event).”

[0037] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that an article or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such an article or device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or device that includes said element.

[0038] It should be noted that any symbols and / or numbers present in the specification that are not marked in the accompanying drawings are not reference numerals.

[0039] Figure 1 This is a schematic diagram of the structure of a semantic processing terminal provided in an embodiment of the present invention. This embodiment is at least applicable to voice interaction scenarios between users in vehicles and smart cockpits, and the terminal can be implemented using software and / or hardware methods. Figure 1 As shown, the semantic processing terminal provided in this embodiment establishes a connection with at least one business execution party.

[0040] The semantic processing terminal includes at least a semantic recognition module 110, an intent analysis module 120, and a processing control module 130.

[0041] The semantic recognition module 110 is at least used to respond to the user's voice input operation and determine the semantic category of the user's input voice; wherein the semantic category includes at least ordinary semantics and streaming semantics.

[0042] The intent analysis module 120 is connected between the semantic recognition module 110 and the processing control module 130. It is used at least to classify and parse multiple user intents contained in the user input speech when the semantic category of the user input speech is streaming semantics, so as to obtain the semantic execution type of each user intent; wherein the semantic execution type includes at least parallel semantics and / or serial semantics.

[0043] The processing control module 130 is connected to at least one business executor and is used at least to control at least one business executor to execute the serial semantics after successfully executing the parallel semantics when the semantic execution type includes parallel semantics and serial semantics, so as to reduce crosstalk between semantics.

[0044] The semantic processing terminal can be integrated into the vehicle's cockpit space center (CSC), for example. Additionally, the service executor can be any functional module or device within the vehicle's cockpit, such as navigation, vehicle control, or entertainment media.

[0045] It is understood that a user's voice input operation can refer to the operation of controlling at least one business executor to perform a corresponding functional action through voice expression. For example, the user's voice input could be "turn off the seat heating," "open the car window," "turn on the music," "call Zhang San," or "turn on the air conditioning." It is also understood that ordinary semantics can be semantics involved in the background technology that have a one-to-one correspondence with the functional actions of the business executor; that is, one ordinary semantic corresponds to one functional action performed by one business executor. For example, the aforementioned "turn off the seat heating" corresponds to the seat cutting off the power to the heating resistance wire inside the seat. Conversely, streaming semantics simultaneously involves multiple functional actions from multiple business executors, such as the aforementioned "open the car window," "turn on the music," "call Zhang San," or "turn on the air conditioning," which will not be elaborated further.

[0046] In one specific implementation, optionally, parallel semantics refers at least to semantics that can be executed directly by the business executor without secondary interaction with the user; serial semantics refers at least to semantics that require secondary interaction with the user before the business executor executes. Specifically, parallel semantics corresponds at least to "EXECUTE" type user intents, which are executed directly by the business executor, such as "open the car window." Serial semantics corresponds at least to "INQUIRY" type user intents, "SELECT" type user intents, and "UNKNOWN" type user intents. "INQUIRY" type user intents represent intents that require multiple rounds of user interaction to satisfy the user's needs, such as "set a custom wake word" via voice input. In this case, the user needs to continue inputting content about the custom wake word to realize the user intent. "SELECT" type user intents represent intents with multiple results to choose from, requiring multiple rounds of user interaction to select one result before execution can continue. For example, if a user inputs "navigate to Xizhimen" via voice, this semantic is sent to the navigation business executor, who will then query multiple locations related to "Xizhimen" and need to select one. The user specifically selects the navigation destination; the "UNKNOWN" type user intent indicates an intent whose business execution direction is unclear and requires a response from the business execution party for confirmation. For example, if the user inputs "call Zhang San" via voice, the semantics corresponding to this voice is sent to the telephone business execution party. In the first case, the telephone business execution party finds multiple Zhang Sans in the phone book and needs the user to select which Zhang San to call, which belongs to the aforementioned "SELECT" type user intent. In the second case, there is only one Zhang San in the phone book. In this case, the telephone business execution party can directly dial Zhang San's phone number, which belongs to the aforementioned "EXECUTE" type user intent.

[0047] Based on this, the working principle of the intent analysis module 120 and the processing control module 130 will be explained using the user input voice – “open the car window, turn on the music, call Zhang San, turn on the air conditioner” – as an example. The intent analysis module 120 categorizes and analyzes the user intent contained in the above user input voice to obtain the following four specific intents: “open the car window,” “turn on the music,” “call Zhang San,” and “turn on the air conditioner.” Clearly, “open the car window,” “turn on the music,” and “turn on the air conditioner” are all “EXECUTE” type user intents, corresponding to parallel semantics, while “call Zhang San” is the “UNKNOWN” type user intent, corresponding to serial semantics. Adaptively, the processing control module 130 can first complete the parallel control of the car window, music, and air conditioner, and then execute the serial semantics of “call Zhang San” after the above parallel control is completed.

[0048] It is understandable that parallel semantics can be directly executed by the corresponding business executor, resulting in high execution efficiency, while serial semantics involves multiple rounds of interaction between the user and the corresponding business executor, leading to slower execution speed. Therefore, if the processing control module 130 simultaneously controls the actions of the business executors corresponding to both parallel and serial semantics, while one business executor is executing serial semantics and further communicating with the user, another business executor has already completed the parallel semantics and needs to report the execution result to the user. This not only causes crosstalk between semantic execution processes but also makes it difficult for the user to determine which specific voice input the feedback from each business executor is referring to, or to ignore the execution result from the business executor responsible for processing parallel semantics, causing confusion for the user. If the processing control module 130 executes parallel semantics after completing serial semantics, it will slow down the entire semantic processing flow, resulting in low efficiency. In contrast, the technical solution of prioritizing the execution of parallel semantics and then processing serial semantics in this embodiment can reduce crosstalk between semantics while ensuring the working efficiency of the semantic processing terminal, thus improving the user's smart cockpit experience.

[0049] In summary, this embodiment provides a semantic processing terminal. First, in response to the user's voice input, the semantic recognition module determines the semantic category of the user's input voice. Then, when the semantic category of the user's input voice is streaming semantics, the intent analysis module categorizes and parses the multiple user intents contained in the user's input voice to obtain the semantic execution type of each user intent. Finally, when the semantic execution type includes parallel semantics and serial semantics, the processing control module controls at least one business execution party to successfully execute parallel semantics before executing serial semantics, thereby reducing semantic interference. Therefore, after introducing streaming semantics into the vehicle's intelligent cockpit, this embodiment can categorize and parse multiple user intents generated by streaming semantics in a short time. By decomposing the semantic execution type of each user intent and rationally arranging the execution logic between parallel and serial semantics, it achieves efficient and accurate processing of user intents while reducing crosstalk between semantics, thus improving the user's intelligent cockpit experience.

[0050] It should be noted that the technical solutions provided in the embodiments of the present invention can also be applied to voice interaction scenarios between users and smart devices in other fields, such as the control scenarios of various home appliances in response to user voice instructions in the field of smart homes, etc., which will not be elaborated further.

[0051] Based on the above embodiments or implementation methods, please continue to refer to Figure 1The following describes the broadcast information generation function of the processing control module 130, but this does not constitute a limitation on the present invention. Optionally, the processing control module 130 is also used to generate broadcast information based on the semantic execution results of all business executors after all the semantics corresponding to each business executor have been executed. The semantic execution results may include whether the corresponding business executor successfully executed the corresponding semantics or failed to execute the corresponding semantics. Furthermore, the broadcast information can inform the user of the semantic execution results of all business executors through voice, pop-up display, or other means.

[0052] However, after careful research, the inventors discovered that existing vehicle smart cockpits, when broadcasting the parallel semantic execution results of various business executors via voice, sometimes lack a business carrier, which can easily confuse users, while other times they mechanically broadcast repetitive and lengthy results, resulting in a poor user listening experience. Therefore, in a scenario where the semantic execution type of each user intent obtained by the intent analysis module 120 is parallel semantic, this embodiment configures the functionality of the processing control module 130 in two different ways.

[0053] In one specific implementation, optionally, when the semantic execution type only includes parallel semantics, if the number of sub-semantics in the parallel semantics is not greater than a preset threshold, the processing control module 130 is specifically used to generate fine broadcast information based on the semantic execution result of each sub-semantics in the parallel semantics; wherein, the fine broadcast information at least points to a specific business carrier.

[0054] The following example illustrates the sub-semantics using the user's voice input: "Open the car window, turn on the music, call Zhang San." In this user input, "open the car window" and "turn on the music" represent the aforementioned "EXECUTE" type user intent and correspond to parallel semantics. Therefore, it can be determined that there are two sub-semantics within this parallel semantics: "open the car window" and "turn on the music." It is understood that the preset threshold can be adaptively adjusted according to the user's personalized listening needs. This embodiment of the invention does not limit the number of thresholds; for example, it can be two.

[0055] Specifically, when the processing control module 130 detects the end of the user's voice input (e.g., no voice input is received from the user within a set time), the semantic processing terminal can determine that the user's voice input has stopped. At this time, the semantic processing terminal can send parallel semantics to each business execution party. The corresponding business execution party executes the sub-semantics in the parallel semantics and feeds back the sub-semantics execution results to the semantic processing terminal. After receiving the execution results of each sub-semantics, the semantic processing terminal can give a precise reply to the sub-semantics execution results according to the pre-set text-to-speech (TTS) content, such as "The car window is open for you, and music is playing for you" (i.e., the aforementioned detailed broadcast information, where "car window" and "music" are the aforementioned specific business carriers). It can be seen that the TTS content replied by the technical solution of this embodiment will point to the specific business of the sub-semantics, rather than giving a general reply (e.g., "opened," "closed," etc., which do not have a business carrier), which can avoid conceptual confusion in business replies, facilitate user understanding, and optimize the user experience.

[0056] In another specific implementation, optionally, when the semantic execution type only includes parallel semantics, if the number of sub-semantics in the parallel semantics is greater than a preset threshold, the processing control module 130 is specifically used to determine the size relationship between the number of failed sub-semantics and the number of successful sub-semantics based on the semantic execution result of each sub-semantics in the parallel semantics, and then generate broadcast information based on the size relationship; wherein, if the number of failed sub-semantics is less than the number of successful sub-semantics, the broadcast information at least provides detailed broadcasting of the failed sub-semantics and provides summary broadcasting of the successful sub-semantics.

[0057] When the number of sub-semantics in parallel semantics exceeds a preset threshold, it indicates that continuing to refine the broadcast of the sub-semantic execution results will make the broadcast content lengthy and repetitive, and at this time it is necessary to simplify the broadcast content.

[0058] Specifically, when parallel semantics involves multiple sub-semantics, the semantic processing terminal, after receiving the response from the first business executor, may not immediately report the corresponding business execution result to the user. Instead, it may wait for all other business executors to report their results before summarizing and broadcasting the result. Generally, most business executors successfully execute the user's intent for the sub-semantics (i.e., the number of failed sub-semantics is less than the number of successful sub-semantics). Therefore, the semantic processing terminal can only broadcast the specific business that failed. For example, if the semantic processing terminal recognizes four parallel sub-semantics based on the user's voice input, and after the corresponding business executors report their execution results, three sub-semantics succeed and one fails, the semantic processing terminal can broadcast to the user: "xxx failed, the rest are done!" Clearly, providing this summary report to the user reduces repetitive, lengthy, and similar-intent-based broadcasts, improving the conciseness and intelligence of the semantic execution result broadcast.

[0059] Understandably, in another specific implementation, if the number of sub-semantics that failed to execute is greater than the number of sub-semantics that succeeded to execute, the broadcast information will at least provide detailed broadcasts of the successfully executed sub-semantics and provide general broadcasts of the failed sub-semantics; or, if the number of sub-semantics that failed to execute is equal to the number of successfully executed sub-semantics, the broadcast information will at least provide detailed broadcasts of the successfully executed sub-semantics and provide general broadcasts of the failed sub-semantics; or, if the number of sub-semantics that failed to execute is equal to the number of successfully executed sub-semantics, the broadcast information will at least provide detailed broadcasts of the failed sub-semantics and provide general broadcasts of the successful sub-semantics.

[0060] In summary, by introducing streaming semantics into the vehicle's intelligent cockpit, this embodiment can categorize and parse multiple user intents generated by streaming semantics in a short period of time. By decomposing the semantic execution type of each user intent and rationally arranging the execution logic between parallel and serial semantics, it achieves efficient and accurate processing of user intents while reducing crosstalk between semantics, thus improving the user's intelligent cockpit experience. Furthermore, this embodiment can adaptively generate broadcast information based on the semantic execution type, the relationship between the number of sub-semantics in parallel semantics and a preset threshold, and the relationship between the number of failed and successful sub-semantics. For specific scenarios, this can avoid conceptual confusion in business responses, making it easier for users to understand and optimizing the user experience. It can also reduce repetitive, lengthy, and similar-intent-based broadcast content, improving the conciseness and intelligence of semantic execution result broadcasts.

[0061] It should be noted that, in one specific implementation, when the semantic recognition module determines that the speech category of the user's input speech is ordinary semantic (similar to a sub-semantic in parallel semantic), the processing control module can directly control the corresponding business execution party to execute the ordinary semantic, which will not be elaborated further.

[0062] Based on the above embodiments or implementation methods, please continue to refer to Figure 1 The following describes the functional configuration of the processing control module 130 in scenarios where the semantic execution type is only serial semantics, and where the semantic execution type includes parallel semantics, serial semantics, and / or termination semantics, but this does not constitute a limitation on the present invention.

[0063] In one specific implementation, optionally, when the semantic execution type includes multiple serial semantics, the processing control module 130 is further configured to control at least one business execution party to execute one by one in the parsing time order of the multiple serial semantics until all the multiple serial semantics are executed.

[0064] The parsing time order of serial semantics can refer to the order in which the intent analysis module 120 determines the serial semantic type. Specifically, the processing control module 130 prioritizes controlling the business executor corresponding to the first serial semantic determined by the intent analysis module 120 to execute the semantic function, then controls the business executor corresponding to the second serial semantic determined by the intent analysis module 120 to execute the semantic function, and so on, until all serial semantics have been executed by their corresponding business executors.

[0065] In another specific implementation, optionally, the semantic execution type also includes at least termination semantics; the processing control module 130 is further configured to control at least one business executor to successfully execute parallel semantics and / or serial semantics before executing termination semantics when the semantic execution type includes termination semantics, as well as parallel semantics and / or serial semantics, so as to ensure that the execution of user intent is maximized.

[0066] In this context, "termination semantics" can refer to a semantic execution type that can interrupt the interaction process between the semantic processing terminal and at least one business executor. Optionally, termination semantics at least refers to semantics that prevents subsequent interaction after the business executor executes the semantics. For example, a user can input "turn off the screen" via voice. After this semantic is executed by the business executor (i.e., the screen), the current screen will be turned off, and the screen cannot be interacted with again until the user manually or by inputting voice to turn it back on.

[0067] It is understandable that when the semantic execution type includes both termination semantics and parallel semantics, the processing control module 130 is specifically used to control at least one business executor to successfully execute parallel semantics before executing termination semantics, in order to ensure that the execution of user intent is maximized; or, when the semantic execution type includes both termination semantics and serial semantics, the processing control module 130 is specifically used to control at least one business executor to successfully execute serial semantics before executing termination semantics, in order to ensure that the execution of user intent is maximized; when the semantic execution type includes termination semantics, as well as parallel semantics and serial semantics, the processing control module 130 is specifically used to control at least one business executor to successfully execute parallel semantics before executing serial semantics (here, the semantic processing logic of the processing control module 130 when parallel semantics and serial semantics coexist is used, which can reduce crosstalk between semantics while ensuring the working efficiency of the semantic processing terminal, and is conducive to improving the user's smart cockpit experience), and finally execute termination semantics to ensure that the execution of user intent is maximized.

[0068] For example, the workflow of the semantic processing terminal is described below based on a real-world interaction scenario.

[0069] 1. The user inputs the voice: "Open the car window, turn on the music, call Zhang San, turn on the air conditioner." After receiving the aforementioned user input voice, the semantic recognition module 110 determines that its semantic category is streaming semantics.

[0070] 2. The intent analysis module 120 will segment the user's voice input into multiple user intents: "Open the car window", "Turn on the music", "Call Zhang San", "Turn off the screen", "Turn on the air conditioner".

[0071] 3. The intent analysis module 120, based on these four user intents, classifies them according to classification rules and confirms their semantic execution types.

[0072] Parallel semantics: "Open the car window", "Turn on the music", "Turn on the air conditioner";

[0073] Sequential semantics: "Call Zhang San";

[0074] End semantics: "Turn off the screen".

[0075] 4. The processing control module 130 first distributes the parallel semantics to the corresponding business execution parties. After the execution results of all parallel semantics are returned, it summarizes and broadcasts the results. For example, if the execution of "open music" fails, but all other parallel semantics succeed, it will broadcast to the user: "The music failed to open, but everything else is done!"; if all executions succeed, it will broadcast to the user: "Everything is done!"

[0076] 5. After the parallel semantic execution is completed, the processing control module 130 sends the serial semantic data to the telephone service executor. The telephone service executor will query the number according to the service execution process. If there are multiple numbers, it will return a message to the semantic processing terminal, which will prompt the user to select the number to use. If there is only one number, it will execute the dialing function.

[0077] 6. After the serial semantics are executed, the screen will be turned off.

[0078] 7. This user input has been executed.

[0079] In summary, by introducing streaming semantics into the vehicle's intelligent cockpit, this embodiment can categorize and parse multiple user intents generated by streaming semantics in a short period of time. By decomposing the semantic execution type of each user intent and rationally arranging the execution logic between parallel and serial semantics, it achieves efficient and accurate processing of user intents while reducing crosstalk between semantics, thus improving the user's intelligent cockpit experience. Furthermore, this embodiment can adaptively generate broadcast information based on the relationship between the semantic execution type, the number of sub-semantics in parallel semantics and a preset threshold, and the relationship between the number of failed and successful sub-semantics. For specific scenarios, this avoids conceptual confusion in business responses, facilitating user understanding and optimizing the user experience. It also reduces repetitive, lengthy, and similar-intent-based broadcast content, improving the conciseness and intelligence of semantic execution result broadcasts. In addition, when the semantic execution type includes an ending semantic, this embodiment prioritizes the execution of other semantic execution types, ensuring maximum execution of user intents.

[0080] This invention also provides a semantic processing method. Figure 2 This is a flowchart of a semantic processing method provided by an embodiment of the present invention. This embodiment is at least applicable to voice interaction scenarios between users in vehicles and smart cockpits. The semantic processing method can be, but is not limited to, executed by the semantic processing terminal in this embodiment as the execution subject, which can be implemented in software and / or hardware. Figure 2 As shown, this semantic processing method includes at least the following steps:

[0081] S1. In response to the user's voice input operation, the semantic recognition module determines the semantic category of the user's voice input.

[0082] S2. When the semantic category of the user's input speech is streaming semantics, the intent analysis module categorizes and parses the multiple user intents contained in the user's input speech to obtain the semantic execution type of each user intent.

[0083] S3. When the semantic execution type includes parallel semantics and serial semantics, the processing control module controls at least one business execution party to successfully execute the parallel semantics before executing the serial semantics, so as to reduce interference between semantics.

[0084] Among them, the semantic categories include at least ordinary semantics and streaming semantics; the semantic execution types include at least parallel semantics and / or serial semantics.

[0085] Optionally, it also includes:

[0086] S4. After all the semantics corresponding to each business executor have been executed, the processing control module generates broadcast information based on the semantic execution results of all business executors.

[0087] Optionally, step S4 includes at least:

[0088] When the semantic execution type only includes parallel semantics, if the number of sub-semantics in the parallel semantics is not greater than a preset threshold, the processing control module generates fine broadcast information based on the semantic execution result of each sub-semantics in the parallel semantics.

[0089] Among them, detailed information broadcasting refers at least to specific business entities.

[0090] Optionally, step S4 includes at least:

[0091] When the semantic execution type only includes parallel semantics, if the number of sub-semantics in the parallel semantics is greater than a preset threshold, the processing control module determines the size relationship between the number of sub-semantics that failed to execute and the number of sub-semantics that succeeded to execute based on the semantic execution result of each sub-semantics in the parallel semantics, and then generates broadcast information based on the size relationship.

[0092] If the number of sub-semantics that fail to execute is less than the number of sub-semantics that succeed, then the broadcast information will at least provide detailed broadcasts of the sub-semantics that fail to execute and provide summary broadcasts of the sub-semantics that succeed.

[0093] Optionally, it also includes:

[0094] S5. When the semantic execution type includes multiple serial semantics, the processing control module controls at least one business execution party to execute them one by one in the parsing time order of the multiple serial semantics until all the multiple serial semantics are executed.

[0095] Optionally, the semantic execution type may also include at least termination semantics;

[0096] Semantic processing methods also include:

[0097] S6. When the semantic execution type includes termination semantics, as well as parallel semantics and / or serial semantics, the termination semantics are executed only after at least one business executor is successfully executed by the processing control module, so as to ensure that the execution of the user's intent is maximized.

[0098] Optionally, parallel semantics at least refers to semantics that can be executed directly by the business executor without requiring secondary interaction with the user;

[0099] Serial semantics refers at least to semantics that require secondary interaction with the user before business operations can be executed;

[0100] Termination semantics at least refer to semantics that prevent subsequent interactions from taking place after the business executor has executed the semantics.

[0101] This embodiment provides a semantic processing method. First, in response to a user's voice input operation, the semantic recognition module determines the semantic category of the user's input voice. Then, when the semantic category of the user's input voice is streaming semantics, the intent analysis module categorizes and parses the multiple user intents contained in the user's input voice to obtain the semantic execution type of each user intent. Finally, when the semantic execution type includes parallel semantics and serial semantics, the processing control module controls at least one business execution party to successfully execute parallel semantics before executing serial semantics, thereby reducing semantic interference. Therefore, after introducing streaming semantics into the vehicle's intelligent cockpit, this embodiment can categorize and parse multiple user intents generated by streaming semantics in a short time. By decomposing the semantic execution type of each user intent and rationally arranging the execution logic between parallel and serial semantics, it achieves accurate processing of user intents while reducing crosstalk between semantics, thus improving the user's intelligent cockpit experience.

[0102] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the semantic processing method provided in all embodiments of this application: in response to a user's voice input operation, a semantic recognition module determines the semantic category of the user's input voice; when the semantic category of the user's input voice is streaming semantics, an intent analysis module categorizes and parses multiple user intents contained in the user's input voice to obtain the semantic execution type of each user intent; when the semantic execution type includes parallel semantics and serial semantics, a processing control module controls at least one business execution party to successfully execute parallel semantics before executing serial semantics, thereby reducing semantic interference.

[0103] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0104] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including—but not limited to—electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0105] The program code contained on a computer-readable medium may be transmitted using any suitable medium, including—but not limited to—wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0106] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0107] This invention also provides a vehicle that integrates a semantic processing terminal as provided in all embodiments of this application.

[0108] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A semantic processing terminal, characterized in that, The semantic processing terminal establishes a connection with at least one business executor. The semantic processing terminal includes at least a semantic recognition module, an intent analysis module, and a processing control module; The semantic recognition module is at least used to respond to the user's voice input operation and determine the semantic category of the user's input voice; wherein, the semantic category includes at least ordinary semantics and streaming semantics; The intent analysis module, connected between the semantic recognition module and the processing control module, is at least used to classify and parse multiple user intents contained in the user input speech when the semantic category of the user input speech is streaming semantics, so as to obtain the semantic execution type of each user intent; wherein, the semantic execution type includes at least parallel semantics and / or serial semantics. The processing control module is connected to at least one of the business execution parties and is used at least to control at least one of the business execution parties to successfully execute the parallel semantics before executing the serial semantics when the semantic execution type includes the parallel semantics and the serial semantics, so as to reduce crosstalk between semantics.

2. The semantic processing terminal according to claim 1, characterized in that, The processing control module is also used to generate broadcast information based on the semantic execution results of all business executors after all the semantics corresponding to each of the business executors have been executed.

3. The semantic processing terminal according to claim 2, characterized in that, When the semantic execution type only includes the parallel semantics, if the number of sub-semantics in the parallel semantics is not greater than a preset threshold, the processing control module is specifically used to generate fine broadcast information based on the semantic execution result of each sub-semantics in the parallel semantics; The detailed broadcast information refers at least to a specific business carrier.

4. The semantic processing terminal according to claim 2, characterized in that, When the semantic execution type only includes the parallel semantics, if the number of sub-semantics in the parallel semantics is greater than a preset threshold, the processing control module is specifically used to determine the size relationship between the number of sub-semantics that failed to execute and the number of sub-semantics that succeeded to execute based on the semantic execution result of each sub-semantics in the parallel semantics, and then generate the broadcast information based on the size relationship. If the number of sub-semantics that failed to execute is less than the number of sub-semantics that succeeded to execute, then the broadcast information shall at least provide detailed broadcasts of the sub-semantics that failed to execute and provide summary broadcasts of the sub-semantics that succeeded to execute.

5. The semantic processing terminal according to claim 1, characterized in that, When the semantic execution type includes multiple serial semantics, the processing control module is further configured to control at least one of the business execution parties to execute sequentially according to the parsing time order of the multiple serial semantics until all the multiple serial semantics are executed.

6. The semantic processing terminal according to claim 1, characterized in that, The semantic execution type also includes at least termination semantics; The processing control module is further configured to, when the semantic execution type includes the termination semantic, the parallel semantic, and / or the serial semantic, control at least one of the business execution parties to successfully execute the parallel semantic and / or the serial semantic before executing the termination semantic, so as to ensure that the execution of the user intent is maximized.

7. The semantic processing terminal according to claim 6, characterized in that, The parallel semantics refer at least to the semantics that the business executor can directly execute without secondary interaction with the user; The serial semantics refer at least to semantics that require secondary interaction with the user before the business can be executed; The term "end semantics" at least refers to the semantics that prevents subsequent interaction after the business executor performs the semantics.

8. A semantic processing method, characterized in that, The method is executed using the semantic processing terminal of claim 1, the method comprising: In response to the user's voice input, the semantic recognition module determines the semantic category of the user's voice input. When the semantic category of the user input speech is streaming semantics, the intent analysis module categorizes and parses the multiple user intents contained in the user input speech to obtain the semantic execution type of each user intent. When the semantic execution type includes parallel semantics and serial semantics, the serial semantics are executed only after at least one business execution party is successfully executed by the processing control module, so as to reduce interference between semantics. The semantic categories include at least ordinary semantics and streaming semantics; the semantic execution types include at least parallel semantics and / or serial semantics.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the semantic processing method of claim 8.

10. A vehicle, characterized in that, It integrates a semantic processing terminal as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Voice interaction method and device, electronic equipment and storage medium

    CN114171016A

  • Control method, device and system, vehicle and storage medium

    CN115762506A