Intelligent cockpit voice processing method and device, cockpit, automobile and storage medium

By building a matching tree and filtering oralized words, the problem of inaccurate matching of voice data in the smart cockpit is solved, and a higher matching success rate and speech interaction accuracy is achieved.

CN120236580APending Publication Date: 2025-07-01GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510581760.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

When facing the voice data with spoken words issued by the user, the existing smart cockpit voice processing method has a low success rate of matching the recall screen elements and cannot accurately hit the user's expected results.

Method used

By building a matching tree, the voice data sent by the user is obtained for initial matching with the scene data, the spoken words are filtered out, and the mask marks are used, and the target matching results are further matched with the initial node information.

Benefits of technology

The matching accuracy between voice data and scene data is improved, and the results of users' expected hits can be matched more accurately, which enhances the voice interaction capabilities of the smart cockpit.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236580A_ABST
    Figure CN120236580A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent cockpit voice processing method and device, a cockpit, an automobile, a storage medium and a program product. The method comprises the following steps: acquiring voice data sent by a user and scene data of a cabin screen; matching the voice data sent by the user with node information of nodes in a matching tree generated according to the scene data to obtain an initial node information matching result corresponding to the voice data sent by the user; filtering the voice data sent by the user according to a preset mode to obtain filtered voice data; and matching the filtered voice data with the initial node information matching result to obtain a target matching result corresponding to the filtered voice data. According to the method, the expected hit result of the user can be more accurately matched, and the success rate of matching the expected hit result for the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving technology, and in particular to an intelligent cockpit voice processing method, device, cockpit, vehicle, storage medium and program product. Background Art

[0002] With the continuous development of autonomous driving technology and intelligent vehicle networking technology, the functions provided by intelligent cockpits in intelligent vehicles are becoming more and more abundant. In the intelligent cockpit of an intelligent vehicle, a visible-and-speakable function can be provided to assist users in interacting with the screen, and users can operate intelligent cockpit-related devices through voice.

[0003] The visible-and-speakable function provided by the intelligent cockpit can adopt the NLU (Natural Language Understanding) algorithm. The NLU algorithm obtains the scenario data uploaded by the client, and matches the label of the node in the scenario data with the voice data issued by the user, so as to obtain the expected hit result, and further perform corresponding interaction processing according to the hit result. The intelligent cockpit voice processing method provided by the related technology can only limit to matching the similarity between the voice data issued by the user and the label in the scenario data. However, the voice data issued by the user may be an irregular voice command, such as containing some colloquial words. The existence of these colloquial words affects the similarity of the matching between the voice data issued by the user and the label in the scenario data, brings certain interference to successfully matching and recalling elements on the screen, and reduces the hit success rate of the hit result. For example Figure 1 in, there is a song list element on the screen of the in-vehicle large screen of the intelligent cockpit. When the user issues voice data of "Play Chrysanthemum Terrace", the song "Chrysanthemum Terrace" on the screen can be hit through the related technology; when the user issues voice data of "Please help me play Chrysanthemum Terrace", since the voice data issued by the user contains more colloquial words such as modal particles, the overall matching degree of the voice data and the label is reduced, and the element that the user expects to hit cannot be hit, resulting in the failure of matching and recall.

[0004] Therefore, the intelligent cockpit voice processing method in the related technology needs to further improve the success rate of matching the expected hit result for the user. Summary of the Invention

[0005] To solve or partially solve the problems existing in the related technology, this application provides an intelligent cockpit voice processing method, device, cockpit, vehicle, storage medium and program product, which can more accurately match the expected hit result of the user and improve the success rate of matching the expected hit result for the user.

[0006] The first aspect of the present application provides an intelligent cockpit voice processing method, including: Obtain the voice data issued by the user and the scene data of the cockpit screen; Match the voice data issued by the user with the node information of the nodes in the matching tree generated according to the scene data to obtain the initial node information matching result corresponding to the voice data issued by the user; Filter the voice data issued by the user in a preset manner to obtain filtered voice data; Match the filtered voice data with the initial node information matching result to obtain the target matching result corresponding to the filtered voice data.

[0007] In one embodiment, the filtering the voice data issued by the user in a preset manner to obtain filtered voice data includes: Obtain the unmatched content that is not matched during the matching process of the voice data issued by the user and the node information of the nodes in the matching tree; Match the unmatched content with the preset words in the preset word library, and mask the matched content; Filter the masked content to obtain filtered voice data.

[0008] In one embodiment, the matching the unmatched content with the preset words in the preset word library and masking the matched content includes: Match the unmatched content at the preset position of the voice data issued by the user with the preset words in the preset word library; Mask the matched content; Wherein the preset position includes at least one of the following positions: prefix position, suffix position, and middle position.

[0009] In one embodiment, the preset words include colloquial words.

[0010] In one embodiment, the matching the voice data issued by the user with the node information of the nodes in the matching tree generated according to the scene data to obtain the initial node information matching result corresponding to the voice data issued by the user includes: Traverse the matching tree generated according to the scene data, match the voice data issued by the user with the node information of the non-root nodes and the parent nodes of the non-root nodes in the matching tree, and obtain the first candidate matching results marked by each node during the traversal process; Perform a first similarity calculation on the first candidate matching result, and filter out the initial node information matching result corresponding to the voice data sent by the user according to the first matching similarity of each node.

[0011] In one embodiment, the matching of the filtered voice data with the initial node information matching result to obtain the target matching result corresponding to the filtered voice data includes: Traverse and match the filtered voice data with the initial node information matching result located at the first preset sorting position, and obtain the second candidate matching results marked by each node during the traversal; Perform a second similarity calculation on the second candidate matching result, and filter out the target matching result corresponding to the filtered voice data according to the second matching similarity of each node.

[0012] In one embodiment, the filtering out the target matching result corresponding to the filtered voice data according to the second matching similarity of each node includes: Filter out the second candidate matching result of the node corresponding to the second matching similarity located at the second preset sorting position from the second matching similarities of each node exceeding the preset threshold as the target matching result corresponding to the filtered voice data.

[0013] The second aspect of the present application provides an intelligent cockpit voice processing device, including: A data acquisition module, configured to acquire voice data sent by a user and scene data of a cockpit screen; A first matching module, configured to match the voice data sent by the user with the node information of the nodes in the matching tree generated according to the scene data, and obtain the initial node information matching result corresponding to the voice data sent by the user; A filtering processing module, configured to filter the voice data sent by the user in a preset manner to obtain filtered voice data; A second matching module, configured to match the filtered voice data with the initial node information matching result to obtain the target matching result corresponding to the filtered voice data.

[0014] In one embodiment, the filtering processing module includes: An unmatched content acquisition sub-module, configured to acquire the unmatched content that is not matched during the matching process of the voice data sent by the user with the node information of the nodes in the matching tree; A mask processing sub-module, configured to match the unmatched content with preset words in a preset word library, and mask the matched content; A filtering sub-module for filtering the content marked by the mask to obtain filtered voice data.

[0015] The third aspect of this application provides an intelligent cockpit, including the intelligent cockpit voice processing device as described above.

[0016] The fourth aspect of this application provides an intelligent vehicle, including: A processor; and A memory storing executable code, which, when executed by the processor, causes the processor to execute the method as described above.

[0017] The fifth aspect of this application provides a computer-readable storage medium storing executable code, which, when executed by a processor of an electronic device, causes the processor to execute the method as described above.

[0018] The sixth aspect of this application provides a computer program product, which includes computer instructions that, when executed by a processor, implement the method as described above.

[0019] The technical solution provided by this application may include the following beneficial effects: The technical solution of this application can match the voice data sent by the user with the node information of the nodes in the matching tree generated according to the scenario data to obtain the initial node information matching result corresponding to the voice data sent by the user; filter the voice data sent by the user in a preset manner to obtain filtered voice data; match the filtered voice data with the initial node information matching result to obtain the target matching result corresponding to the filtered voice data. Through the above processing, when processing the scenario data information of the client, based on the first match, for example, the initial node information matching result, the voice data sent by the user can be further filtered in a preset manner to reduce the interference of some voice words, and then the filtered voice data is matched with the initial node information matching result again to obtain the target matching result, which can improve the matching accuracy, more accurately match the result that the user expects to hit, and improve the success rate of matching the result that the user expects to hit.

[0020] Further, the technical solution of the present application can obtain the unmatched content that has not been matched during the matching process between the voice data sent by the user and the node information of the nodes in the matching tree; match the unmatched content with the preset words in the preset word library, and mask the matched content; filter the masked content to obtain the filtered voice data. The preset words may include colloquial words, etc. Through the above processing of colloquial words, etc., the present application can achieve compatibility with voice data with colloquial words, improve the generalization ability of understanding the voice data (voice input) sent by the user, so that even when there are many colloquial words in the voice data sent by the user, the result that the user expects to hit can be accurately matched, and the success rate of matching the result that the user expects to hit for the user is improved.

[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] By describing the exemplary embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features and advantages of the present application will become more obvious, wherein, in the exemplary embodiments of the present application, the same reference numerals generally represent the same components.

[0023] Figure 1 is an exemplary schematic diagram of the scene data of the screen in the intelligent cockpit shown in the present application; Figure 2 is the first process schematic diagram of the intelligent cockpit voice processing method shown in the present application; Figure 3 is the application framework schematic diagram of the intelligent cockpit voice processing method shown in the present application; Figure 4 is the second process schematic diagram of the intelligent cockpit voice processing method shown in the present application; Figure 5 is the first structural schematic diagram of the intelligent cockpit voice processing device shown in the present application; Figure 6 is the second structural schematic diagram of the intelligent cockpit voice processing device shown in the present application; Figure 7 is the structural schematic diagram of the intelligent cockpit shown in the present application; Figure 8 is the structural schematic diagram of the intelligent vehicle shown in the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] Embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.

[0025] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0026] It should be understood that although the terms "first", "second", "third", etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.

[0027] The present application provides an intelligent cockpit voice processing method, which can more accurately match the results expected by the user and improve the success rate of matching the results expected by the user.

[0028] The technical solution of the present application will be described in detail below with reference to the accompanying drawings.

[0029] See Figure 2 , an intelligent cockpit voice processing method shown in the present application includes: S211, obtaining voice data issued by the user and scene data of the cockpit screen.

[0030] In this step S211, voice data (query) issued by the user can be obtained. For example, the voice data (query) issued by the user is "Play Chrysanthemum Terrace". Among them, the voice data issued by the user can be collected through a voice collection device such as a microphone.

[0031] In this step S211, scene data of the cockpit screen transmitted by the client of the intelligent cockpit of the intelligent vehicle can also be obtained.

[0032] S212. Match the voice data sent by the user with the node information of the nodes in the matching tree generated according to the scenario data to obtain the initial node information matching result corresponding to the voice data sent by the user.

[0033] Among them, the generation process of the matching tree may include: parsing the scenario data; constructing the matching tree recursively according to the parsing result. For example, the scenario data can be uniformly processed into preset format data; the uniformly processed preset format data is used to construct a tree structure object according to the original format of the scenario data; the tree structure object is used to construct the matching tree recursively, where the root node of the matching tree is created according to the root node of the scenario data, and the non-root nodes of the matching tree with the same structure are constructed in turn by traversing the elements of the scenario data in a depth-first manner.

[0034] In step S212, matching the voice data sent by the user with the node information of the nodes in the matching tree generated according to the scenario data to obtain the initial node information matching result corresponding to the voice data sent by the user may include: Traverse the matching tree generated according to the scenario data, match the voice data sent by the user with the node information of the non-root nodes and the parent nodes of the non-root nodes in the matching tree to obtain the first candidate matching results marked by each node during the traversal process; perform the first similarity calculation on the first candidate matching results, and filter out the initial node information matching result corresponding to the voice data sent by the user according to the first matching similarities of each node.

[0035] S213. Filter the voice data sent by the user in a preset manner to obtain the filtered voice data.

[0036] This step S213 may include: Obtain the unmatched content that is not matched during the matching process of the voice data sent by the user with the node information of the nodes in the matching tree; match the unmatched content with the preset words in the preset word library, and mask the matched content; filter the masked content to obtain the filtered voice data.

[0037] Among them, matching the unmatched content with the preset words in the preset word library and masking the matched content may include: Match the unmatched content at the preset positions of the voice data sent by the user with the preset words in the preset word library; mask the matched content; where the preset positions include at least one of the following positions: prefix position, suffix position, and middle position.

[0038] The preset words may include colloquial words, such as some modal particles, etc.

[0039] S214. Match the filtered voice data with the matching result of the initial node information to obtain the target matching result corresponding to the filtered voice data.

[0040] This step S214 may include: Traverse and match the filtered voice data with the matching result of the initial node information selected at the first preset sorting position to obtain the second candidate matching results marked by each node during the traversal; perform a second similarity calculation on the second candidate matching results, and filter out the target matching result corresponding to the filtered voice data according to the second matching similarities of each node.

[0041] Among them, filtering out the target matching result corresponding to the filtered voice data according to the second matching similarities of each node may include: Filter out the second candidate matching result of the node corresponding to the second matching similarity at the second preset sorting position from the second matching similarities of each node exceeding the preset threshold as the target matching result corresponding to the filtered voice data.

[0042] It can be found that this application can match the voice data sent by the user with the node information of the nodes in the matching tree generated according to the scene data to obtain the matching result of the initial node information corresponding to the voice data sent by the user; filter the voice data sent by the user in a preset manner to obtain the filtered voice data; match the filtered voice data with the matching result of the initial node information to obtain the target matching result corresponding to the filtered voice data. Through the above processing, when processing the scene data information of the client, this application can, on the basis of the first matching, such as the matching result of the initial node information, further filter the voice data sent by the user in a preset manner to reduce the interference of some voice words, and then match the filtered voice data with the matching result of the initial node information again to obtain the target matching result, which can improve the matching accuracy and can more accurately match the result expected by the user, and improve the success rate of matching the result expected by the user.

[0043] Figure 3 It is a schematic diagram of the application framework of the intelligent cockpit voice processing method shown in this application; Figure 4 It is a second process schematic diagram of the intelligent cockpit voice processing method shown in this application.

[0044] In the scenario data obtained by this application, such as the scenario data transmitted or uploaded by the client, it contains scenario tree structure information, such as some element information of each button on the screen. After the functional module of the visible-and-speakable service provided by this application obtains the scenario data, it parses the scenario data and constructs a matching tree structure in a recursive manner. During the NLU matching process, by traversing the matching tree search method, the matching similarity calculation is performed on the nodes of the tree structure of the matching tree in turn, and then several node elements with relatively high matching similarity are selected for secondary matching with the voice data (query) issued by the user. Among them, the way of mask processing can be used to mark the colloquial words at the prefix position, suffix position, middle position, etc. that need to be filtered out in the query, and then continue to perform secondary matching on the query after mask processing with the node information (such as the text of the node) selected for the first time, and calculate the matching degree between the query after filtering out the colloquial words and the node information. Based on the initial node information matching result selected in the first matching, this application then selects the element with a relatively high matching degree, such as the highest matching degree, between the query and the node information through secondary matching as the target matching result (recall result), and returns it to the subsequent link, so as to achieve compatibility with the voice data with colloquial words. Even when there are relatively many colloquial words in the voice data issued by the user, the result that the user expects to hit can be accurately matched, and the success rate of matching the result that the user expects to hit is improved.

[0045] That is to say, this application performs the similarity calculation on the nodes of the tree structure of the matching tree in turn, which is equivalent to the first matching. The matching results can be sorted according to the matching similarity, and the top K results in the front row of the sorting are selected for secondary matching. Among them, K is greater than or equal to 1. This application can select the top K results in the front row of the sorting in the results of the first matching for secondary matching. The matching result with a relatively high matching degree, such as the highest matching degree, in the results of the second matching can be used as the matching result of the visible-and-speakable service finally obtained. In the first matching, mainly the similarity between the label in the node information and the query is calculated. The matching results of the label are sorted according to the matching similarity, and the top K results in the front row of the sorting are selected for secondary matching. In the second matching, based on the results of the first matching, after continuing to mask the colloquial words at the prefix position, suffix position, and middle position of the query, the matching similarity between the query and the node information is calculated continuously, and the result with a relatively high matching degree, such as the highest matching degree, in the results is selected as the final target matching result (recall result), so as to more accurately match the result that the user expects to hit.

[0046] See Figure 4 , and at the same time, it can also be seen Figure 3 , the intelligent cockpit voice processing method of this application may include: S411. Obtain the voice data issued by the user and the scene data of the cockpit screen.

[0047] In this step S411, the voice data (query) issued by the user can be obtained. For example, the voice data (query) issued by the user is "Play Chrysanthemum Terrace". Among them, the voice data issued by the user can be collected through a voice collection device such as a microphone.

[0048] In this step S411, the scene data of the cockpit screen transmitted by the client of the intelligent cockpit of the intelligent vehicle can be obtained. Among them, the client can convert the UI (User Interface) information on the cockpit screen, such as the car machine screen, into scene data. The client can upload the scene data of the cockpit screen through the SDK (Software Development Kit). The SDK generally refers to the tool kit developed to implement a certain function of the product software.

[0049] Among them, the scene data can be in JSON (JavaScript Object Notation, JS key-value pair data) format but is not limited to this. Among them, the scene element in the scene data generally corresponds to the control on the screen. Among them, the client can be, for example, an application installed and running on the car machine system but is not limited to this. The application can be a system-built-in application or a third-party application, etc. The client can be a multimedia application, such as a music application or a video application, etc. Among them, JSON is an open standard file format and data exchange format. It is easy to read and write, and is also easy for machines to parse and generate.

[0050] In this step, the voice data (query) issued by the user is obtained. For example, the query is "Play Chrysanthemum Terrace by Zhou XX", "Play Chrysanthemum Terrace", "Please help me play Chrysanthemum Terrace", or "Please help me play Chrysanthemum Terrace by Zhou XX".

[0051] Among them, the voice data issued by the user can be collected through a voice collection device such as a microphone. Among them, for the obtained voice data issued by the user, speech recognition technology such as ASR (Automatic Speech Recognition) can be used to convert the speech information into the corresponding text. This application does not limit this.

[0052] S412. Parse the scene data.

[0053] In this step S412, after obtaining the voice data issued by the user and the scene data uploaded by the client, the scene data is parsed.

[0054] Among them, the scenario data is generally in JSON format and has a nested relationship layer by layer, so it carries tree structure information and parent-child relationships. The functional module of the visible-and-speak service provided by this application can read and parse the obtained scenario data, and can construct a tree structure object according to the original format of JSON for subsequent use in constructing a matching tree. That is to say, the parsing performed in this step includes uniformly processing the scenario data into preset format data, such as converting it into a dictionary, and then constructing a tree structure object from the uniformly processed preset format data according to the original format of the scenario data, such as JSON format.

[0055] Among them, in a data structure, a tree is a hierarchical data structure composed of nodes, where each node is connected to other nodes through edges to form parent-child relationships. A tree is a non-linear data structure widely used to represent data with hierarchical relationships.

[0056] Among them, a node is the basic element in a tree, containing data and pointers (or references) to child nodes; an edge is a connection line connecting two nodes in a tree; the root node is the top-level node of the tree and has no parent node; a child node is a lower-level node of a certain node; a parent node is a higher-level node of a certain node; a leaf node is a node without child nodes; the depth is the length of the path from a certain node in the tree to the root node.

[0057] It should be noted that the scenario data obtained from the scenario service is generally in JSON format, but not necessarily all in JSON format. It may also be in other formats, such as string and other type formats. Other formats may not be directly parsable. Therefore, the purpose of the parsing process of the scenario data in this step is to uniformly process and convert the obtained scenario data of different types into a unified format, such as converting it into a strict dictionary format, so as to maintain the tree structure in the original data and facilitate the subsequent construction of a matching tree. For example, if the scenario data is in JSON format, the json.loads( ) method can be used to convert it into a dictionary; if the scenario data is in key-value pair format, the string can be split and converted into a dictionary, etc.

[0058] S413. Construct a matching tree according to the parsing result.

[0059] In this step S413, for the tree structure object constructed in the previous step, a matching tree is constructed by means of recursion.

[0060] For example, obtain the tree structure object constructed in the previous step to construct the root node of the matching tree, and recursively obtain the root node on the original scene data and the elements on nodes such as element in a depth-first manner until no more elements can be found. Then return to the previous level and continue traversing other element elements until the entire tree structure object is traversed. Here, element represents a single element, and elements usually represents a collection of multiple elements. elements will be inserted as a key value on the root node or recursively inserted on the elements node of the previous level. The so-called tree structure is recursively constructed with elements. One function of elements is to be a collection of elements at one level in the scene data, including each element, and these elements contain various node information of each node, including label (tag), type, etc.; another function of elements is to be the implementation carrier of the nested tree structure of the scene data tree and the constructed matching tree, and this nesting is achieved through the tree structure nested by elements.

[0061] For example, the label (tag) of a child node includes: "The First Song ∣ Chrysanthemum Terrace", "The Second Song ∣ Jasmine"; the label (tag) of the parent node of this child node includes "Zhou XX".

[0062] This step can use the outermost layer of the obtained tree structure as the root node, that is, use the root node of the scene data tree structure as the input to construct the root node of the matching tree.

[0063] Among them, the construction of the root node of the matching tree can include the following information: 1) scene ID, 2) screen position information, 3) voice position, etc. These information are the global information of the entire matching tree, and the child nodes on the root node are built on the root node with elements as the key value. If there are multiple tree structures in the scene data, multiple root nodes of the matching tree can be constructed respectively, so that different matching trees can be traversed separately during matching.

[0064] Among them, during the recursive traversal process, create a Node (node) and insert it into the matching tree structure. At the same time, create a path from the root node to the current node in the Node, and this path contains node information such as label.

[0065] The matching tree created in this application can be based on the JSON format of the scenario data, and gradually add some information required in the subsequent matching process on the basis of the tree structure of the scenario data. For example, first create the root node of the matching tree according to the root node of the scenario data, and then traverse the elements of the scenario data in a depth-first manner to construct the elements nodes (non-root nodes) of the matching tree with the same structure in sequence. The node information included in the elements node, that is, the non-root node, can include: element ID (identification code), label, element priority, resource name, parent node, child node, visibility, etc. Among them, the element priority indicates the priority of each element hit, that is, when there are two or more elements with the same label on the page, the hit priority can be adjusted by setting the value of the priority; the visibility indicates the visibility of the element. For example, when the value of the visibility is set to TRUE, it means that the element is visible, etc.

[0066] It should be noted that the creation of the Node node (non-root node) here is basically the same as the creation of the aforementioned root node, but the node information of the non-root node and the root node will be different.

[0067] S414, traverse the matching tree, match the voice data issued by the user with the matching tree, and obtain the first candidate matching results marked on each node of the matching tree.

[0068] In this step, obtain the matching tree constructed in the previous step, and match the voice data issued by the user with the matching tree in a recursive traversal manner. Among them, the voice data issued by the user can be recognized by using voice recognition technology such as ASR to convert the voice information into corresponding text for matching.

[0069] That is to say, traverse and match the voice data issued by the user with the node information of the non-root node and the parent node of the non-root node of the matching tree. Among them, the matching tree can be traversed upward from the non-root node according to the preset hierarchical parent node.

[0070] Among them, it includes matching each traversed node with the voice data issued by the user, matching the node label saved in the path of the node with the voice data issued by the user, and marking the single characters that are matched.

[0071] Since the structure information of the scenario data cannot determine the depth of the node, there may be nodes with a very deep depth. If it is completely matched from the root node to the leaf node, it may be unreasonable. Therefore, the matching of the node label and the voice data issued by the user can be to match a preset level, for example, set to match the length of 2 parent nodes, that is, traverse and match from a non-root node upward for the length of 2 parent nodes.

[0072] Among them, the content in the node label can be the information on the cockpit screen UI. The matching with the voice data (query) issued by the user can include the matching of text and some supported action words, such as click, select, etc. For example, the labels on the cockpit screen UI can include: "The first song ∣ Chrysanthemum Terrace", "The second song ∣ Jasmine Flower", "Zhou XX", etc.

[0073] S415, perform a first similarity calculation on the first candidate matching results, and filter out the initial node information matching results corresponding to the voice data issued by the user according to the first matching similarities of each node.

[0074] In this step, the first candidate matching results corresponding to the first matching similarities located at the first preset sorting position can be filtered out from the first matching similarities of each node exceeding the preset threshold as the initial node information matching results corresponding to the voice data issued by the user.

[0075] Among them, the preset sorting position can be, for example, the top K positions, that is, the top K positions in the front row of the sorting.

[0076] In this step, according to the marks of the single characters matched by the nodes and the parent nodes on the path in the previous step, the first matching similarity of each node can be calculated respectively, and the results exceeding the preset threshold and having a relatively large matching similarity, such as the top K results in the front row of the sorting, are selected as the initial node information matching results.

[0077] Among them, the first matching similarity can refer to the similarity of text matching. Among them, the preset threshold can be set to 1 but is not limited to this. At this time, it means that the voice data (query) issued by the user is completely matched with the label information. The preset threshold can also be adjusted as needed.

[0078] Assume that the screen includes a song list of a singer. When the user says "Play Chrysanthemum Terrace", the relevant technical solutions can match and hit the song "Chrysanthemum Terrace" on the screen after matching processing. When the user says "Play Chrysanthemum Terrace by Zhou XX" or "Play Zhou XX's Chrysanthemum Terrace", the relevant technical solutions cannot match and hit the song "Chrysanthemum Terrace" on the screen after matching processing. The technical solution provided by this application, through the processing of the above steps, can, when processing the scenario data information uploaded by the client, take into account the content of the parent-child relationship nodes at the same time. By constructing a matching tree of scenario information, the information of the parent-child relationship nodes can be discovered at the same time during matching, that is, the content of the parent-child relationship nodes can be processed at the same time during the NLU understanding of the visible-and-speakable service, so that the song "Chrysanthemum Terrace" by Zhou XX on the screen can be hit, and the matching of the content containing the parent-child relationship nodes in the voice data (query) issued by the user can be solved, improving the understanding ability of the visible-and-speakable service for the screen, more accurately matching the result that the user hopes to hit according to the voice data issued by the user, and improving the accuracy of voice interaction matching.

[0079] S416. Mask the voice data issued by the user in a preset manner to obtain the filtered voice data after masking.

[0080] This application can filter the voice data issued by the user in a preset manner, for example, by masking, to obtain the filtered voice data after masking.

[0081] For example, it may include: Obtain the unmatched content that is not matched during the matching process between the voice data issued by the user and the node information of the nodes in the matching tree; Match the unmatched content with the preset words in the preset word library, and mask the matched content; Filter the masked content to obtain the filtered voice data.

[0082] Among them, matching the unmatched content with the preset words in the preset word library and masking the matched content may include: Match the unmatched content at the preset position of the voice data issued by the user with the preset words in the preset word library, and mask the matched content; where the preset position includes at least one of the following positions: prefix position, suffix position, and middle position.

[0083] Among them, for the unmatched content that is not matched during the matching process between the voice data issued by the user and the node information of the nodes in the matching tree, the colloquial word list loaded can be used to perform colloquial word matching at the prefix position, suffix position, and middle position respectively, and the matched content is masked.

[0084] A mask is a technique widely used in computer science and data processing. It can be used to mask and filter data, enabling selective operations on target objects. A mask achieves specific functions by screening and manipulating certain parts of the data.

[0085] Among them, the preset words can include colloquial words, such as some modal particles, etc. This application can pre-load a colloquial word list on the vehicle side and / or the server side and can be updated regularly.

[0086] For example, if the voice data query sent by the user is "Quickly play Chrysanthemum Terrace" and the unmatched content that is not matched during the matching process with the node information of the nodes in the matching tree includes "Quickly play". Use the loaded colloquial word list to match colloquial words at the prefix position, and mask and mark the matched content "Quickly play". Then, filter out the masked content "Quickly play" to obtain the filtered voice data "play Chrysanthemum Terrace" for subsequent matching processing.

[0087] This application summarizes several different types of possible colloquial words, which generally mainly appear in the following positions of the voice data: the prefix position, the suffix position, and the middle position. Among them, the ones in the middle position are generally middle stop words.

[0088] The colloquial word list can be stored in a preset word library, configured into the project in the form of a file, and this file can be loaded during the project operation and used as a global variable when the corresponding content is not successfully matched in the query.

[0089] Prefix position: Generally refers to the position at the beginning of the query.

[0090] For example, if the query is "Quickly play Chrysanthemum Terrace", the words "Quickly play", "play", and "play for me" among them are colloquial words.

[0091] Some other common colloquial words such as "help", "help me", "help me play", "ah help", "ah help me", "ah play for me", "ah please give" etc.

[0092] Suffix position: Generally refers to the position at the end of the query.

[0093] For example, if the query is "Play Chrysanthemum Terrace, please", the word "please" among them is a colloquial word.

[0094] Some other common colloquial words such as "ah", "ah", "ba", "ha", "is that okay", "okay", "alright", "is it good", "is it okay" etc.

[0095] Middle position: Generally, it is a middle stop word, usually referring to the position in the middle of the query.

[0096] For example, if the query is "Can Qi Li Xiang be played?", the phrase "Can" is an oral expression.

[0097] Some other common oral expressions such as "Can you help", "Can you help me", "Can you take", "Can you do it for me", "Can you give", "Can you give me", "Can", "Can you", "Can help", etc.

[0098] This step performs masking processing to obtain the filtered voice data after masking. For example, if the query is "Quickly play Chrysanthemum Terrace for me", the filtered voice data obtained through masking is "Play Chrysanthemum Terrace"; if the query is "Play Chrysanthemum Terrace", the filtered voice data obtained through masking is "Play Chrysanthemum Terrace".

[0099] S417, traverse and match the filtered voice data with the matching results of the initial node information located at the first preset sorting position to obtain the second candidate matching results marked by each node during the traversal.

[0100] Based on the matching results of the first matching similarity between the label and the query calculated in the previous steps of this application, the matching results of the initial node information located at the first preset sorting position can be screened out. Then, the filtered voice data is traversed and matched with the matching results of the initial node information located at the first preset sorting position to obtain the second candidate matching results marked by each node during the traversal.

[0101] For example, the first few nodes with higher first matching similarity, that is, higher matching degree, can be selected from the matching results of the initial node information, that is, the top K nodes with higher matching degree in the front row of the sorting are selected. Then, these nodes are traversed, and the node information of these nodes is matched with the query after masking again in a recursive traversal manner to obtain the second candidate matching results marked by each node during the traversal.

[0102] Among them, it includes matching each traversed node with the query after masking, that is, matching the node label saved in the path of the node with the query after masking, and marking the single characters that match.

[0103] S418, perform a second similarity calculation on the second candidate matching results, and screen out the target matching results corresponding to the filtered voice data according to the second matching similarity of each node.

[0104] In step S418, from the second matching similarities of each node that exceed the preset threshold, the second candidate matching result of the node corresponding to the second matching similarity located at the second preset sorting position can be screened out as the target matching result corresponding to the filtered voice data. This target matching result can be used as the final matching result and returned to the subsequent link.

[0105] The second matching similarity located at the second preset sorting position can be, for example, the numerically largest second matching similarity ranked first.

[0106] In this step, according to the markings of the single characters matched on the nodes and paths in the previous step, the second matching similarity of each node can be calculated respectively, and the nodes that exceed the preset threshold and have a relatively large (such as the largest) second matching similarity can be screened out as the final target matching result and returned to the subsequent link, thereby achieving a more accurate matching of the results desired by the user and improving the matching accuracy. According to the final target matching result, that is, the natural language understanding result, the subsequent link can generate an operation instruction, such as selecting the song "Juhuatai" by Zhou XX for playback.

[0107] The second matching similarity can refer to the similarity of text matching. The preset threshold can be set to 1, but it is not limited to this. At this time, it means that the filtered voice data (query) is completely matched with the node information. The preset threshold can also be adjusted as needed.

[0108] It can be found that when the NLU of the visible-and-speak service in this application is understanding, it can process the content of the parent-child relationship nodes at the same time, and solve the matching of the content of the parent-child relationship nodes in the voice data (query) issued by the user. This application also realizes the compatibility of the voice data with colloquial words through the processing of colloquial words, etc. This application can load the colloquial word list through global configuration; the content that cannot be matched during the matching process of the query and the node label can be masked by the colloquial word list, and then the matching degree between the masked query and the node information can be calculated. By processing the colloquial words in the voice data, this application can improve the generalization ability of the understanding of the voice data (voice input) issued by the user, so that even when there are many colloquial words in the voice data issued by the user, the results desired by the user can be accurately matched, and the success rate of matching the results desired by the user can be improved.

[0109] Corresponding to the foregoing application function implementation method embodiments, this application also provides an intelligent cockpit voice processing device, an intelligent cockpit, an intelligent vehicle, and corresponding embodiments.

[0110] Figure 5 It is the first structural schematic diagram of the intelligent cockpit voice processing device shown in this application.

[0111] As shown Figure 5 in the figure, the intelligent cockpit voice processing device 600 provided by the present application includes: a data acquisition module 601, a first matching module 602, a filtering processing module 603, and a second matching module 604.

[0112] The data acquisition module 601 is configured to acquire voice data emitted by a user and scenario data of a cockpit screen; The first matching module 602 is configured to match the voice data emitted by the user with the node information of the nodes in the matching tree generated according to the scenario data, and obtain an initial node information matching result corresponding to the voice data emitted by the user; The filtering processing module 603 is configured to filter the voice data emitted by the user in a preset manner to obtain filtered voice data; The second matching module 604 is configured to match the filtered voice data with the initial node information matching result to obtain a target matching result corresponding to the filtered voice data.

[0113] The device provided by the present application can, when processing the scenario data information of the client, on the basis of the first matching, such as the initial node information matching result, further filter the voice data emitted by the user in a preset manner to reduce the interference of some voice words, and then match the filtered voice data with the initial node information matching result again to obtain a target matching result, which can improve the matching accuracy, can more accurately match the result that the user expects to hit, and improve the success rate of matching the result that the user expects to hit.

[0114] Figure 6 is the second structural schematic diagram of the intelligent cockpit voice processing device shown in the present application.

[0115] As shown Figure 6 in the figure, the intelligent cockpit voice processing device 600 provided by the present application includes: a data acquisition module 601, a first matching module 602, a filtering processing module 603, and a second matching module 604.

[0116] Among them, the functions of the data acquisition module 601, the first matching module 602, the filtering processing module 603, and the second matching module 604 can be referred to Figure 5 as described, and will not be elaborated here.

[0117] Among them, the filtering processing module 603 may include: an unmatched content acquisition sub-module 6031, a mask processing sub-module 6032, and a filtering sub-module 6033.

[0118] The unmatched content acquisition sub-module 6031 is configured to acquire the unmatched content that is not matched during the matching process of the voice data emitted by the user with the node information of the nodes in the matching tree; The mask processing sub-module 6032 is used to match the unmatched content with the preset words in the preset word library and mark the matched content with a mask; The filtering sub-module 6033 is used to filter the content marked with a mask to obtain the filtered voice data.

[0119] Among them, the mask processing sub-module 6032 can match the unmatched content at the preset positions of the voice data sent by the user with the preset words in the preset word library; mark the matched content with a mask; where the preset positions include at least one of the following positions: the prefix position, the suffix position, and the middle position.

[0120] Among them, the first matching module 602 can traverse the matching tree generated according to the scenario data, match the voice data sent by the user with the node information of the non-root nodes and the parent nodes of the non-root nodes in the matching tree, and obtain the first candidate matching results marked by each node during the traversal process; Perform a first similarity calculation on the first candidate matching results, and filter out the initial node information matching results corresponding to the voice data sent by the user according to the first matching similarities of each node.

[0121] Among them, the second matching module 604 can perform a traversal match on the filtered voice data and the initial node information matching results located at the first preset sorting position selected, and obtain the second candidate matching results marked by each node during the traversal process; Perform a second similarity calculation on the second candidate matching results, and filter out the target matching results corresponding to the filtered voice data according to the second matching similarities of each node.

[0122] The second matching module 604 can select the second candidate matching results of the nodes corresponding to the second matching similarities located at the second preset sorting position from the second matching similarities exceeding the preset threshold of each node as the target matching results corresponding to the filtered voice data.

[0123] The device provided by this application can, when processing the scenario data information uploaded by the client, take into account the content of the parent-child relationship nodes at the same time. By constructing a matching tree for the scenario information, the information of the parent-child relationship nodes can be discovered simultaneously during matching, that is, the content of the parent-child relationship nodes can be processed simultaneously during the NLU understanding of the visible-and-speakable service. Thus, it can solve the matching of the content of the parent-child relationship nodes in the voice data (query) issued by the user, and improve the screen understanding ability of the visible-and-speakable service. This application also processes colloquial words, etc., to achieve compatibility with voice data with colloquial words, and improve the generalization ability of understanding the voice data (voice input) issued by the user. In this way, even if there are many colloquial words in the voice data issued by the user, the result expected to be hit by the user can be accurately matched, and the success rate of matching the result expected to be hit for the user can be improved.

[0124] Figure 7 It is a schematic structural diagram of the intelligent cockpit shown in this application.

[0125] As Figure 7 shown, an intelligent cockpit 800 provided by this application may include the intelligent cockpit voice processing device 600 as shown above Figure 5 or Figure 6 shown.

[0126] It should be noted that this application may also provide a server, and this server may also include the intelligent cockpit voice processing device 600 as shown above Figure 5 or Figure 6 shown.

[0127] The server can obtain the scenario data of the cockpit screen uploaded by the client of the intelligent cockpit of the intelligent vehicle, and obtain the voice data (query) issued by the user uploaded by the intelligent cockpit of the intelligent vehicle. After processing to obtain the matching result corresponding to the voice data issued by the user, the server sends it to the intelligent vehicle cockpit for corresponding processing.

[0128] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0129] Figure 8 It is a schematic structural diagram of the intelligent vehicle shown in the embodiments of this application.

[0130] See Figure 8 , the intelligent vehicle 1000 includes a memory 1010 and a processor 1020.

[0131] The processor 1020 can be a Central Processing Unit (CPU), or it can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0132] The memory 1010 can include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM can store static data or instructions required by the processor 1020 or other modules of the computer. The permanent storage device can be a read-write storage device. The permanent storage device can be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (such as a floppy disk, optical drive). The system memory can be a read-write storage device or a volatile read-write storage device, such as dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during operation. In addition, the memory 1010 can include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks can also be used. In some embodiments, the memory 1010 can include a removable storage device that is readable and / or writable, such as a compact disc (CD), read-only digital versatile disc (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray disc, super density disc, flash memory card (such as SD card, min SD card, Micro-SD card, etc.), magnetic floppy disk, etc. The computer-readable storage medium does not include carrier waves and instantaneous electronic signals transmitted wirelessly or by wire.

[0133] An executable code is stored on the memory 1010. When the executable code is processed by the processor 1020, it can cause the processor 1020 to execute some or all of the methods described above.

[0134] In addition, the method according to the present application can also be implemented as a computer program or a computer program product, which includes computer program code instructions for performing some or all of the steps in the above-mentioned method of the present application.

[0135] Alternatively, the present application can also be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium), on which executable code (or a computer program or computer instruction code) is stored. When executed by a processor of an electronic device (or a server, etc.), the processor is caused to execute some or all of the steps of the above-mentioned method according to the present application.

[0136] The present application also provides a computer program product, which includes computer instructions that, when executed by a processor, implement the method as described above.

[0137] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the disclosed embodiments.

Claims

1. A smart cockpit voice processing method, characterized in that: include: Acquire the voice data sent by the user and the scene data of the cockpit screen; Matching the voice data sent by the user with the node information of the nodes in the matching tree generated according to the scene data to obtain an initial node information matching result corresponding to the voice data sent by the user; Filtering the voice data sent by the user in a preset manner to obtain filtered voice data; The filtered voice data is matched with the initial node information matching result to obtain a target matching result corresponding to the filtered voice data.

2. The method according to claim 1, characterized in that The filtering of the voice data sent by the user in a preset manner to obtain filtered voice data includes: Acquire unmatched content that is not matched in the process of matching the voice data sent by the user with the node information of the nodes in the matching tree; Matching the unmatched content with preset words in a preset word library, and masking the matched content; The content marked by the mask is filtered to obtain filtered voice data.

3. The method according to claim 2, characterized in that The step of matching the unmatched content with preset words in a preset word library and performing mask marking on the matched content includes: Matching the unmatched content at a preset position of the voice data sent by the user with preset words in a preset word library; Mask the matched content; The preset position includes at least one of the following positions: a prefix position, a suffix position, and a middle position.

4. The method according to claim 2, characterized in that: The preset words include colloquial words.

5. The method according to claim 1, characterized in that The matching of the voice data sent by the user with the node information of the nodes in the matching tree generated according to the scene data to obtain the initial node information matching result corresponding to the voice data sent by the user includes: Traversing a matching tree generated according to the scene data, matching the voice data sent by the user with node information of a non-root node and a parent node of the non-root node in the matching tree, and obtaining a first candidate matching result of each node mark in the traversal process; A first similarity calculation is performed on the first candidate matching result, and the initial node information matching result corresponding to the voice data sent by the user is screened out according to the first matching similarity of each node.

6. The method according to claim 1, characterized in that The step of matching the filtered voice data with the initial node information matching result to obtain a target matching result corresponding to the filtered voice data includes: Traversing and matching the filtered voice data with the selected initial node information matching results located at the first preset sorting position to obtain second candidate matching results for each node mark in the traversal process; A second similarity calculation is performed on the second candidate matching result, and a target matching result corresponding to the filtered voice data is screened out according to the second matching similarity of each node.

7. The method according to claim 6, characterized in that The step of screening out target matching results corresponding to the filtered voice data according to the second matching similarity of each node includes: From the second matching similarities of the nodes exceeding the preset threshold, the second candidate matching results of the nodes corresponding to the second matching similarities at the second preset sorting positions are screened out as the target matching results corresponding to the filtered voice data.

8. An intelligent cockpit voice processing device, characterized in that: include: A data acquisition module, used to acquire voice data sent by the user and scene data of the cockpit screen; A first matching module, used to match the voice data sent by the user with the node information of the nodes in the matching tree generated according to the scene data, to obtain an initial node information matching result corresponding to the voice data sent by the user; A filtering processing module, used for filtering the voice data sent by the user in a preset manner to obtain filtered voice data; The second matching module is used to match the filtered voice data with the initial node information matching result to obtain a target matching result corresponding to the filtered voice data.

9. The device according to claim 8, characterized in that The filtering processing module comprises: An unmatched content acquisition submodule, used to acquire unmatched content that is not matched in the process of matching the voice data sent by the user with the node information of the nodes in the matching tree; A mask processing submodule, used for matching the unmatched content with preset words in a preset word library, and performing mask marking on the matched content; The filtering submodule is used to filter the contents marked by the mask to obtain filtered voice data.

10. A smart cockpit, characterized in that: It comprises the intelligent cockpit voice processing device as described in any one of claims 8-9.

11. A smart car, characterized in that: include: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to execute the method according to any one of claims 1 to 7.

12. A computer-readable storage medium having executable codes stored thereon, which, when executed by a processor of an electronic device, causes the processor to execute the method according to any one of claims 1 to 7.

13. A computer program product, characterized in that The computer program product comprises computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Voice interaction method, server and computer readable storage medium

    CN120913561A