Interaction control method and device, equipment and storage medium

By configuring command tags in the on-board terminal, using the tag body and extended tag to match the text information of voice requests, the command execution error caused by inaccurate user statements is solved, and more accurate voice control is achieved.

CN120496522APending Publication Date: 2025-08-15GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510773594.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When a user is interacting with the on-board terminal, the on-board terminal cannot correctly execute the user's instructions due to inaccurate statements.

Method used

By configuring the instruction tag of preset instructions in the vehicle terminal, including the tag body and the extended tag, the text information of the voice request matches the instruction tag, the target instruction is determined, and the instruction is executed.

Benefits of technology

It improves the generalization of voice requests by vehicle terminals, ensures that the user's intentions are executed more accurately, and solves the problem of mismatch between voice requests and instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496522A_ABST
    Figure CN120496522A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an interaction control method and device, equipment and a storage medium. The method comprises the steps that a voice request is acquired; a target instruction is determined according to text information corresponding to the voice request and an instruction label of a preset instruction, the instruction label of the preset instruction comprises a label body and at least one extension label, the label body is determined according to display content in a target interface, and the at least one extension label is used for describing the label body through different text information; the at least one preset instruction corresponds to display content in the target interface; and executing the target instruction. Generalized understanding of the vehicle-mounted terminal on the voice request can be improved, and further the vehicle-mounted terminal can execute the target instruction more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to smart cockpit technology, and relate to but are not limited to an interactive control method and apparatus, equipment, and storage medium. Background Art

[0002] When users interact with the vehicle terminal through voice, they usually need to use corresponding sentences to control the vehicle terminal to execute a certain instruction, such as: opening and closing windows, opening and closing doors, or switching to playing music.

[0003] In the related art, the voice interaction between the user and the vehicle terminal often results in inaccurate user statements, which causes the vehicle terminal to be unable to correctly feedback the user's voice and thus unable to control the vehicle terminal to execute corresponding instructions. Summary of the Invention

[0004] The interactive control method, apparatus, device, and storage medium provided in the embodiments of the present application are implemented as follows:

[0005] In one aspect of an embodiment of the present application, an interactive control method is provided, which is applied to a vehicle-mounted terminal. The vehicle-mounted terminal includes a display screen, and a target interface is displayed on the display screen. The method includes:

[0006] Get voice request;

[0007] Determining a target instruction based on the text information corresponding to the voice request and an instruction tag of a preset instruction, wherein the instruction tag of the preset instruction includes a tag body and at least one extended tag, wherein the tag body is determined based on the content displayed on the target interface, and the at least one extended tag is used to describe the tag body using different text information, and the at least one preset instruction corresponds to the content displayed on the target interface;

[0008] Execute the target instruction.

[0009] In another aspect of the embodiment of the present application, an interactive control device is provided, which is applied to a vehicle-mounted terminal, wherein the vehicle-mounted terminal includes a display screen, and a target interface is displayed on the display screen. The device includes: an acquisition module, a determination module, and an execution module;

[0010] Acquisition module, used to obtain voice requests;

[0011] a determination module, configured to determine a target instruction based on text information corresponding to the voice request and an instruction tag of a preset instruction, wherein the instruction tag of the preset instruction includes a tag body and at least one extended tag, wherein the tag body is determined based on content displayed on the target interface, and the at least one extended tag is configured to describe the tag body using different text information, and the at least one preset instruction corresponds to content displayed on the target interface;

[0012] The execution module is used to execute the target instructions.

[0013] The computer device provided in the embodiment of the present application includes a memory and a processor. The memory stores a computer program that can be run on the processor. When the processor executes the program, the method of the embodiment of the present application is implemented.

[0014] The computer-readable storage medium provided in the embodiment of the present application stores a computer program thereon, and when the computer program is executed by a processor, the method provided in the embodiment of the present application is implemented.

[0015] In the interactive control method, apparatus, device, and storage medium provided in the embodiments of the present application, a voice request can be obtained; a target instruction can be determined based on the text information corresponding to the voice request and the instruction label of the preset instruction; and the target instruction can be executed. The instruction label of the preset instruction includes a label body and at least one extended label, the label body is determined based on the content displayed in the target interface, the at least one extended label is used to describe the label body through different text information, and the at least one preset instruction corresponds to the content displayed in the target interface. Since the different extended tags corresponding to each instruction label can describe the label body through different text information, there can be multiple description methods for the same instruction label. In the process of determining the target instruction, the text information corresponding to the voice request can be matched with the multiple description methods, so that the target instruction indicated by the text information corresponding to the voice request can be determined more accurately, thereby improving the vehicle-mounted terminal's generalized understanding of the voice request, and further allowing the vehicle-mounted terminal to execute the target instruction more accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 A schematic diagram of an application scenario provided in an embodiment of the present application;

[0018] Figure 2 This is a schematic diagram of the interface of the vehicle terminal provided in an embodiment of the present application;

[0019] Figure 3 A flow chart of the interactive control method provided in an embodiment of the present application;

[0020] Figure 4 A schematic diagram of a process for determining the instruction tag with the highest matching degree provided in an embodiment of the present application;

[0021] Figure 5 This is another flowchart of determining the instruction tag with the highest matching degree provided in an embodiment of the present application;

[0022] Figure 6 This is another flowchart of determining the instruction tag with the highest matching degree provided in an embodiment of the present application;

[0023] Figure 7 This is a schematic diagram of the structure of the target matching tree provided in the embodiments of the present application;

[0024] Figure 8 A schematic diagram of a process for determining the degree of matching of matching nodes based on a matching tree provided in an embodiment of the present application;

[0025] Figure 9 This is another flowchart of determining the matching degree of a matching node based on a matching tree provided in an embodiment of the present application;

[0026] Figure 10 This is a schematic diagram of the structure of the interactive control device provided in an embodiment of the present application;

[0027] Figure 11 This is a schematic diagram of the structure of the computer device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0028] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the specific technical solutions of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0030] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0031] It should be pointed out that the terms "first\second\third" involved in the embodiments of the present application are used to distinguish similar or different objects, and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0032] The following explains one feasible application scenario of the interactive control method provided in the embodiment of the present application.

[0033] Figure 1 This is a schematic diagram of the application scenario provided in the embodiment of this application, please refer to Figure 1 The scenario may include: a vehicle-mounted terminal 100, wherein the vehicle-mounted terminal 100 may be, for example, an intelligent vehicle machine. The vehicle-mounted terminal 100 is set in a smart cockpit and can have an intelligent dialogue with a user in the smart cockpit, for example: it can perform corresponding tasks based on the user's instructions.

[0034] For example, the user may also request the vehicle terminal 100 to perform a task, such as requesting the vehicle terminal 100 to open the window, requesting the vehicle terminal 100 to turn on the air conditioner, and requesting the vehicle terminal 100 to play music.

[0035] Correspondingly, the vehicle-mounted terminal 100 can execute corresponding instructions according to the user's request. For example, the vehicle-mounted terminal 100 can open the window at the corresponding position according to the user's window opening request, or the vehicle-mounted terminal 100 can also open the door at the corresponding position according to the user's door opening request, etc. There is no specific restriction here, and it can be any instruction that the vehicle-mounted terminal 100 can implement.

[0036] In this scenario, one user can interact with the vehicle terminal 100, or multiple users can interact with the vehicle terminal 100, without specific restrictions. For example, the user in the main driver's seat can interact with the vehicle terminal 100, and the user in the co-pilot seat can also interact with the vehicle terminal 100.

[0037] Among them, the process of user interaction with the vehicle terminal 100 can be achieved through voice request, for example: the user speaks the corresponding voice request, the vehicle terminal 100 collects the corresponding audio signal, and performs voice recognition to determine the user's intention, and then executes the corresponding instruction.

[0038] In one embodiment, the vehicle-mounted terminal 100 includes a display screen 110 , and a target interface is displayed on the display screen 110 .

[0039] It should be noted that a variety of optional interfaces can be displayed in the display screen 110. For example, when the user plays audio, the interface displayed in the display screen 110 may be a music playback interface; when the user is reversing, the display screen 110 may display the picture taken by the reversing camera, etc. There is no specific limitation here, and the target interface displayed in the display screen 110 can be set according to the actual needs of the user.

[0040] The following is an explanation of the interface diagram provided by the vehicle-mounted terminal in the embodiment of the present application using a specific display interface on the display screen.

[0041] Figure 2 This is a schematic diagram of the interface of the vehicle terminal provided in the embodiment of this application, please refer to Figure 2 The content displayed in this interface may be an interface for assisted driving, for example, it may include: obstacle avoidance assistance when starting, blind spot safety assistance, door opening warning, rear collision warning, etc.

[0042] Optionally, the user can use the target interface to turn on or off the auxiliary function or warning function corresponding to the vehicle terminal.

[0043] As for the obstacle avoidance assistance at start, it mainly monitors the surrounding environment of the vehicle in real time during the vehicle's start phase through the sensors equipped in the vehicle (such as radar, camera, etc.). When obstacles are detected around the vehicle and the vehicle has a tendency to move forward or backward, the system will promptly issue an alarm to remind the driver that there is a potential collision risk around it. If necessary, it can also actively intervene and take measures such as braking to prevent the vehicle from colliding with obstacles and ensure driving safety. Its working principle is mainly that when the vehicle is stationary, the sensor continuously scans the area around the vehicle body. Once it is identified that the obstacle is too close and the vehicle gear is not in neutral (such as engaged in forward gear or reverse gear to start), the system will be activated and issue warnings through sound, image or vibration. Some advanced on-board terminals can also calculate a safe driving trajectory to assist the driver in starting smoothly and avoiding obstacles.

[0044] Blind spot safety assistance primarily monitors vehicles, pedestrians, or other objects in the blind spots on both sides of the vehicle (i.e., areas not covered by the rearview mirror, generally a sector-shaped area approximately 30-50 degrees behind the vehicle on either side). When other road users enter the blind spot, the onboard terminal promptly alerts the driver to avoid collisions caused by visual blind spots during lane changes, turns, and other maneuvers. It works by relying on radar sensors installed around the vehicle body to monitor dynamic conditions within the blind spot in real time. When an object enters the blind spot, the sensor transmits a signal to the vehicle's electronic control unit (ECU), which, after processing, triggers the corresponding warning device to alert the driver to the situation within the blind spot.

[0045] Door opening warnings primarily monitor traffic conditions to the side and rear of the vehicle when the vehicle is stopped or about to stop and the occupants are about to open the door. If a collision risk is detected with a rapidly approaching vehicle, bicycle, or pedestrian, an alarm will sound, reminding occupants not to open the door immediately to avoid traffic accidents caused by sudden door opening. The system uses the vehicle's sensors (such as millimeter-wave radar and cameras) to continuously monitor the area to the side and rear of the vehicle. When the door handle is pulled or the door attempts to open, the onboard terminal quickly determines the traffic conditions to the side and rear. If a dangerous object approaches, it immediately alerts occupants through sound, light, and other warnings, preventing the door from opening.

[0046] Rear collision warning mainly monitors traffic conditions behind the vehicle while the vehicle is in motion, especially when driving at low speed or reversing. When it detects other vehicles, pedestrians, bicycles, etc. approaching at a high speed from behind and there is a risk of collision, the system will promptly issue an alarm to remind the driver to pay attention to the danger behind and take measures such as slowing down, braking, or adjusting the driving direction in advance. Its working principle is to use sensors at the rear of the vehicle (such as reversing radar, rear camera, etc.) to monitor information such as the distance and relative speed of objects behind in real time. When it is determined that the approaching speed and distance of the rear object reach the preset danger threshold, the on-board terminal will alert the driver of the dangerous situation behind through the vehicle's warning devices (such as buzzers, display prompts, etc.), allowing the driver to react in time to avoid collision.

[0047] for Figure 2 The four auxiliary warning functions shown in the example can be turned on or off by users through voice requests. For example, users can use the voice request "turn on rear collision warning" to control the vehicle terminal to turn on the corresponding function.

[0048] However, in actual applications, users may not be able to fully remember the voice request corresponding to each function. For example, the user may say a few words incorrectly when using the voice request of "turn on rear collision warning", or express the meaning through synonyms or antonyms, such as "turn on reverse collision warning", "turn on reverse warning", etc. Such situations will cause the vehicle terminal to be unable to determine the command corresponding to the user's voice request, and thus unable to execute the command that the user actually needs to execute.

[0049] In order to solve the above problems existing in the related art, an interactive control method is provided in an embodiment of the present application. This interactive control method can solve the problem of mismatch between requests and instructions existing in the related art. The following explains one of the feasible implementation methods of the interactive control method provided in the embodiment of the present application.

[0050] Figure 3For a flow chart of the interactive control method provided in the embodiment of this application, please refer to Figure 3 , the method comprising:

[0051] S310: Acquire a voice request.

[0052] It should be noted that the execution subject of this method can be the above-mentioned vehicle-mounted terminal, for example, it can be the smart car machine in the smart cockpit.

[0053] Optionally, the voice request may be a request initiated by the user, and may be a sentence spoken by the user in the smart cockpit, for example, a request from the user to the vehicle terminal.

[0054] In one embodiment, after the vehicle-mounted terminal obtains the voice request, it may perform audio-to-text conversion on the voice request, thereby obtaining text information of the voice request.

[0055] For example, a voice recognition module can be set in the vehicle terminal, through which the above-mentioned voice request can be converted into text information. In addition, an intention recognition module can also be set in the vehicle terminal, which can determine the user's intention based on the corresponding text information and thus guess the user's thoughts.

[0056] It should be noted that the above two modules can be set separately, or can also be set in one. For example, a semantic recognition module is set to determine the user's intention by recognizing the voice request.

[0057] S320: Determine a target instruction according to the text information corresponding to the voice request and the instruction label of the preset instruction.

[0058] It should be noted that the text information corresponding to the voice request can be determined through the above-mentioned voice recognition module.

[0059] Multiple preset instructions can be configured in the vehicle terminal. The instruction tag of each preset instruction includes a tag body and at least one extended tag. The tag body is determined according to the content displayed in the target interface, and at least one extended tag is used to describe the tag body through different text information.

[0060] Among them, the preset instructions can be executable instructions pre-configured in the vehicle terminal, such as: door opening instructions, door closing instructions, window opening instructions, window closing instructions, etc. There is no specific restriction here, and the corresponding preset instructions can be configured according to actual needs.

[0061] For each preset instruction, an instruction label can be set. For example, for the instruction "open the main driver's window", the corresponding instruction label can be the text "open the main driver's window". The text of each instruction label can be composed of multiple keywords. For example, "open the main driver's window" can include the two keywords "open" and "main driver's window". The keywords of the label body are the above-mentioned "open" and "main driver's window". The keywords can be extended: the extended description of "open" can include "open", "pull open", "pull down", etc.; the extended description of "main driver's window" can include "left front window", "window No. 1", "main window", etc.; the extended description of each of the above keywords can be combined as an extended tag of the instruction tag, for example, "open window No. 1", "pull open the left front window", etc., can all be used as one of the extended tags.

[0062] It should be noted that for each instruction tag, its tag body can be determined based on the specific text of the content displayed in the target interface, and its extended tag can be obtained in advance after expansion based on the tag body, can be stored after manual expansion, or can be generated by synonyms, antonyms and other generation tools and then stored. No specific restrictions are made here.

[0063] In one embodiment, at least one preset instruction corresponds to content displayed in the target interface.

[0064] For example, if the content displayed in the target interface is a music playback interface, the corresponding preset instructions may include: switch to the previous song, switch to the next song, start / pause playback, switch to the lyrics interface, return to the music list interface, increase the volume, decrease the volume, and other preset instructions; accordingly, the content displayed in the target interface is as follows: Figure 2 In the case of the assisted driving interface shown in , the corresponding instructions may include: turning on / off the start-up obstacle avoidance assist, turning on / off the blind spot safety assist, turning on / off the door opening warning, turning on / off the rear collision warning, returning to the previous level interface and other preset instructions.

[0065] That is to say, at least one preset instruction can be determined based on the content displayed in the target interface. Each preset instruction corresponds to an instruction tag. The instruction tag can include a tag body and an extended tag. The tag body and the extended tag can express the preset instruction through different text descriptions.

[0066] In one embodiment, a target instruction may be determined by matching text information corresponding to the voice request with an instruction tag of each preset instruction in at least one preset instruction, wherein the target instruction may be one of the preset instructions.

[0067] S330: Execute the target instruction.

[0068] It should be noted that after determining the target instruction from the preset instructions, the target instruction can be executed. For example, when the target instruction is determined to be the instruction "open the main driver's window", the instruction can be executed and the vehicle-mounted terminal controls the opening of the main driver's window.

[0069] In one embodiment, the target instruction may be a single instruction or multiple consecutive instructions, which is not specifically limited herein. In actual implementation, the target instruction may be executed according to actual needs.

[0070] Optionally, before executing the target instruction, the vehicle terminal can determine whether the target instruction can be executed at present. For example, when the vehicle is driving at high speed, the target instruction is to open the door. After the execution of the target instruction, it may cause harm to the driver and other people in the vehicle. It can be determined that the target instruction cannot be executed at present.

[0071] In one embodiment, it is possible to first determine whether the target instruction can be executed. If the current vehicle state meets the execution conditions, the target instruction can be executed; if the current vehicle state does not meet the execution conditions, the vehicle terminal can output a prompt to inform the user that the conditions for executing the target instruction are not currently met.

[0072] In the interactive control method provided in the embodiment of the present application, a voice request can be obtained; a target instruction can be determined based on the text information corresponding to the voice request and the instruction label of the preset instruction; and the target instruction can be executed. Among them, the instruction label of the preset instruction includes a label body and at least one extended label, the label body is determined based on the content displayed in the target interface, at least one extended label is used to describe the label body through different text information, and at least one preset instruction corresponds to the content displayed in the target interface. Since different extended tags corresponding to each instruction label can describe the label body through different text information, there can be multiple description methods for the same instruction label. In the process of determining the target instruction, the text information corresponding to the voice request can be matched with the multiple description methods, so that the target instruction indicated by the text information corresponding to the voice request can be determined more accurately, which improves the generalized understanding of the voice request by the vehicle-mounted terminal, and further allows the vehicle-mounted terminal to execute the target instruction more accurately.

[0073] In one embodiment, a target instruction is determined based on text information corresponding to a voice request and an instruction tag of each preset instruction in at least one preset instruction, including: matching the text information corresponding to the voice request and the tag body and / or extended tag of each preset instruction in at least one preset instruction to determine the instruction tag with the highest degree of matching, and the instruction corresponding to the instruction tag with the highest degree of matching is the target instruction.

[0074] It should be noted that the matching degree may be the text similarity between the text information corresponding to the voice request and the text information of the tag body, or may be the text similarity between the text information corresponding to the voice request and the text information of the extended tag.

[0075] By comparing the text similarities, the instruction label that has the highest degree of matching with the text information corresponding to the voice request can be determined, and the instruction corresponding to the instruction label is the target instruction.

[0076] In one embodiment, the text similarity may be determined based solely on matching the text information corresponding to the voice request with the text information of the tag body.

[0077] In another embodiment, the text information corresponding to the voice request can be matched with the text information of the tag body. If the tag body with the highest similarity is matched, the text similarity can be determined. If the tag body with the highest similarity is not matched, the text information corresponding to the voice request can be matched with the text information of the extended tag to determine the text similarity.

[0078] In another embodiment, the text similarity may be determined based solely on matching the text information corresponding to the voice request with the text information of the extended tag.

[0079] It should be noted that the above-mentioned target instruction can be determined by selecting the above-mentioned multiple ways of determining text similarity according to actual needs, and no specific limitation is made here.

[0080] In the interactive control method provided in the embodiments of the present application, a match is performed based on the text information corresponding to the voice request and the tag body and / or extended tag of each preset instruction in at least one preset instruction to determine the instruction tag with the highest degree of matching. The instruction corresponding to the instruction tag with the highest degree of matching is the target instruction. Specifically, by matching the text information corresponding to the voice request with the tag body and / or extended tag of each preset instruction in at least one preset instruction, the target instruction can be determined quickly, accurately, and comprehensively, thereby improving the accuracy and diversity of the target instruction determination.

[0081] The following explains one feasible implementation method of determining the instruction tag with the highest matching degree provided in the embodiment of the present application.

[0082] Figure 4 For a flow chart of determining the instruction tag with the highest matching degree provided in the embodiment of this application, please refer to Figure 4 , matching the text information corresponding to the voice request and the tag body and / or extended tag of each preset instruction in at least one preset instruction to determine the instruction tag with the highest matching degree, including:

[0083] S410: Matching is performed based on text information corresponding to the voice request and a tag body of each preset instruction in at least one preset instruction.

[0084] S420: If there is no tag body with a matching degree greater than or equal to a preset threshold, determine an instruction tag with the highest matching degree according to the extended tag of each preset instruction in the at least one preset instruction.

[0085] It should be noted that Figure 4 The matching method provided in the embodiment shown in is an implementation method for matching the tag body and the extended tag in sequence. It can first be matched according to the text information corresponding to the voice request and the tag body of each preset instruction in at least one preset instruction.

[0086] If a tag body with a matching degree greater than or equal to a preset threshold is matched, the instruction corresponding to the tag body with the greatest matching degree can be used as the target instruction; if a tag body with a matching degree greater than or equal to the preset threshold is not matched, the instruction tag with the greatest matching degree can be determined based on the extended tag of each preset instruction in at least one preset instruction.

[0087] Among them, the tag body and the extended tag are both composed of text, and the text information corresponding to the voice request is also a piece of text. A text-to-text matching method can be used to determine whether there is a match.

[0088] Taking the comparison between the text information corresponding to the voice request and the label body as an example, for example, the text information corresponding to the voice request can be used as a benchmark to compare the text similarity between the text information corresponding to the voice request and the label body of each preset instruction in at least one preset instruction.

[0089] In the process of comparing text similarity, taking the text information corresponding to the voice request as "turn on the air conditioner" as an example, the voice request can be first split into multiple keywords, such as "turn on" and "air conditioner", and then each word of each keyword can be compared in turn to see if it is the same as the keyword in the label body, and then the text similarity can be determined based on the proportion of the same words.

[0090] For example, the text information corresponding to the voice request is "turn on the air conditioner", and the tag body is "turn on the air conditioner". It can be determined that three of the four characters are the same, and the text similarity can be determined to be 75%.

[0091] It should be noted that the above example uses the comparison of the text information corresponding to the voice request with the tag body as an example for explanation. In the actual implementation process, a similar comparison method can also be used when comparing the text information corresponding to the voice request with the extended tag, and no specific restrictions are made here.

[0092] In one embodiment, the preset threshold can be set according to actual needs, for example, set to 90% or 80%, etc., without specific limitation here. Different preset thresholds can be determined according to the number of words in the text information corresponding to the voice request, or a fixed threshold can be set.

[0093] In the interactive control method provided in the embodiments of the present application, a match can be performed based on the text information corresponding to the voice request and the tag body of each preset instruction in at least one preset instruction; if there is no tag body with a matching degree greater than or equal to a preset threshold, the instruction tag with the highest matching degree is determined based on the extended tag of each preset instruction in the at least one preset instruction. Specifically, by first comparing the text information corresponding to the voice request with the tag body, and then comparing the text information corresponding to the voice request with the extended tag if there is no tag body with a matching degree greater than or equal to the preset threshold, the matching efficiency can be improved, the number of unnecessary matches can be reduced, and the speed of determining the target instruction can be increased.

[0094] Another feasible implementation method for determining the instruction tag with the highest matching degree provided in the embodiment of the present application is explained below.

[0095] Figure 5 For another flow chart of determining the instruction tag with the highest matching degree provided in the embodiment of the present application, please refer to Figure 5 , matching the text information corresponding to the voice request and the tag body and / or extended tag of each preset instruction in at least one preset instruction to determine the instruction tag with the highest matching degree, including:

[0096] S510: Matching is performed based on text information corresponding to the voice request and a tag body of each preset instruction in at least one preset instruction.

[0097] S520: Matching is performed according to the text information corresponding to the voice request and the extended tag of each preset instruction in the at least one preset instruction.

[0098] S530: The tag body or the instruction corresponding to the extended tag with the highest matching degree is used as the target instruction.

[0099] It should be noted that Figure 5 The matching method provided in the embodiment shown in is an implementation method for matching the tag body and the extended tag separately, and can be matched according to the text information corresponding to the voice request, and the tag body and extended tag of each preset instruction in at least one preset instruction.

[0100] In one embodiment, the tag body and the extended tag may be matched synchronously to determine the matching degree of all tag bodies and extended tags, and the instruction corresponding to the tag body or extended tag with the highest matching degree is used as the target instruction.

[0101] Optionally, if the highest matching degree is one of the tag bodies, the instruction corresponding to the tag body can be used as the target instruction; if the highest matching degree is one of the extended tags, the instruction corresponding to the extended tag can be used as the target instruction.

[0102] In the interactive control method provided in the embodiments of the present application, a match can be performed based on the text information corresponding to the voice request and the tag body of each preset instruction in at least one preset instruction; a match can be performed based on the text information corresponding to the voice request and the extended tag of each preset instruction in at least one preset instruction; and the instruction corresponding to the tag body or extended tag with the highest degree of matching is selected as the target instruction. By synchronously comparing the text information corresponding to the voice request with the tag body and the extended tag, the tag body or extended tag with the highest degree of matching can be more comprehensively and accurately determined, thereby more accurately obtaining the target instruction.

[0103] The following explains another feasible implementation process of determining the instruction tag with the highest matching degree provided in the embodiment of the present application.

[0104] Figure 6 For another flow chart of determining the instruction tag with the highest matching degree provided in the embodiment of the present application, please refer to Figure 6 , matching the text information corresponding to the voice request and the tag body and / or extended tag of each preset instruction in at least one preset instruction to determine the instruction tag with the highest matching degree, including:

[0105] S610: Matching is performed according to the text information corresponding to the voice request and the tag body and / or extended tag of each preset instruction in at least one preset instruction to obtain one or more candidate instruction tags with the highest matching degree.

[0106] It should be noted that, through the above matching method, the matching degree between the text information corresponding to the voice request and each tag body and extended tag can be determined, thereby obtaining one or more candidate instruction tags with the highest matching degree.

[0107] Among them, since the matching degree is based on whether the text is the same, it may happen that the text information corresponding to the voice request has the same matching degree with multiple tag bodies or extended tags and all of them have the highest matching degree. In this case, it can be determined based on the number of candidate instruction tags with the highest matching degree.

[0108] S620: When there are multiple candidate instruction tags with the highest matching degree, determine the candidate instruction tag with the shortest tag length among the multiple candidate instruction tags with the highest matching degree as the instruction tag with the highest matching degree.

[0109] It should be noted that when there are multiple instruction labels with the highest degree of matching, the label length of the instruction label can be further determined, where the label length can be expressed by the number of corresponding characters. For example: if the number of characters corresponding to the label "turn on the air conditioner" is 4, then the label length of the label can be determined to be 4.

[0110] Optionally, the tag length of each candidate instruction tag with the highest matching degree may be determined, and then the candidate instruction tag with the shortest tag length may be used as the instruction tag with the highest matching degree.

[0111] In one implementation, after obtaining the voice request, the method further includes: when there are multiple candidate instruction tags with the highest matching degree and the shortest tag length, outputting prompt information, where the prompt information is used to indicate that the request corresponding to the voice request is unclear.

[0112] It should be noted that if there is only one candidate instruction label with the highest degree of match and the shortest label length, the instruction corresponding to the candidate instruction label can be used as the above-mentioned target instruction; if there are multiple candidate instruction labels with the highest degree of match and the shortest label length, a prompt message can be output to prompt the user that the voice request is unclear and the voice request needs to be re-entered.

[0113] For example, the prompt message may be, for example, "Your request is not clear, please repeat it."

[0114] S630: When there is only one instruction tag to be selected with the highest matching degree, determine the instruction tag to be selected as the instruction tag with the highest matching degree.

[0115] It should be noted that if there is only one instruction tag with the highest degree of matching, there is no need to calculate the tag length of the instruction tag. The instruction tag can be used as the instruction tag with the highest degree of matching, that is, the instruction corresponding to the instruction tag can be used as the target instruction.

[0116] In the interactive control method provided in the embodiment of the present application, when there are multiple candidate instruction tags with the highest degree of matching, the candidate instruction tag with the shortest label length among the multiple candidate instruction tags with the highest degree of matching can be determined as the instruction tag with the highest degree of matching; when there are multiple candidate instruction tags with the highest degree of matching and the shortest label length, a prompt message is output, and the prompt message is used to indicate that the request corresponding to the voice request is unclear. Among them, when multiple candidate instruction tags are determined, the most suitable target instruction for execution can be determined by comparing the label lengths. When the degree of matching is the same, the shorter the label length, the more similar the text is. In this way, the instruction tag with the highest degree of matching can be screened out in the case of multiple candidate instruction tags, so that the only instruction tag with the highest degree of matching can be accurately determined, so that the corresponding target instruction can be accurately executed. In addition, in the case where the target instruction cannot be executed, a prompt message can also be generated to prompt the user why the execution cannot be performed.

[0117] It should be noted that the interactive control method provided in the embodiment of the present application can realize the comparison between the text information corresponding to the voice request and the instruction label by setting a matching tree. The matching process is explained below using a specific matching tree as an example.

[0118] Figure 7 For a schematic diagram of the target matching tree provided in the embodiment of this application, please refer to Figure 7 , Figure 7 The matching tree shown can be based on Figure 2 A target matching tree is obtained by the interface shown, and multiple matching nodes can be set in the target matching tree. For example, the highest-level node can be an assisted driving node, which can include multiple sub-nodes, such as a starting obstacle avoidance assistance node, a blind spot safety assistance node, a door opening warning node, and a rear collision warning node.

[0119] For the above four nodes, each node may include two sub-nodes, for example: open or close.

[0120] In addition, the target matching tree may also include other nodes, such as an exit assisted driving interface node.

[0121] In the target matching tree, each node may correspond to one or more instruction labels. For example, taking the first node at the bottom layer as an example, its corresponding instruction label may be "turn on obstacle avoidance assistance when starting."

[0122] It should be noted that Figure 7The target matching tree shown in the figure is only a part of it. In actual implementation, different interfaces may correspond to different matching trees, and the matching nodes in the matching tree may also be increased or decreased according to actual needs, which is not specifically limited here.

[0123] During the matching process, the target matching tree may be traversed for matching, that is, the instruction label corresponding to each matching node may be matched in turn, thereby determining the instruction label with the highest matching degree.

[0124] The following explains a feasible implementation process of determining the instruction label with the highest matching degree based on the target matching tree provided in an embodiment of the present application.

[0125] In one embodiment, matching is performed based on the text information corresponding to the voice request and the tag body and / or extended tag of each preset instruction in at least one preset instruction to determine the instruction tag with the highest degree of matching, including: matching the text information corresponding to the voice request with the text information of each matching node in the target matching tree to determine the degree of matching of each matching node that matches the text information corresponding to the voice request, wherein each matching node in the target matching tree corresponds to at least one instruction tag, and the instruction tag corresponding to the matching node with the highest degree of matching is the instruction tag with the highest degree of matching.

[0126] Optionally, after obtaining the voice request, each matching node in the target matching tree can be traversed to match the text information of the voice request with the text information of the matching node, where the text information of the matching node refers to at least one instruction tag corresponding to the matching node.

[0127] After matching in sequence through traversal, the matching degree corresponding to the instruction label in each matching node in the target matching tree can be determined.

[0128] Then, the instruction label corresponding to the matching node with the highest matching degree can be determined, and the instruction corresponding to the instruction label is used as the target instruction.

[0129] In the interactive control method provided in the embodiments of the present application, the text information corresponding to the voice request can be matched with the text information of each matching node in the target matching tree to determine the matching degree of each matching node that matches the text information corresponding to the voice request, wherein each matching node in the target matching tree corresponds to at least one instruction tag, and the instruction tag corresponding to the matching node with the highest matching degree is the instruction tag with the highest matching degree. By traversing multiple matching nodes in the target matching tree, the instruction tag corresponding to each matching node can be matched in turn, thereby comprehensively and quickly matching the text information corresponding to the voice request and accurately determining the target instruction.

[0130] The following explains the specific implementation process of matching the text information corresponding to the voice request with each matching node provided in the embodiments of the present application.

[0131] Figure 8 For a flow chart of determining the matching degree of a matching node based on a matching tree provided in an embodiment of the present application, please refer to Figure 8 , determining the matching degree of each matching node that matches the text information corresponding to the voice request, including:

[0132] S810: Determine the degree of matching between the text information corresponding to the voice request and the target matching node based on the text similarity between the text information corresponding to the voice request and the instruction label in each target matching node.

[0133] The target matching node is any matching node in the matching tree.

[0134] It should be noted that text similarity refers to the text matching between the text information corresponding to the voice request and the instruction label. For example, the match can be determined starting from the first word of the text information corresponding to the voice request. If the text information corresponding to the voice request is N words, and the number of identical words matched is M, then the text similarity is M / N.

[0135] It should be noted that the matched identical characters refer to two identical characters in the same position. For example, if the first character is the same, the first character may be the matched identical character; if the second character is different, the second character may not be the matched identical character.

[0136] The text similarity can be determined in the above manner, and the text similarity can be used to represent the degree of matching between the text information corresponding to the voice request and the target matching node.

[0137] Based on this approach, the matching degree between the text information corresponding to the voice request and each target matching node can be compared, thereby determining the target matching node with the highest matching degree.

[0138] Accordingly, the method further includes:

[0139] S820: The instruction corresponding to the instruction label in the target matching node with the highest matching degree is used as the target instruction.

[0140] It should be noted that after determining the degree of matching between the text information corresponding to the voice request and each target matching node in the above manner, the degree of matching with multiple target matching nodes can be sorted in a sequential manner to obtain the target matching node with the highest degree of matching.

[0141] In one embodiment, the instruction corresponding to the target matching node with the highest matching degree may be used as the target instruction.

[0142] For example, Figure 7 Taking the matching tree shown as an example, if it is determined that the target matching node with the highest matching degree is the exit assisted driving interface node, the instruction corresponding to the instruction label in the node can be used as the target instruction.

[0143] Another feasible implementation process of matching the text information corresponding to the voice request with each matching node provided in an embodiment of the present application is explained below.

[0144] Figure 9 This is another flow chart of determining the matching degree of matching nodes based on the matching tree provided in the embodiment of the present application. Please refer to Figure 9 , determining the matching degree of each matching node that matches the text information corresponding to the voice request, including:

[0145] S910: Determine the degree of matching between the text information corresponding to the voice request and the target matching node based on the text similarity between the text information corresponding to the voice request and the instruction label in each target matching node, and the text similarity between the text information corresponding to the voice request and the instruction label in at least one first-level parent node of each target matching node.

[0146] Figure 8 The method shown in is a solution for determining the degree of matching based on the text similarity of the target matching node itself. In the actual implementation process, in addition to using this method, other methods can also be used to determine the degree of matching between the text information corresponding to the voice request and the target matching node.

[0147] It should be noted that the degree of matching between the text information corresponding to the voice request and the target matching node may also be determined by combining the text similarity of the target matching node and the text similarity of at least one first-level parent node of the target matching node.

[0148] For example, Figure 7 Taking the first node "Start obstacle avoidance assistance node" at the bottom shown in the figure as an example, in the process of calculating the matching degree, not only the text similarity between the text information corresponding to the voice request and the instruction label of the target matching node can be calculated, but also the text similarity between the text information corresponding to the voice request and the instruction label of at least one first-level parent node of the target matching node can be calculated.

[0149] For example, the text similarity between the two nodes "Start Obstacle Avoidance Assist Node" and "Start Obstacle Avoidance Assist Node" can be calculated to determine the degree of matching between the text information corresponding to the voice request and the target matching node.

[0150] It should be noted that the degree of matching between the text information corresponding to the voice request and the target matching node can be determined by setting a text similarity weight.

[0151] Taking the above two nodes as an example, the child node can be given a higher weight, for example, 0.8, and the parent node can be given a lower weight, for example, 0.2, so as to calculate the matching degree between the text information corresponding to the voice request and the target matching node.

[0152] Optionally, you can use Figure 8 or Figure 9 Any of the above methods is used to calculate the matching degree between the text information corresponding to the voice request and the target matching node.

[0153] It should be noted that a corresponding configuration file can be set for the calculation method of the matching degree of each target matching node, and the calculation method can be stored in the configuration file. When calculating the corresponding node, the corresponding calculation method can be used to determine the matching degree of the target matching node.

[0154] In the interactive control method provided in the embodiment of the present application, the degree of matching between the text information corresponding to the voice request and the target matching node can be determined based on the text similarity between the text information corresponding to the voice request and the instruction label in each target matching node; or, the degree of matching between the text information corresponding to the voice request and the instruction label in each target matching node can be determined based on the text similarity between the text information corresponding to the voice request and the instruction label in each target matching node, and the text similarity between the text information corresponding to the voice request and the instruction label in at least one first-level parent node of each target matching node. Among them, the degree of matching of each target matching node can be calculated in different ways according to the actual structural requirements of different target matching nodes, so that a more accurate calculation result of the degree of matching can be obtained, thereby improving the accuracy of determining the target instruction.

[0155] It should be noted that Figure 7 The target matching tree shown is only one example. In actual implementation, a variety of methods can be used to obtain the matching tree used for matching.

[0156] In one embodiment, before matching the text information corresponding to the voice request with the text information of each matching node in the target matching tree, the method further includes: determining the target matching tree according to the interface identifier of the target interface.

[0157] Among them, each interface in the vehicle terminal corresponds to a different matching tree.

[0158] It should be noted that multiple matching trees can be preset in the vehicle terminal, each target interface can correspond to a matching tree, and the mapping relationship between the interface identifier of each target interface and the matching tree can be pre-configured. After obtaining the interface identifier of the target interface, the corresponding matching tree can be determined based on the mapping relationship, so that the matching tree corresponding to the interface identifier of the target interface can be used as the above-mentioned target matching tree.

[0159] Optionally, in actual implementation, in addition to being called from preset matching tree types, the matching tree may also be established in real time.

[0160] In one embodiment, before matching the text information corresponding to the voice request with the text information of each matching node in the target matching tree, the method further includes: establishing a target matching tree based on a configuration file of each instruction tag corresponding to the content displayed in the target interface.

[0161] The configuration file of each instruction tag includes an executable interface of the instruction corresponding to the instruction tag.

[0162] It should be noted that each instruction tag can set a corresponding configuration file, and each configuration file can set the interface for the adaptive execution of the instruction corresponding to the instruction tag. For example, the executable interface of the instruction corresponding to each instruction tag can be recorded, and the instructions corresponding to which instruction tags are executable can be determined based on the interface identifier of the current interface, so that a target matching tree can be established based on the instruction tags of these executable instructions.

[0163] It should be noted that in addition to recording the executable interface of the instructions, the configuration file can also record the relationship between instructions. Through these relationships, the node position relationship of different instruction tags in the matching tree can be determined, so that the target matching tree can be established more conveniently.

[0164] In the interactive control method provided in the embodiments of the present application, a target matching tree can be determined based on the interface identifier of the target interface. Alternatively, a target matching tree can be established based on the configuration file for each instruction tag corresponding to the content displayed in the target interface. In actual implementation, different methods can be used to determine or establish the target matching tree based on actual needs, thereby more accurately and quickly determining the degree of matching and further quickly and accurately determining the target instruction.

[0165] In one embodiment, the vehicle-mounted terminal can convert the display content of the display screen into scene data, which is in JSON format. The scene elements in the scene data correspond to the controls on the screen. After receiving the voice request, the scene data can be parsed.

[0166] Since the scene data is in JSON format and nested, it naturally contains tree structure information and parent-child relationships. Based on these relationships, the matching tree can be quickly constructed using the original JSON format.

[0167] Optionally, the matching tree can be established recursively. For example, a tree structure object can be obtained first to construct a root node, and scene elements on the node can be recursively obtained in a depth-first manner until no scene elements are found. Then, the matching tree can be established by returning to the previous layer and continuing to traverse other scene elements until the entire tree structure object is traversed.

[0168] It should be noted that, in the process of establishing the matching tree, corresponding matching nodes may be established by inserting nodes, and each matching node may also include information such as the path of the node in the matching tree.

[0169] During the matching process, since the structural information of the scene data cannot determine the depth of the node, there may be nodes with very deep depths. It is unreasonable to match completely from the root node to the leaf node. Therefore, during the matching process, the text similarity with the target matching node can be determined to determine the degree of matching, or the text similarity with the target matching node and at least one first-level parent node of the target matching node can be matched to determine the degree of matching.

[0170] In one embodiment, for the above configuration file, corresponding generalization capabilities can be set according to actual needs, so as to generalize the corresponding instruction tags in the target interface, and add corresponding generalization parameters to the configuration file. Through the generalization parameters, for example, the number of extended tags can be configured.

[0171] According to different generalization requirements, generalization configuration can be implemented at the global level, application level, and page level. The details are as follows:

[0172] When the application scenario configuration corresponding to the configuration file is empty, the relevant configuration of the instruction tag can be effective globally; when the application scenario configuration corresponding to the configuration file is a unified prefix name of the application, the relevant configuration of the instruction tag can be effective at the application level, such as music, navigation, etc.; when the application scenario configuration corresponding to the configuration file is a specific page, the relevant configuration of the instruction tag can be effective for a unique interface.

[0173] In one embodiment, the extended tag may be set after the tag body by masking.

[0174] It should be understood that, although the steps in the above-mentioned flowcharts are shown in sequence according to the instructions of the arrows, these steps are not necessarily performed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the above-mentioned flowcharts may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0175] Based on the foregoing embodiments, an embodiment of the present application provides an interactive control device, which includes the modules included and the units included in each module, and can be implemented by a processor; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.

[0176] Figure 10 This is a schematic diagram of the structure of the interactive control device provided in the embodiment of the present application. Please refer to Figure 10 , the device includes: an acquisition module 1010, a determination module 1020 and an execution module 1030;

[0177] An acquisition module 1010 is configured to acquire a voice request;

[0178] Determining module 1020, configured to determine a target instruction based on the text information corresponding to the voice request and an instruction tag of a preset instruction, wherein the instruction tag of the preset instruction includes a tag body and at least one extended tag, wherein the tag body is determined based on the content displayed on the target interface, and the at least one extended tag is used to describe the tag body using different text information, and the at least one preset instruction corresponds to the content displayed on the target interface;

[0179] The execution module 1030 is used to execute the target instruction.

[0180] In one embodiment, the determination module 1020 is specifically used to match the text information corresponding to the voice request and the tag body and / or extended tag of each preset instruction in at least one preset instruction to determine the instruction tag with the highest matching degree, and the instruction corresponding to the instruction tag with the highest matching degree is the target instruction.

[0181] In one embodiment, the determination module 1020 is specifically used to match the text information corresponding to the voice request and the tag body of each preset instruction in at least one preset instruction; if there is no tag body with a matching degree greater than or equal to a preset threshold, the instruction tag with the highest matching degree is determined based on the extended tag of each preset instruction in at least one preset instruction.

[0182] In one embodiment, the determination module 1020 is specifically configured to determine, when there are multiple candidate instruction tags with the highest matching degree, the candidate instruction tag with the shortest tag length among the multiple candidate instruction tags with the highest matching degree as the instruction tag with the highest matching degree.

[0183] In one embodiment, the determination module 1020 is further configured to output a prompt message when there are multiple candidate instruction tags with the highest matching degree and the shortest tag length, where the prompt message is used to indicate that the request corresponding to the voice request is unclear.

[0184] In one embodiment, the determination module 1020 is specifically used to match the text information corresponding to the voice request with the text information of each matching node in the target matching tree, and determine the matching degree of each matching node that matches the text information corresponding to the voice request, wherein each matching node in the target matching tree corresponds to at least one instruction tag, and the instruction tag corresponding to the matching node with the highest matching degree is the instruction tag with the highest matching degree.

[0185] In one embodiment, the determination module 1020 is specifically used to determine the degree of matching between the text information corresponding to the voice request and the target matching node based on the text similarity between the text information corresponding to the voice request and the instruction label in each target matching node, where the target matching node is any matching node in the matching tree; or, to determine the degree of matching between the text information corresponding to the voice request and the target matching node based on the text similarity between the text information corresponding to the voice request and the instruction label in each target matching node, and the text similarity between the text information corresponding to the voice request and the instruction label in at least one first-level parent node of each target matching node.

[0186] In one embodiment, the determination module 1020 is further configured to determine a target matching tree according to the interface identifier of the target interface, where each interface in the vehicle-mounted terminal corresponds to a different matching tree.

[0187] In one embodiment, the determination module 1020 is further configured to establish a target matching tree based on a configuration file of each instruction tag corresponding to the content displayed in the target interface, wherein the configuration file of each instruction tag includes an executable interface of the instruction corresponding to the instruction tag.

[0188] In the interactive control device provided in the embodiment of the present application, a voice request can be obtained; a target instruction can be determined based on the text information corresponding to the voice request and the instruction label of the preset instruction; and the target instruction can be executed. Among them, the instruction label of the preset instruction includes a label body and at least one extended label, the label body is determined based on the content displayed in the target interface, at least one extended label is used to describe the label body through different text information, and at least one preset instruction corresponds to the content displayed in the target interface. Since different extended tags corresponding to each instruction label can describe the label body through different text information, there can be multiple description methods for the same instruction label. In the process of determining the target instruction, the text information corresponding to the voice request can be matched with the multiple description methods, so that the target instruction indicated by the text information corresponding to the voice request can be determined more accurately, which improves the generalized understanding of the voice request by the vehicle-mounted terminal, and further allows the vehicle-mounted terminal to execute the target instruction more accurately.

[0189] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of this application, please refer to the description of the method embodiment of this application for understanding.

[0190] It should be noted that in the embodiments of this application Figure 10 The division of modules in the interactive control device shown is schematic and is only a logical functional division. In actual implementation, other division methods may be used. In addition, the functional units in the various embodiments of the present application can be integrated into a processing unit, or can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or software functional units. It can also be implemented in the form of a combination of software and hardware.

[0191] It should be noted that, in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling an electronic device to execute all or part of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0192] Figure 11For a schematic diagram of the structure of the computer device provided in the embodiment of this application, please refer to Figure 11 , an embodiment of the present application provides a computer device, which may be the above-mentioned vehicle-mounted terminal, and its internal structure diagram may be as shown in FIG. Figure 11 As shown. The computer device includes a processor 1120, a memory, and a network interface 1140 connected via a system bus 1110. The processor 1120 of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium 1131 and an internal memory 1132. The non-volatile storage medium 1131 stores an operating system, a computer program, and a database. The internal memory 1132 provides an environment for the operation of the operating system and computer program in the non-volatile storage medium 1131. The database of the computer device is used to store data. The network interface 1140 of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor 1120, the above method is implemented.

[0193] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method provided in the above embodiment are implemented.

[0194] An embodiment of the present application provides a computer program product containing instructions, which, when executed on a computer, enables the computer to execute the steps of the method provided in the above method embodiment.

[0195] Those skilled in the art will understand that Figure 11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0196] In one embodiment, the interactive control device provided by the present application can be implemented in the form of a computer program. The computer program can be used in Figure 11 The computer device is operated on the computer device shown. The memory of the computer device can store various program modules that constitute the above-mentioned device. The computer program composed of each program module enables the processor to execute the steps of the method of each embodiment of the present application described in this specification.

[0197] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium, and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0198] It should be understood that "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments. The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other. For the sake of brevity, they will not be repeated here.

[0199] The term "and / or" in this article is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, object A and / or object B can mean: object A exists alone, object A and object B exist at the same time, and object B exists alone.

[0200] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0201] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical or other forms.

[0202] The modules described above as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules; they may be located in one place or distributed across multiple network units; some or all of the modules may be selected according to actual needs to achieve the purpose of this embodiment.

[0203] In addition, all functional modules in the embodiments of the present application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the above-mentioned integrated modules can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0204] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.

[0205] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks or optical disks.

[0206] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0207] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0208] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0209] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. An interactive control method, characterized in that: Applied to a vehicle-mounted terminal, the vehicle-mounted terminal includes a display screen, and a target interface is displayed on the display screen. The method includes: Get voice request; Determining a target instruction based on the text information corresponding to the voice request and an instruction tag of a preset instruction, wherein the instruction tag of the preset instruction includes a tag body and at least one extended tag, the tag body being determined based on the content displayed on the target interface, the at least one extended tag being used to describe the tag body using different text information, and the at least one preset instruction corresponding to the content displayed on the target interface; Execute the target instruction.

2. The method according to claim 1, characterized in that The determining the target instruction according to the text information corresponding to the voice request and the instruction label of the preset instruction includes: Matching is performed based on the text information corresponding to the voice request and the tag body and / or extended tag of each preset instruction in at least one preset instruction to determine the instruction tag with the highest matching degree, and the instruction corresponding to the instruction tag with the highest matching degree is the target instruction.

3. The method according to claim 2, characterized in that The matching of the text information corresponding to the voice request and the tag body and / or the extended tag of each preset instruction in at least one preset instruction to determine the instruction tag with the highest matching degree includes: Matching based on text information corresponding to the voice request and a tag body of each preset instruction in at least one preset instruction; In the case that there is no tag body with a matching degree greater than or equal to a preset threshold, an instruction tag with the highest matching degree is determined according to the extended tag of each preset instruction in the at least one preset instruction.

4. The method according to claim 2, characterized in that The matching of the text information corresponding to the voice request and the tag body and / or the extended tag of each preset instruction in at least one preset instruction to determine the instruction tag with the highest matching degree includes: In the case that there are multiple candidate instruction tags with the highest matching degree, the candidate instruction tag with the shortest tag length among the multiple candidate instruction tags with the highest matching degree is determined as the instruction tag with the highest matching degree.

5. The method according to claim 4, characterized in that After obtaining the voice request, the method further includes: In the case where there are multiple candidate instruction labels with the highest matching degree and the shortest label length, prompt information is output, where the prompt information is used to indicate that the request corresponding to the voice request is unclear.

6. The method according to claim 2, characterized in that The matching of the text information corresponding to the voice request and the tag body and / or the extended tag of each preset instruction in at least one preset instruction to determine the instruction tag with the highest matching degree includes: According to the text information corresponding to the voice request, the text information of each matching node in the target matching tree is matched to determine the matching degree of each matching node that matches the text information corresponding to the voice request, wherein each matching node in the target matching tree corresponds to at least one instruction tag, and the instruction tag corresponding to the matching node with the highest matching degree is the instruction tag with the highest matching degree.

7. The method according to claim 6, characterized in that Determining the matching degree of each matching node that matches the text information corresponding to the voice request includes: Determining the degree of matching between the text information corresponding to the voice request and the target matching node based on the text similarity between the text information corresponding to the voice request and the instruction label in each target matching node, where the target matching node is any matching node in the matching tree; or The degree of matching between the text information corresponding to the voice request and the target matching node is determined based on the text similarity between the text information corresponding to the voice request and the instruction label in each target matching node, and the text similarity between the text information corresponding to the voice request and the instruction label in at least one first-level parent node of each target matching node.

8. The method according to claim 6, characterized in that Before matching the text information corresponding to the voice request with the text information of each matching node in the target matching tree, the method further includes: The target matching tree is determined according to the interface identifier of the target interface, and each interface in the vehicle-mounted terminal corresponds to a different matching tree.

9. The method according to claim 6, characterized in that Before matching the text information corresponding to the voice request with the text information of each matching node in the target matching tree, the method further includes: The target matching tree is established according to the configuration file of each instruction tag corresponding to the displayed content in the target interface, and the configuration file of each instruction tag includes an executable interface of the instruction corresponding to the instruction tag.

10. An interactive control device, characterized in that: Applied to a vehicle-mounted terminal, the vehicle-mounted terminal includes a display screen, and a target interface is displayed on the display screen. The device includes: an acquisition module, a determination module, and an execution module; The acquisition module is used to acquire the voice request; The determining module is configured to determine a target instruction based on the text information corresponding to the voice request and an instruction tag of a preset instruction, wherein the instruction tag of the preset instruction includes a tag body and at least one extended tag, the tag body being determined based on content displayed on the target interface, the at least one extended tag being configured to describe the tag body using different text information, and the at least one preset instruction corresponding to the content displayed on the target interface; The execution module is used to execute the target instruction.

11. A computer device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 9 are implemented.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Terminal control method based on voice and voice control device

    CN105869643A

  • Voice control system and method thereof

    CN107293298A

  • Software development method, apparatus, and computer readable storage medium

    CN109271157A

  • Terminal voice control method and device, storage medium and electronic equipment

    CN117059100A

  • Voice interaction method, vehicle, device and storage medium

    CN118173086A