Vehicle control method and device, vehicle and storage medium

CN122501142APending Publication Date: 2026-08-04IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2026-04-23
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

[0004]本发明提供一种车辆控制方法、装置、车辆及存储介质,用以解决现有技术中当语音指令中使用的关系代词或称呼词存在歧义时,容易造成个性化车控指令执行错误或控制失效的缺陷,实现在关系指代存在歧义时,通过多模态特征对服务对象身份与物理位置进行精准识别,进而提升车辆个性化代理控制的准确性、鲁棒性、灵活性和交互体验

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122501142A_ABST
    Figure CN122501142A_ABST
Patent Text Reader

Abstract

The application provides a vehicle control method and device, a vehicle and a storage medium. The method comprises the following steps: determining first identity information in a relationship topology network according to a voiceprint feature and a sound source position of voice data in a target vehicle; performing semantic analysis on the voice data to obtain control description information of the target vehicle and object description information of a service object; when relationship description information in the object description information corresponds to relationship reference ambiguity, determining second identity information and target position information according to appearance feature description information and / or position description information in the object description information, and the first identity information and the relationship description information; and generating a vehicle control instruction for the service object according to the first identity information, the second identity information, the target position information and the control description information to control the target vehicle. The application improves the accuracy, robustness, flexibility and interactive experience of vehicle personalized agent control when there is ambiguity in the relationship reference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent vehicle control technology, and in particular to a vehicle control method, device, vehicle, and storage medium. Background Technology

[0002] With the rapid development of automotive intelligence, especially the continued growth in market demand for multi-seat vehicles, family travel scenarios are becoming increasingly common. During family trips, there are often multiple passengers in the car, including users of different ages such as the elderly and children, each with varying abilities and needs regarding vehicle functions. Therefore, how to provide personalized vehicle control services for different passengers is a crucial issue that the industry urgently needs to address.

[0003] Traditional technologies often rely on parsing specific terms of address in conversations for vehicle control. This mechanism is highly dependent on preset, precise terms of address or the target passenger's own active voice verification. In multi-occupant scenarios, when the relative pronouns or terms of address used in voice commands are ambiguous, they cannot be effectively distinguished, leading to an inability to accurately determine the actual service recipient and their physical location. This results in errors in the execution of personalized vehicle control commands or control failure. Summary of the Invention

[0004] This invention provides a vehicle control method, device, vehicle, and storage medium to address the shortcomings of existing technologies where ambiguous relational pronouns or appellations used in voice commands can easily lead to errors in the execution of personalized vehicle control commands or control failures. The invention achieves accurate identification of the service object's identity and physical location through multimodal features when relational references are ambiguous, thereby improving the accuracy, robustness, flexibility, and interactive experience of personalized vehicle agent control.

[0005] This invention provides a vehicle control method, comprising: Based on the voiceprint characteristics of the voice data inside the target vehicle and the location of the sound source of the voice data, the first identity information of the speaker corresponding to the voice data is determined in a pre-constructed relational topology network. Semantic parsing is performed on the voice data to obtain the control description information of the target vehicle and the object description information of the service object; When there is ambiguity in the relational reference corresponding to the relational description information in the object description information, the second identity information and target location information of the service object are determined based on the appearance feature description information and / or location description information in the object description information, as well as the first identity information and the relational description information; Based on the first identity information, the second identity information, the target location information, and the control description information, a vehicle control command is generated for the service object, and the target vehicle is controlled according to the vehicle control command.

[0006] According to a vehicle control method provided by the present invention, determining the second identity information and target location information of the service object based on the appearance feature description information and / or location description information in the object description information, as well as the first identity information and the relationship description information, includes: The target location information is determined based on the appearance feature description information and / or the location description information; The second identity information is determined in the relationship topology network based on the target location information, the first identity information, and the relationship description information.

[0007] According to a vehicle control method provided by the present invention, determining the second identity information in the relationship topology network based on the target location information, the first identity information, and the relationship description information includes: Based on the relationship description information and the first identity information, a target relationship subnet is determined in the relationship topology; Based on the target location information, the biometric features of the service recipient are obtained; the biometric features include at least one of vital signs, facial features, multimodal appearance features, and voiceprint features. Match the biometric characteristics of the service recipient with the feature profiles of each object in the target relationship subnet; Based on the matching results, obtain the second identity information.

[0008] According to a vehicle control method provided by the present invention, determining the target location information based on the appearance feature description information includes: Extract multimodal appearance features of each passenger in the target vehicle from the visual perception image inside the target vehicle; The multimodal appearance features of each passenger are matched with the appearance feature description information to obtain the appearance matching degree corresponding to each passenger. The target location information is determined based on the seat position information of the passenger with the highest appearance matching degree in the target vehicle.

[0009] According to a vehicle control method provided by the present invention, determining the target location information based on the appearance feature description information and the location description information includes: When there are multiple target seats corresponding to the location description information, at least one candidate seat is determined from the multiple target seats based on the seat perception data in the target vehicle and / or the speaker's gesture instructions; Extract multimodal appearance features of candidate passengers occupying each of the candidate seats from the visual perception image inside the target vehicle; The multimodal appearance features of each candidate passenger are matched with the appearance feature description information to obtain the appearance matching degree corresponding to each candidate passenger. The target location information is determined based on the seating position information of the candidate passenger with the highest appearance matching degree in the target vehicle.

[0010] According to a vehicle control method provided by the present invention, determining the target location information based on the location description information includes: If the number of target seats corresponding to the location description information is one, verify whether the target seat is occupied based on the seat perception data inside the target vehicle. When the target seat is occupied, the target location information is determined based on the physical location information corresponding to the target seat.

[0011] According to a vehicle control method provided by the present invention, the step of generating vehicle control commands for the service object based on the first identity information, the second identity information, the target location information, and the control description information includes: Based on the control description information, the control intent is obtained; Based on the control intent and the target location information, obtain the vehicle-mounted components to be controlled inside the target vehicle; Based on the first identity information and the second identity information, determine whether the speaker has proxy control authority over the vehicle-mounted component to be controlled; When the speaker has proxy control authority over the vehicle-mounted component to be controlled, the vehicle control command is generated based on the control intent, the second identity information, and the vehicle-mounted component to be controlled.

[0012] According to a vehicle control method provided by the present invention, the step of generating the vehicle control command based on the control intention, the second identity information, and the vehicle-mounted component to be controlled includes: Based on the second identity information, determine the feature profile of the service object in the relationship topology network; Based on the control parameters and control actions of the vehicle-mounted component to be controlled, the control intent, and the historical preference information in the feature profile of the service object, the control parameters and control actions of the vehicle-mounted component to be controlled are obtained. The vehicle control command is generated based on the control parameters, the control action, and the vehicle-mounted component to be controlled.

[0013] According to a vehicle control method provided by the present invention, the method further includes: Receive target feedback data from the service recipient; the target feedback data is service evaluation data input by the service recipient after the target vehicle executes the vehicle control command. Based on the target feedback data, the control records corresponding to the vehicle control commands, and the multimodal perception data within the target vehicle, the feature profiles of each object in the relational topology network are updated; the multimodal perception data includes at least one of voice data, visual perception images, and seat perception data.

[0014] The present invention also provides a vehicle control device, comprising: The first identification unit is used to determine the first identity information of the speaker corresponding to the voice data in a pre-constructed relational topology network based on the voiceprint features of the voice data in the target vehicle and the sound source location of the voice data. The parsing unit is used to perform semantic parsing on the voice data to obtain the control description information of the target vehicle and the object description information of the service object; The second identification unit is used to determine the second identity information and target location information of the service object based on the appearance feature description information and / or location description information in the object description information, as well as the first identity information and the relationship description information, when there is ambiguity in the relationship reference corresponding to the relationship description information in the object description information. The control unit is configured to generate vehicle control commands for the service object based on the first identity information, the second identity information, the target location information, and the control description information, and to control the target vehicle according to the vehicle control commands.

[0015] The present invention also provides a vehicle, including: an on-board controller and a multi-source data acquisition device; the multi-source data acquisition device includes an audio sensor, a visual sensor, a seat sensor and a gesture recognition sensor; The on-board controller is used to execute any of the vehicle control methods described above.

[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the vehicle control method as described above.

[0017] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the vehicle control method as described above.

[0018] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the vehicle control method as described above.

[0019] The vehicle control method, device, vehicle, and storage medium provided by this invention, when semantic ambiguity arises from parsing relational referents, introduces appearance feature descriptions or location descriptions for multimodal disambiguation matching. This breaks the reliance of existing technologies on a single voice or fixed address, achieving accurate identity and location confirmation in silent scenarios and with ambiguous addresses. Furthermore, by fusing the identity of the instruction initiator, the identity of the service object after precise disambiguation, and its physical location, highly structured and precisely addressed vehicle control instructions can be generated. This ensures that vehicle control operations strictly and accurately act on the correct physical partitions and service objects, avoiding misoperation. Consequently, the accuracy, robustness, and flexibility of personalized vehicle agent control are improved, providing users with a safer and more convenient personalized cockpit agent control service experience. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating the vehicle control method provided by the present invention.

[0022] Figure 2 This is a schematic diagram of the vehicle control device provided by the present invention.

[0023] Figure 3 This is a structural schematic diagram of the vehicle provided by the present invention.

[0024] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0026] All actions involving the acquisition of signal information or data in this application are carried out in accordance with the relevant data protection laws and policies of the country where the application is located, and with the authorization granted by the owner of the relevant device.

[0027] With the rapid development of automotive intelligence, especially the continued growth in market demand for multi-seat vehicles, family travel scenarios are becoming increasingly common. During family trips, there are often multiple passengers in the car, including users of different ages such as the elderly and children, each with varying abilities and needs regarding vehicle functions. Therefore, how to provide personalized vehicle control services for different passengers is a crucial issue that the industry urgently needs to address.

[0028] Traditional technologies typically control vehicles by parsing specific terms of address in conversations. However, this mechanism relies heavily on preset, precise terms of address or the target passenger's own active voice verification. In multi-occupant scenarios, when the relative pronouns or terms of address used in voice commands are ambiguous, they cannot be effectively distinguished, leading to an inability to accurately determine the actual service recipient and their physical location. This results in errors in the execution of personalized vehicle control commands or control failure.

[0029] In response, this application provides a vehicle control method to effectively and accurately identify the identity and location of service objects when there is ambiguity, thereby improving the accuracy, robustness, flexibility and interactive experience of vehicle control.

[0030] Figure 1 This is a flowchart illustrating the vehicle control method provided by the present invention.

[0031] The method provided in this application can be applied to intelligent vehicle control scenarios with multi-occupant and multimodal interaction capabilities, especially complex interaction scenarios such as family travel, intelligent cockpit partition control, and proxy vehicle control. The executing entity of the method provided in this application can be a vehicle control device, which can be an in-vehicle controller, intelligent vehicle terminal, or cloud server in a vehicle-cloud collaborative control system, etc. This embodiment does not specifically limit it in this way.

[0032] like Figure 1 As shown, the method includes steps 110, 120, 130 and 140.

[0033] Step 110: Based on the voiceprint characteristics of the voice data in the target vehicle and the location of the sound source of the voice data, determine the first identity information of the speaker corresponding to the voice data in a pre-constructed relational topology network.

[0034] Optionally, a relational topology needs to be constructed before performing step 110. Specifically, the construction process of the relational topology here mainly covers the creation of user (also known as object) feature profiles (i.e., the user registration phase), the generation of relationships between users (i.e., the relationship establishment phase), and the updating of the relational topology.

[0035] In the user profile creation process, each user can register an account through the vehicle control terminal (such as an in-vehicle infotainment system) or a related mobile application. During account registration, users are simultaneously guided to enter their basic identity information, including but not limited to name, nickname, age, gender, and work information; this embodiment does not specifically limit this. Furthermore, while entering basic information, multimodal biometrics and personalized preference information can be collected to generate a feature profile for each user node in the relational topology network based on the basic identity information, multimodal biometrics, and personalized preference information. When collecting multimodal biometrics, users can be guided to read designated text to extract and store their voiceprint features. Simultaneously, facial images of users can be collected using in-vehicle visual sensors (such as in-vehicle cameras) to extract facial features, and images of users' appearance can be collected to extract multimodal appearance feature preferences. Additionally, in-vehicle seat sensors can collect users' vital signs, such as weight. These multimodal appearance feature preferences include, but are not limited to, clothing color, clothing style, and accessory type.

[0036] The personalized preference information here includes, but is not limited to, seat preference, temperature control preference, seat angle preference, and music genre preference, etc., which are not specifically limited in this embodiment. Furthermore, this personalized preference information can be not only actively entered and set by the user through the vehicle control terminal or mobile application, but also configured by other related users in the relationship topology network authorized by the user. It can also be automatically generated in the background through continuous learning and extraction of algorithms based on the user's historical interaction habits, vehicle control execution records, and multimodal perception data during the daily operation of the target vehicle. For example, it can automatically record the user's preferred air conditioning temperature in a specific season or their preferred seat back angle during daily riding. Therefore, this multi-channel preference information acquisition and update mechanism not only greatly reduces the initial configuration threshold for users and fully considers the ease of use for special groups (such as the elderly and children), but also ensures the richness and real-time nature of personalized data in the feature profile, laying a complete data foundation for the subsequent accurate generation of personalized vehicle control commands.

[0037] In the process of generating relationships between users, registered users can actively add other registered users as related parties, thereby establishing connection edges between user nodes in the relationship topology. When establishing a connection, registered users can define the type of relationship between each other, such as parents, children, spouses, friends, etc., and set exclusive nicknames or relationship terms for the corresponding related parties, such as "son," "daughter," "Xiaoming," "baby," etc. These preset terms can serve as important reference tags for matching the identity of service objects in subsequent semantic parsing. Furthermore, this relationship topology supports bidirectional confirmation and bidirectional association mechanisms when establishing connections between users. On one hand, the initiator of the relationship needs to send a binding invitation to the target user, and the connection and its corresponding title tag only become effective in the relationship topology after the recipient (the invited party) confirms authorization through the vehicle control terminal or mobile application. On the other hand, under the bidirectional association mechanism, once the positive relationship type between the initiator and the recipient is defined (e.g., the initiator defines the recipient as "son"), it can automatically map or guide the user to set the reverse relationship type and corresponding title for the recipient towards the initiator (e.g., the associated recipient addresses the initiator as "father" or "mother"), thus forming a closed-loop bidirectional directed edge in the topology. Through this bidirectional mechanism, not only is the accuracy of the relationship network data ensured, but malicious binding of user profiles by others is also effectively prevented, fully protecting the privacy and data security of multiple users within the vehicle. Furthermore, it lays a solid data foundation for subsequent identity resolution and bidirectional permission verification of proxy control commands between different users.

[0038] Furthermore, the update process of the relationship topology network supports dynamic updates of relationship connections and feature profiles, meaning the network possesses continuous learning and adaptive evolution capabilities. Specifically, in daily vehicle use scenarios, based on multimodal perception data within the target vehicle (e.g., recording changes in clothing style and common accessories in different seasons or scenarios through visual perception images, and establishing a user's seating position preference model through seat perception data), combined with vehicle control command execution records and user feedback data on service objectives (e.g., voice evaluation of vehicle control execution results or manual fine-tuning), the feature profiles corresponding to each object node in the relationship topology network are continuously learned and automatically updated in the background. In addition, the relationship topology network also allows users to manually supplement, correct, or unbind established relationship types, name tags, and basic personal profile information through the vehicle control interface. Through this continuous learning and dynamic feature update mechanism, not only can changes in user appearance characteristics and vehicle control preferences over time be accurately captured, ensuring high accuracy of multimodal identity matching during long-term use, but the maintainability of the entire relationship topology network and the quality of proactive personalized services are also greatly improved.

[0039] It should be noted that, to ensure the timeliness of multimodal feature matching and the accuracy of identity recognition, and to avoid recognition biases caused by changes in user appearance, seating preferences, or relationship networks, the pre-constructed relationship topology network here is the latest version. In other words, this relationship topology network is a network model that has been iteratively updated synchronously after incorporating recent multimodal perception data, vehicle control execution records, and target feedback data through the aforementioned continuous learning and dynamic feature update mechanism. This ensures that accurate decisions can be made based on the user's most authentic and real-time feature profile during identity recognition.

[0040] After obtaining the relational topology network, if voice data is received from inside the target vehicle, the voiceprint features contained in the voice data are extracted, and the sound source location of the voice data is located using the audio sensor inside the target vehicle. Then, the first identity information of the speaker corresponding to the voice data is determined in the pre-constructed relational topology network.

[0041] The target vehicle here refers to a vehicle equipped with multimodal data acquisition capabilities (such as multiple sensors including audio sensors, visual sensors, and seat sensors) and intelligent cockpit zoning control capabilities. The voice data here refers to the full-domain voice signals or interactive commands collected by audio sensors such as microphone arrays within the target vehicle's cockpit at the current time. The voiceprint feature here is a data vector characterizing the unique biometric features of the speaker's voice. The sound source location here refers to the location information of the physical source emitting the voice data within the three-dimensional space of the target vehicle, such as specific seat identifiers like the driver's seat, front passenger seat, or rear left seat. The speaker here refers to the user currently issuing the voice command; the speaker's corresponding real identity node in the relational topology network is the primary identity information.

[0042] During the specific matching process in step 110, voiceprint features can be extracted from the collected speech data in real time to obtain the voiceprint features of the current speech data. The location of the sound source in the current speech data is then calculated using algorithms such as microphone array beamforming or sound source localization. After obtaining these two pieces of information, multiple candidate objects are determined in a pre-constructed relational topology network based on the sound source location. Specifically, objects whose seating preferences fall within the range of the sound source location are selected as candidate objects. The extracted voiceprint features are then compared and matched with the voiceprint features in the feature profiles of the candidate objects. This allows the speaker's primary identity information to be obtained from the relational topology network.

[0043] For example, when the microphone array locates the current sound source as the "driver's seat", it can prioritize retrieving the voiceprint features of candidates whose historical seating preferences include the "driver's seat" (such as "father" or "mother") from the relationship topology network for comparison. If the comparison finds that the voiceprint features of the current speech data match the voiceprint features of the "father" to a preset threshold, then the first identity information of the current speaker is determined to be the identity information corresponding to the "father". The first identity information includes, but is not limited to, the speaker's identity identifier and / or role in the relationship topology network.

[0044] Therefore, by combining the speaker's location information located by the microphone array with the voiceprint feature vector for multimodal identity matching, it is possible not only to effectively narrow the scope of voiceprint feature retrieval and comparison and significantly improve the recognition speed of the first identity information, but also to effectively filter out interference from in-vehicle environmental noise or other occupant conversations by utilizing spatial location information, thereby greatly improving the accuracy and robustness of speaker identity recognition. This enables the speaker's identity information to be determined quickly and accurately, providing an absolutely reliable identity benchmark for subsequent determination of agent control permissions and intent parsing.

[0045] Step 120: Perform semantic parsing on the voice data to obtain the control description information of the target vehicle and the object description information of the service object.

[0046] Optionally, after acquiring the voice data, speech recognition technology can be used to convert the voice data into text, and natural language understanding technology can be used to extract key information from the text to obtain control description information of the target vehicle and object description information of the service object.

[0047] The control description information here refers to the descriptive information related to vehicle control contained in the voice commands. This can include explicit control commands, such as clear control actions like "turn on," "turn off," or "adjust," and specific controlled vehicle components like "air conditioning," "seat," and "windows." It can also include more ambiguous control information, such as contextual or subjective intent words like "massage," "relax," or "it's too hot." The service recipient here refers to the user who requires the service indicated by the voice data. The object description information here refers to the characteristic descriptive words in the voice data used to refer to the service recipient, including but not limited to descriptions of appearance features and / or location, as well as relational descriptions.

[0048] Step 130: When there is ambiguity in the relational reference corresponding to the relational description information in the object description information, determine the second identity information and target location information of the service object based on the appearance feature description information and / or location description information in the object description information, as well as the first identity information and the relational description information.

[0049] The relationship description information here refers to the label words extracted from semantic parsing to represent interpersonal relationships, such as "son," "daughter," "child," or "colleague." Relationship ambiguity here refers to situations where the semantic relationship label alone cannot uniquely identify the service recipient within the target vehicle. For example, the instruction might be "turn on the seat heater for the child," but two children matching that relationship label are currently seated in the vehicle, thus triggering relationship ambiguity. The appearance feature description information here refers to the visual feature descriptions contained in the voice data, such as "a child wearing red clothes, glasses, or a hat." The location description information here is the description of spatial orientation within the vehicle contained in the voice data, such as the front passenger seat, the left rear seat, etc. The secondary identity information here refers to the service recipient's identity and role within the relationship topology. The target location information here refers to the specific physical seat position the service recipient is currently occupying within the target vehicle.

[0050] Optionally, after parsing and obtaining the relationship description information from the object description information of the service object in the voice data, object matching can be performed first in the relationship subnet corresponding to the first identity information in the relationship topology network based on the relationship description information. If a unique object is matched, it indicates that there is no ambiguity in the relationship reference information. The uniquely matched object can be directly taken as the service object, and its corresponding identity information in the relationship topology network can be determined as the second identity information. Subsequently, the biometric features of the service object in the target vehicle, such as at least one of vital signs, facial features, multimodal appearance features, and voiceprint features, can be combined to match the physical seat of the service object currently actually seated in the feature file corresponding to the second identity information, thereby obtaining the target location information.

[0051] Conversely, if multiple objects are matched, for example, if two objects labeled "son" exist simultaneously in the relationship subnet of the first identity information "father," it indicates that the current relationship description information has ambiguity. In this case, a multimodal disambiguation mechanism is triggered. This involves further extracting appearance feature descriptions and / or location descriptions from the object description information, combining them with the first identity information and relationship description information to perform multimodal feature recognition, thereby accurately determining the service object's second identity information and target location information. For example, the speaker's first identity information is used to define the subnet range of the relationship search, and the relationship description information is used to initially filter out candidate objects with specific relationships. Then, the appearance feature descriptions and / or location descriptions are matched and verified against the actual riding status inside the target vehicle, ultimately uniquely determining the service object's true second identity information and its target location information among multiple candidate objects.

[0052] It should be noted that, in specific embodiments, due to differences in users' voice command expression habits, the utilization of the above-described information has a high degree of flexibility: Method A1 can be to match and verify the appearance feature description information with the actual riding status inside the target vehicle (such as the visual perception results of the in-vehicle camera) to finally uniquely determine the real second identity information of the service object and its target location information among multiple candidate objects. Method A2 can be used to match and verify only the location description information with the actual occupancy status inside the target vehicle (such as the occupancy status of the seat sensor), and finally uniquely determine the true second identity information of the service object and its target location information among multiple candidate objects; Method A3 can be a comprehensive approach that combines appearance feature description information and location description information with the actual riding status inside the target vehicle to verify and ultimately uniquely determine the true second identity information of the service recipient and its target location information among multiple candidate objects.

[0053] In practical implementation, if the object description information contains only a single modality type of description information, that is, only appearance feature description information or only location description information, then one of the methods in method A1 and method A2 can be adaptively determined to determine the second identity information and target location information of the service object based on the modality type of the description features parsed in the object description information. If the object description information contains both appearance feature description information and location description information, then method A3 can be directly used to determine the second identity information and target location information of the service object, or one of the methods in method A1 and method A2 can be adaptively determined based on the priority level of each modality description information, etc. This embodiment does not specifically limit this.

[0054] Step 140: Generate vehicle control instructions for the service object based on the first identity information, the second identity information, the target location information, and the control description information, and control the target vehicle according to the vehicle control instructions.

[0055] The vehicle control commands here refer to structured command signals that can be directly recognized and responded to by the target vehicle's underlying domain controller (DC) or corresponding hardware actuators.

[0056] Optionally, after clarifying the first identity information, second identity information, target location information, and control description information, specific vehicle control commands can be generated by integrating this information. Specifically, the first identity information, second identity information, target location information, and control description information can be directly filled into the corresponding command template to generate the corresponding vehicle control command; or, based on the first identity information, second identity information, target location information, and control description information, further information reasoning can be performed to obtain accurate control parameters, control actions, control authority information, and the vehicle components to be controlled, and then this information can be combined to generate vehicle control commands, etc. This embodiment does not specifically limit this, and it can be adaptively determined according to the command interface specifications, communication protocol types, and underlying data parsing capabilities supported by the target vehicle's underlying domain controller or corresponding hardware execution mechanism. For example, if the underlying domain controller has strong semantic parsing computing power, it can directly issue templated intent commands; if the underlying hardware only supports basic bus signals, then all parameters need to be accurately reasoned and converted on the vehicle control device side before issuing the lowest-level device control message.

[0057] The method provided in this embodiment introduces appearance feature descriptions or location descriptions for multimodal disambiguation matching when semantic ambiguity is found in relational referencing. This breaks the dependence of existing technologies on a single voice or fixed address, achieving accurate identity and location confirmation in silent scenarios and under ambiguous addresses. Furthermore, by fusing the identity of the instruction initiator, the identity of the service object after accurate disambiguation, and its physical location, highly structured and precisely addressed vehicle control instructions can be generated. This ensures that vehicle control operations strictly and accurately act on the correct physical partitions and service objects, avoiding misoperation. As a result, the accuracy, robustness, and flexibility of personalized vehicle agent control are improved, providing users with a safer and more convenient personalized cockpit agent control service experience.

[0058] In some embodiments, determining the second identity information and target location information of the service object based on the appearance feature description information and / or location description information in the object description information, and the first identity information and the relationship description information, includes: The target location information is determined based on the appearance feature description information and / or the location description information; The second identity information is determined in the relationship topology network based on the target location information, the first identity information, and the relationship description information.

[0059] Optionally, when determining the secondary identity information and target location information of the service recipient, the appearance features or spatial location descriptions in the voice data can be visualized first, prioritizing the identification of the target's physical seat (i.e., target location information) at the physical space level. Subsequently, based on the passenger characteristics at that physical seat, combined with the speaker's primary identity information and relationship description information, a match is performed in the relationship topology network to determine their true identity. Thus, adopting a two-step strategy of first determining the physical location and then determining the network identity can effectively narrow the search range of candidates, filter out data interference from irrelevant passengers, and significantly improve the computational efficiency and accuracy of identity determination.

[0060] In some embodiments, determining the target location information based on the appearance feature description information includes: Extract multimodal appearance features of each passenger in the target vehicle from the visual perception image inside the target vehicle; The multimodal appearance features of each passenger are matched with the appearance feature description information to obtain the appearance matching degree corresponding to each passenger. The target location information is determined based on the seat position information of the passenger with the highest appearance matching degree in the target vehicle.

[0061] The visual perception image here refers to the cabin view captured in real time by visual sensors inside the target vehicle (such as a panoramic camera). The multimodal appearance features here refer to the external feature attributes of passengers extracted from the visual perception image through multimodal visual analysis, including but not limited to clothing features (such as colors like red and blue, and styles like T-shirts and jackets), accessory features (such as glasses, hats, and scarves), age features (such as children, youths, and the elderly), and other external features (such as hairstyles and body types).

[0062] Optionally, when using appearance feature description information to determine the target location information of the service object, an existing multimodal visual model with multimodal visual feature extraction can be used to perform deep analysis on the visual perception image inside the target vehicle to extract the appearance features of each passenger in each seat. These features are then semantically matched with the appearance feature descriptions parsed from the voice data (e.g., "a child wearing red clothes and glasses") to calculate the appearance matching degree for each passenger. Finally, the physical coordinates or identifier of the seat where the passenger with the highest appearance matching degree is located is selected to determine the target location information. It should be noted that this matching process supports fuzzy semantic matching; for example, "red" can be matched to similar color features such as dark red or pink to adapt to the generalization of natural interactions.

[0063] The method provided in this embodiment effectively solves the problem of locating the target passenger when the passenger is silent by directly corresponding multimodal visual features with natural language descriptions. This allows the initiator of the instruction to provide personalized services to a specific passenger based on appearance descriptions, thereby improving the flexibility of vehicle control.

[0064] In some embodiments, determining the target location information based on the location description information includes: If the number of target seats corresponding to the location description information is one, verify whether the target seat is occupied based on the seat perception data inside the target vehicle. When the target seat is occupied, the target location information is determined based on the physical location information corresponding to the target seat.

[0065] The seat perception data here refers to the physical sensing data collected by seat sensors (such as seat pressure sensors, gravity sensors, or seat belt status sensors) inside the target vehicle.

[0066] Optionally, when parsing the location description information reveals a unique target seat (e.g., "front passenger"), the seat perception data of that target seat can be directly retrieved to determine if the target seat is occupied. If the target seat is confirmed to be occupied, its physical location is directly determined as the target location information; if it is not occupied, the matching process can be terminated promptly, and the user can be notified via voice announcement or other means.

[0067] The method provided in this embodiment, when determining the target location information of the service object through location description information, uses seat perception data to verify the seat status, which can prevent invalid vehicle control operations from being performed on empty seats and improve the reliability of vehicle control.

[0068] In some embodiments, determining the target location information based on the appearance feature description information and the location description information includes: When there are multiple target seats corresponding to the location description information, at least one candidate seat is determined from the multiple target seats based on the seat perception data in the target vehicle and / or the speaker's gesture instructions; Extract multimodal appearance features of candidate passengers occupying each of the candidate seats from the visual perception image inside the target vehicle; The multimodal appearance features of each candidate passenger are matched with the appearance feature description information to obtain the appearance matching degree corresponding to each candidate passenger. The target location information is determined based on the seating position information of the candidate passenger with the highest appearance matching degree in the target vehicle.

[0069] Optionally, when parsing the location description information reveals that there are multiple target seats (e.g., the location description information is "rear row," corresponding to the left and right rear seats), empty seats can be filtered out first by combining seat perception data within the target vehicle, and / or by combining the speaker's gesture commands (e.g., pointing to a location in the rear row) recognized by the gesture sensor within the target vehicle for initial region screening, thereby identifying at least one candidate seat. Next, multimodal appearance features of candidate passengers on these candidate seats are extracted, and deep visual semantic matching calculations are performed by combining appearance feature description information from the speech. The seat position information of the candidate passenger with the highest matching degree within the target vehicle is then determined as the target position.

[0070] The method provided in this embodiment deeply integrates gesture commands, physical position sensing, and visual appearance features in multiple dimensions, which can cope with extremely complex cockpit referencing environments and significantly improve the robustness and accuracy of recognition when faced with ambiguous or broad voice commands.

[0071] In some embodiments, determining the second identity information in the relationship topology network based on the target location information, the first identity information, and the relationship description information includes: Based on the relationship description information and the first identity information, a target relationship subnet is determined in the relationship topology; Based on the target location information, the biometric features of the service recipient are obtained; the biometric features include at least one of vital signs, facial features, multimodal appearance features, and voiceprint features. Match the biometric characteristics of the service recipient with the feature profiles of each object in the target relationship subnet; Based on the matching results, obtain the second identity information.

[0072] The target relational subnet here refers to a set of specific associated nodes extending outward from the primary identity information as the central node, based on the relational description information. For example, if the primary identity information is the main driver and the relational description information includes the relational label "son," then the set of all nodes of the main driver labeled as "son" can be considered the target relational subnet. Biometric features here refer to vital sign data that can uniquely identify or assist in confirming a passenger's identity, including vital sign features obtained by the seat sensor (such as weight range, height range, etc.), facial and appearance features obtained by the visual sensor, and voiceprint features obtained by the audio sensor.

[0073] Optionally, after obtaining the target location information of the service recipient, the first identity information can be used as an anchor point to extract a list of possible related persons, forming a target relationship subnet, based on the relationship description. Then, using the determined target location information, real-time biometric features collected at that physical location are retrieved and compared with the pre-stored biometric information in the feature files of each candidate object in the target relationship subnet. The comprehensive matching confidence score corresponding to the maximum matching degree is obtained. If the comprehensive matching confidence score is higher than a preset safety threshold, the identity information of the candidate object corresponding to the maximum matching degree in the target relationship subnet is determined as the second identity information. Otherwise, the speaker can be proactively asked for confirmation through the target vehicle's audio sensor (voice broadcast) or visual interaction interface. For example, a confirmation script can be automatically generated and broadcast: "Do you want to turn on the seat heater for the child wearing red clothes sitting on the left side of the back seat?" Upon receiving a positive confirmation instruction from the speaker, such as a voice reply "yes" or a nod, the corresponding identity recognition and location determination of the service recipient, as well as the generation of subsequent vehicle control commands, are then executed. Here, by introducing an active inquiry and confirmation mechanism under low confidence, not only is the vehicle control device given fault tolerance in complex and extreme cockpit environments, but it also effectively avoids misoperation caused by identification deviation, further strengthening the safety defense of vehicle agent control.

[0074] Furthermore, if an interactive process of active inquiry and confirmation occurs, the interactive information generated by this process can be fed back to the system's continuous learning module. This information can then be used as feedback data to update the feature profiles of the corresponding service objects in the relational topology network, thereby significantly improving the robustness and accuracy of recognition during long-term use.

[0075] The method provided in this application effectively eliminates data interference from strangers or irrelevant occupants in the vehicle by combining local subnet search based on graph network relationships with real-time biometric verification of the target location, thereby further ensuring the accuracy of relationship identification and the privacy and security of in-vehicle agent control.

[0076] In some embodiments, generating vehicle control commands for the service object based on the first identity information, the second identity information, the target location information, and the control description information includes: Based on the control description information, the control intent is obtained; Based on the control intent and the target location information, obtain the vehicle-mounted components to be controlled inside the target vehicle; Based on the first identity information and the second identity information, determine whether the speaker has proxy control authority over the vehicle-mounted component to be controlled; When the speaker has proxy control authority over the vehicle-mounted component to be controlled, the vehicle control command is generated based on the control intent, the second identity information, and the vehicle-mounted component to be controlled.

[0077] The controlled in-vehicle components here refer to independent hardware modules within the target vehicle responsible for performing specific physical actions, such as independent temperature zone air conditioning, independent ventilated / heated seats, corresponding side window motors, or entertainment screens. The delegated control permissions here refer to granting specific users the authority to control components in other areas based on preset security rules or licensing mechanisms. Typically, the driver has the highest control authority in the entire vehicle, while other occupants can only control components within their own seats or other seat components mutually authorized through a network of relationships.

[0078] Optionally, when generating vehicle control commands, the control description information can be parsed first to determine the control intent. Based on the parsed control intent and target location information, the vehicle component to be controlled can be accurately located. Before issuing the command, to ensure the effectiveness of vehicle control, an authorization verification mechanism can be triggered to check the hierarchical relationship between the first identity information (speaker) and the second identity information (service object), determining whether they have the authority to perform proxy operations on the vehicle component to be controlled for the service object. If the verification passes, the vehicle control command is generated based on the control intent, the second identity information, and the vehicle component to be controlled; if the verification fails, execution can be refused and a voice prompt can be given.

[0079] The steps of generating vehicle control commands through control intent, secondary identity information, and the vehicle-mounted component to be controlled can be tailored to various command generation mechanisms, depending on the clarity of the control intent and the specific attributes of the secondary identity information. Specific embodiments include, but are not limited to, the following methods: Method B1 is a command generation mechanism based on implicit control intent and personalized preference completion. Specifically, when the control intent is an implicit intent that only includes the operation action without specific target parameters, such as the intent being only "turn on the air conditioner" or "play music", and the vehicle component to be controlled requires specific operating parameters, the personalized feature profile corresponding to the service object is retrieved from the relational topology network using the second identity information, the corresponding historical preference parameters are extracted, and then the preference parameters are used to complete the control parameters to generate a complete vehicle control command.

[0080] Method B2 is based on a command generation mechanism that weights relative control intent and identity sensitivity. Specifically, when the control intent is a relative adjustment intent (e.g., the intent is "turn it up a bit," "it's too cold," or "turn it down a bit"), the sensitivity weight of the service recipient to environmental changes can be determined based on the age group, gender, or physical characteristics of the second identity information. Combined with the current operating status of the vehicle components to be controlled (e.g., current real-time temperature or current brightness), control parameters are dynamically calculated to generate vehicle control commands.

[0081] Method B3 is a command generation mechanism based on explicit control intent and identity security constraints. Specifically, when the control intent is an explicit intent containing clear parameters, such as "lower all the windows," but the action of the controlled vehicle component may involve occupant safety, the corresponding security constraint rule base can be triggered based on the second identity information. The security constraint rules are used to verify the legality of the control intent and correct the control action, thereby generating a vehicle control command that conforms to the security level of the service object.

[0082] Method B4 is an instruction generation mechanism based on contextualized fuzzy intent and multi-component collaborative reasoning. Specifically, when the control intent is a contextualized or subjective fuzzy intent, such as "want to rest" or "feeling carsick," the control intent, the feature profile of the second identity information, and all available vehicle components to be controlled within the target location area can be input into an existing vehicle control reasoning model or an existing large language model pre-deployed in the target vehicle. The model then infers and outputs a combined sequence of vehicle control instructions for the service object, thereby achieving multi-modal collaborative control across components.

[0083] The method provided in this embodiment effectively prevents malicious interference with control or unauthorized operations between occupants by introducing a proxy permission verification mechanism based on relationship networks and identity features, which greatly improves the security and logical rationality of vehicle function control in multi-occupant scenarios.

[0084] In some embodiments, generating the vehicle control command based on the control intent, the second identity information, and the vehicle-mounted component to be controlled includes: Based on the second identity information, determine the feature profile of the service object in the relationship topology network; Based on the control parameters and control actions of the vehicle-mounted component to be controlled, the control intent, and the historical preference information in the feature profile of the service object, the control parameters and control actions of the vehicle-mounted component to be controlled are obtained. The vehicle control command is generated based on the control parameters, the control action, and the vehicle-mounted component to be controlled.

[0085] The control parameters here include temperature values, angle values, etc., but this embodiment does not specifically limit them.

[0086] Optionally, after verifying that the speaker has the authority to control the vehicle component under control, the feature profile specific to the service recipient can be retrieved directly from the latest updated relational topology based on the determined second identity information. For fuzzy or generalized control intentions (e.g., the intention is simply "turn on the air conditioning"), historical preference information from the feature profile is automatically extracted (e.g., the "son" prefers a temperature of 26 degrees Celsius and a fan speed of level 2), thereby deriving precise control parameters (26 degrees Celsius, level 2) and control actions (e.g., turning on). Finally, the control parameters, control actions, and the vehicle component under control are integrated to generate a complete structured vehicle control command.

[0087] The method provided in this embodiment, based on the deep reuse of the historical identity data of the served user, automatically maps generalized control intentions to specific parameters, avoiding the tedious specification of various adjustment details by the instruction initiator, effectively improving the proactive service of personalized vehicles, and greatly enhancing the accuracy, robustness, flexibility and interactive experience of personalized vehicle agent control.

[0088] In some embodiments, the vehicle control method further includes the following feedback and update steps: receiving target feedback data from the service object; the target feedback data is service evaluation data input by the service object after the target vehicle executes the vehicle control command; updating the feature profiles of each object in the relational topology network based on the target feedback data, the control record corresponding to the vehicle control command, and the multimodal perception data in the target vehicle; the multimodal perception data includes at least one of voice data, visual perception images, and seat perception data.

[0089] The target feedback data here refers to the evaluation feedback results given by the service recipient after the vehicle control command is executed, through voice evaluation (such as "the temperature is just right" or "it's still a bit cold"), manual fine-tuning operation, or visual expression.

[0090] Optionally, after executing the vehicle control command, the system continuously monitors the user's subsequent reactions, i.e., target feedback data, through multi-source sensors. It then integrates this target feedback data, the control records of the current execution, and the multimodal perception data collected, sending this incremental data into the learning module to automatically correct or enrich the feature profiles of corresponding objects in the relationship topology. For example, if the user reports that the temperature is too cold and manually adjusts it, their temperature preference data is updated; if the camera detects that they changed a specific accessory today, the new accessory is added to the appearance feature database.

[0091] The method provided in this embodiment, through a multimodal driven continuous learning and feature update mechanism, can dynamically capture the daily changes in a user's appearance features and the long-term drift of personalized preferences, thereby ensuring high matching accuracy and strong adaptive iteration capability of the entire vehicle control device during its long life cycle.

[0092] In summary, the method provided in this application has the following significant technical advantages and beneficial effects compared to the prior art: First, it greatly enriches the methods of identity recognition and successfully solves the challenge of identity recognition in silent scenarios. The method provided in this application supports recognition based on a combination of multiple dimensions, including nicknames, relationship tags, appearance feature descriptions (such as clothing color, style, pattern, and accessories such as glasses, hats, and scarves), and location. Users do not need to remember specific titles or command formats; they can complete the interaction simply by using natural language descriptions, making the expression more natural and flexible. In particular, by introducing visual feature recognition technology, identity verification can be accurately completed even if the service recipient does not actively speak, completely breaking the limitations of existing technologies that heavily rely on voiceprint features or require verification by the target passenger's voice. This is especially suitable for providing a convenient proxy vehicle control mechanism for the primary driver, facilitating services for users such as the elderly and children who are unable to speak or are unfamiliar with vehicle functions.

[0093] Secondly, it significantly improves the accuracy of service object identification, exhibiting strong adaptability and robustness. The method provided in this application deeply integrates relationship networks and appearance features, comprehensively analyzing and calculating decision confidence based on multiple modal information such as voiceprint (voiceprint, semantics), vision (face, appearance), location (seat sensor), and gestures. Compared to existing technologies that rely solely on single basic features such as voiceprint or weight, the method provided in this application achieves multi-mode collaborative work and mutual verification. Furthermore, the method provided in this application also supports multi-feature combination matching and fuzzy matching, effectively disambiguating and maintaining extremely high recognition accuracy even in complex interference scenarios such as noisy in-vehicle environments, changes in lighting, and temporary changes in the appearance of the target object, while significantly reducing the false recognition rate.

[0094] Secondly, the vehicle control execution efficiency has been optimized, enabling precise and personalized zoned services. The method provided in this application can not only perform complex identity recognition but also directly output structured vehicle control commands. These commands fully encompass complete information such as the controlled object, control actions, target parameters, and the location of the service object, effectively simplifying the vehicle control logic processing flow and improving the overall response speed. After accurately locating the seat position of the service object, it can precisely execute vehicle function controls for that specific physical location, such as independently adjusting zoned temperature, seat angle, and music entertainment. This not only meets the diverse needs of different users but also completely avoids accidental operation of devices in other parts of the vehicle.

[0095] Secondly, a continuous learning mechanism for feature profiles has been established, along with robust privacy protection. Over long-term use, the system continuously records and learns users' dynamic characteristics and preferences, including evolution of appearance features (such as seasonal clothing styles and preferred accessories), location preferences, and personalized vehicle control settings. This establishes and dynamically updates user feature profiles, enabling continuous iteration and optimization of recognition accuracy and service quality over time, significantly improving vehicle control precision. Simultaneously, users' multimodal feature data can be stored locally or in a cloud environment authorized by the user, with encrypted storage and transmission throughout the entire process. Furthermore, the establishment of relationship networks and profile associations strictly adhere to a two-way confirmation mechanism, fully protecting the personal privacy and data security of multiple users within the vehicle from the system's underlying layer.

[0096] The vehicle control device provided by the present invention is described below. The vehicle control device described below can be referred to in correspondence with the vehicle control method described above.

[0097] Figure 2 This is a structural schematic diagram of the vehicle control device provided by the present invention; as shown. Figure 2 As shown, the device includes: The first identification unit 210 is used to determine the first identity information of the speaker corresponding to the voice data in a pre-constructed relational topology network based on the voiceprint features of the voice data in the target vehicle and the sound source location of the voice data. The parsing unit 220 is used to perform semantic parsing on the voice data to obtain the control description information of the target vehicle and the object description information of the service object; The second identification unit 230 is used to determine the second identity information and target location information of the service object based on the appearance feature description information and / or location description information in the object description information, as well as the first identity information and the relationship description information, when there is ambiguity in the relationship reference corresponding to the relationship description information in the object description information. The control unit 240 is used to generate vehicle control instructions for the service object based on the first identity information, the second identity information, the target location information, and the control description information, and to control the target vehicle according to the vehicle control instructions.

[0098] The device provided in this embodiment introduces appearance feature descriptions or location descriptions for multimodal disambiguation matching when semantic ambiguity is found in relational referencing. This breaks the dependence of existing technologies on a single voice or fixed title, achieving accurate identity and location confirmation in silent scenarios and under ambiguous titles. Furthermore, by fusing the identity of the instruction initiator, the identity of the service object after accurate disambiguation, and its physical location, it can generate highly structured and precisely addressed vehicle control instructions. This ensures that vehicle control operations strictly and accurately act on the correct physical partitions and service objects, avoiding misoperation. As a result, it improves the accuracy, robustness, and flexibility of personalized vehicle agent control, providing users with a safer and more convenient personalized cockpit agent control service experience.

[0099] The apparatus provided by the present invention is used to execute the above-described method embodiments. For specific processes and details, please refer to the above embodiments, which will not be repeated here.

[0100] Figure 3 This is a structural schematic diagram of the vehicle provided by the present invention; as shown below. Figure 3 As shown, the vehicle includes an onboard controller 310 and a multi-source data acquisition unit; the multi-source data acquisition unit includes an audio sensor 320 for audio data acquisition, a visual sensor 330 for visual image acquisition, a seat sensor 340 for seat perception data acquisition, and a gesture recognition sensor 350 for gesture command acquisition; the onboard controller is used to execute a vehicle control method, which includes: determining the first identity information of the speaker corresponding to the voice data in a pre-constructed relational topology network based on the voiceprint characteristics of the voice data in the target vehicle and the sound source location of the voice data; processing the voice data... The data undergoes semantic parsing to obtain control description information of the target vehicle and object description information of the service object. When there is ambiguity in the relational reference corresponding to the relational description information in the object description information, the second identity information and target location information of the service object are determined based on the appearance feature description information and / or location description information in the object description information, as well as the first identity information and the relational description information. Based on the first identity information, the second identity information, the target location information, and the control description information, a vehicle control command is generated for the service object, and the target vehicle is controlled according to the vehicle control command.

[0101] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communications bus 440. The processor 410 can call logical instructions in the memory 430 to execute a vehicle control method, which includes: determining the first identity information of the speaker corresponding to the voice data in a pre-constructed relational topology network based on the voiceprint features of the voice data in the target vehicle and the sound source location of the voice data; performing semantic parsing on the voice data to obtain control description information of the target vehicle and object description information of the service object; when there is ambiguity in the relational reference information in the object description information, determining the second identity information and target location information of the service object based on the appearance feature description information and / or location description information in the object description information, as well as the first identity information and the relational description information; generating vehicle control instructions for the service object based on the first identity information, the second identity information, the target location information, and the control description information, and controlling the target vehicle according to the vehicle control instructions.

[0102] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0103] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the vehicle control method provided by the above methods. The method includes: determining the first identity information of the speaker corresponding to the voice data in a pre-constructed relational topology network based on the voiceprint features of the voice data in the target vehicle and the sound source location of the voice data; performing semantic parsing on the voice data to obtain control description information of the target vehicle and object description information of the service object; when there is ambiguity in the relational reference corresponding to the relational description information in the object description information, determining the second identity information and target location information of the service object based on the appearance feature description information and / or location description information in the object description information, as well as the first identity information and the relational description information; generating vehicle control instructions for the service object based on the first identity information, the second identity information, the target location information, and the control description information, and controlling the target vehicle according to the vehicle control instructions.

[0104] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the vehicle control method provided by the above methods. The method includes: determining, based on the voiceprint features of voice data within a target vehicle and the sound source location of the voice data, a first identity information of the speaker corresponding to the voice data in a pre-constructed relational topology network; performing semantic parsing on the voice data to obtain control description information of the target vehicle and object description information of the service object; when there is ambiguity in the relational reference corresponding to the relational description information in the object description information, determining, based on the appearance feature description information and / or location description information in the object description information, as well as the first identity information and the relational description information, a second identity information and target location information of the service object; generating vehicle control instructions for the service object based on the first identity information, the second identity information, the target location information, and the control description information, and controlling the target vehicle according to the vehicle control instructions.

[0105] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0106] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A vehicle control method, characterized in that, include: Based on the voiceprint characteristics of the voice data inside the target vehicle and the location of the sound source of the voice data, the first identity information of the speaker corresponding to the voice data is determined in a pre-constructed relational topology network. Semantic parsing is performed on the voice data to obtain the control description information of the target vehicle and the object description information of the service object; When there is ambiguity in the relational reference corresponding to the relational description information in the object description information, the second identity information and target location information of the service object are determined based on the appearance feature description information and / or location description information in the object description information, as well as the first identity information and the relational description information; Based on the first identity information, the second identity information, the target location information, and the control description information, a vehicle control command is generated for the service object, and the target vehicle is controlled according to the vehicle control command.

2. The vehicle control method according to claim 1, characterized in that, The step of determining the second identity information and target location information of the service object based on the appearance feature description information and / or location description information in the object description information, as well as the first identity information and the relationship description information, includes: The target location information is determined based on the appearance feature description information and / or the location description information; The second identity information is determined in the relationship topology network based on the target location information, the first identity information, and the relationship description information.

3. The vehicle control method according to claim 2, characterized in that, Determining the second identity information in the relationship topology network based on the target location information, the first identity information, and the relationship description information includes: Based on the relationship description information and the first identity information, a target relationship subnet is determined in the relationship topology; Based on the target location information, the biometric features of the service recipient are obtained; the biometric features include at least one of vital signs, facial features, multimodal appearance features, and voiceprint features. Match the biometric characteristics of the service recipient with the feature profiles of each object in the target relationship subnet; Based on the matching results, obtain the second identity information.

4. The vehicle control method according to claim 2, characterized in that, Based on the appearance feature description information, the target location information is determined, including: Extract multimodal appearance features of each passenger in the target vehicle from the visual perception image inside the target vehicle; The multimodal appearance features of each passenger are matched with the appearance feature description information to obtain the appearance matching degree corresponding to each passenger. The target location information is determined based on the seat position information of the passenger with the highest appearance matching degree in the target vehicle.

5. The vehicle control method according to claim 2, characterized in that, Determining the target location information based on the appearance feature description information and the location description information includes: When there are multiple target seats corresponding to the location description information, at least one candidate seat is determined from the multiple target seats based on the seat perception data in the target vehicle and / or the speaker's gesture instructions; Extract multimodal appearance features of candidate passengers occupying each of the candidate seats from the visual perception image inside the target vehicle; The multimodal appearance features of each candidate passenger are matched with the appearance feature description information to obtain the appearance matching degree corresponding to each candidate passenger. The target location information is determined based on the seating position information of the candidate passenger with the highest appearance matching degree in the target vehicle.

6. The vehicle control method according to claim 2, characterized in that, Determining the target location information based on the location description information includes: If the number of target seats corresponding to the location description information is one, verify whether the target seat is occupied based on the seat perception data inside the target vehicle. When the target seat is occupied, the target location information is determined based on the physical location information corresponding to the target seat.

7. The vehicle control method according to any one of claims 1-6, characterized in that, The step of generating vehicle control commands for the service object based on the first identity information, the second identity information, the target location information, and the control description information includes: Based on the control description information, the control intent is obtained; Based on the control intent and the target location information, obtain the vehicle-mounted components to be controlled inside the target vehicle; Based on the first identity information and the second identity information, determine whether the speaker has proxy control authority over the vehicle-mounted component to be controlled; When the speaker has proxy control authority over the vehicle-mounted component to be controlled, the vehicle control command is generated based on the control intent, the second identity information, and the vehicle-mounted component to be controlled.

8. The vehicle control method according to claim 7, characterized in that, The step of generating the vehicle control command based on the control intent, the second identity information, and the vehicle-mounted component to be controlled includes: Based on the second identity information, determine the feature profile of the service object in the relationship topology network; Based on the control parameters and control actions of the vehicle-mounted component to be controlled, the control intent, and the historical preference information in the feature profile of the service object, the control parameters and control actions of the vehicle-mounted component to be controlled are obtained. The vehicle control command is generated based on the control parameters, the control action, and the vehicle-mounted component to be controlled.

9. The vehicle control method according to any one of claims 1-6, characterized in that, The method further includes: Receive target feedback data from the service recipient; the target feedback data is service evaluation data input by the service recipient after the target vehicle executes the vehicle control command. Based on the target feedback data, the control records corresponding to the vehicle control commands, and the multimodal perception data within the target vehicle, the feature profiles of each object in the relational topology network are updated; the multimodal perception data includes at least one of voice data, visual perception images, and seat perception data.

10. A vehicle control device, characterized in that, include: The first identification unit is used to determine the first identity information of the speaker corresponding to the voice data in a pre-constructed relational topology network based on the voiceprint features of the voice data in the target vehicle and the sound source location of the voice data. The parsing unit is used to perform semantic parsing on the voice data to obtain the control description information of the target vehicle and the object description information of the service object; The second identification unit is used to determine the second identity information and target location information of the service object based on the appearance feature description information and / or location description information in the object description information, as well as the first identity information and the relationship description information, when there is ambiguity in the relationship reference corresponding to the relationship description information in the object description information. The control unit is configured to generate vehicle control commands for the service object based on the first identity information, the second identity information, the target location information, and the control description information, and to control the target vehicle according to the vehicle control commands.

11. A vehicle, characterized in that, include: The vehicle controller and multi-source data acquisition unit; the multi-source data acquisition unit includes an audio sensor, a visual sensor, a seat sensor, and a gesture recognition sensor; The on-board controller is used to perform the vehicle control method as described in any one of claims 1 to 9.

12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the vehicle control method as described in any one of claims 1 to 9.