Voice interaction method and system, and vehicle

By binding occupant voiceprint features and identities to the in-vehicle voice system, a family relationship model is generated, which solves the problem of ambiguous command sources in multi-occupant scenarios of the in-vehicle voice system, realizes accurate identity recognition and personalized services, and improves the convenience and security of voice interaction.

CN121938360APending Publication Date: 2026-04-28GREAT WALL MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GREAT WALL MOTOR CO LTD
Filing Date
2026-01-15
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing in-vehicle voice systems struggle to accurately pinpoint the source of commands in multi-occupant scenarios, lacking an understanding of the social relationships and interaction context among occupants. This leads to mis-triggering and chaotic responses, preventing the provision of intelligent and personalized services.

Method used

When the vehicle starts, the voiceprint features and identities of the occupants are bound together to generate a family relationship model. The identity of the voice command issuer is determined through the voiceprint features and family relationship model, realizing decentralized voice interaction and supporting accurate identity recognition and personalized services in multi-occupant scenarios.

Benefits of technology

It achieves accurate identity verification without manual operation, reduces the risk of false responses, improves the convenience of voice interaction and driving safety, adapts to natural dialogue in family scenarios, and provides intelligent and personalized services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121938360A_ABST
    Figure CN121938360A_ABST
Patent Text Reader

Abstract

The invention relates to a voice interaction method and system and a vehicle, and belongs to the technical field of vehicle voice interaction, and the method comprises the steps: binding the voiceprint feature of each passenger in the vehicle with the identity of the passenger when the current driving is started, and generating a family relation model according to the identity of each passenger in the vehicle; when the voice instruction is received, the identity of a voice instruction sender is confirmed according to the voiceprint feature of the voice instruction; recognizing an associated subject in the voice instruction, and determining a target passenger pointed by the associated subject according to a family relationship model based on the associated subject and the identity of a voice instruction sender; and controlling to execute an operation corresponding to the voice instruction and the target passenger pointed by the associated main body. According to the invention, intelligent and personalized voice interaction service is provided, and a decentralized interaction effect is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle voice interaction technology, and in particular to a voice interaction method and system, and a vehicle. Background Technology

[0002] With the development of intelligent connected vehicles, in-vehicle voice systems have become standard equipment. Existing in-vehicle voice systems typically employ a centralized interaction model, meaning they usually respond to the driver's commands by default, or require a wake word or button press to specify the interaction target. When multiple passengers are in the vehicle, on the one hand, the in-vehicle voice system struggles to distinguish the source of the command from the system's intended meaning, easily leading to mis-triggering or confusing responses; on the other hand, the semantic understanding capabilities of in-vehicle voice systems are limited to the recognition and execution of isolated commands, lacking a deep understanding of context and the social relationships between passengers. Therefore, existing in-vehicle voice systems cannot provide intelligent and personalized services based on the sender of the command. Summary of the Invention

[0003] To address the aforementioned problems in existing technologies, this application provides a voice interaction method and system, storage medium, and vehicle. Upon vehicle startup, a family relationship model is generated based on the identities of each passenger, and the source of the command is accurately located based on voiceprint features. The target object referred to by the associated subject in the voice command is determined by combining the family relationship model. This eliminates reliance on a single main user for voice interaction; any passenger with a confirmed identity can issue voice commands. Intelligent and personalized services can be provided based on the voice command issuer, achieving a decentralized interaction effect of "whoever speaks, serves them."

[0004] Firstly, a voice interaction method is provided for use in vehicles. The method includes: When the vehicle starts, the voiceprint characteristics of each passenger in the vehicle are bound to their identity, and a family relationship model is generated based on the identities of each passenger in the vehicle. Upon receiving a voice command, the identity of the voice command sender is confirmed based on the voiceprint characteristics of the voice command. Identify the associated subject in the voice command, and determine the target occupant referred to by the associated subject based on the identity of the associated subject and the voice command issuer, according to the family relationship model; Controls the execution of operations corresponding to voice commands and the target occupant indicated by the associated entity.

[0005] It should be noted that this driving session refers to the complete driving process from starting the engine (or turning on the power) to turning it off (or turning off the power). Specifically, the time frame of this driving session is from the start of the vehicle's power system to the end of the time frame, such as a one-way drive from home to the office, or a one-way drive from home to the mall or from the mall back home during a shopping trip.

[0006] In this embodiment, upon vehicle startup, the voiceprint characteristics of each occupant are bound to their identity. A family relationship model is generated based on the identities of each occupant, and the source of the voice command is accurately located based on the voiceprint characteristics. The target object referred to by the associated subject (e.g., husband, wife, son) in the voice command is determined by combining the family relationship model. On one hand, this eliminates reliance on a single main user for voice interaction; any occupant with a confirmed identity can issue voice commands. Intelligent and personalized services can be provided based on the voice command issuer, achieving a decentralized interaction effect of "whoever speaks, serves them." On the other hand, it reduces the burden of manual operation and explicit target designation for the driver, making voice interaction more natural and suitable for family scenarios, thus improving convenience and driving safety.

[0007] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: If a new passenger boards the vehicle during this trip, or if a new passenger boards the vehicle at the next start of the trip, the identity of the new passenger shall be identified. The relationship between the new passenger and other passengers is incorporated into the family relationship model, and the family relationship model is updated.

[0008] In this embodiment, when a new passenger boards the vehicle during the current trip or upon the next trip, the relationship between the new passenger and existing passengers is automatically added to the family relationship model based on the passenger's identification result, eliminating the need for manual editing by the user. This addresses the technical pain point of outdated family relationship model information in scenarios such as family member additions / removals or temporary passenger arrivals, breaking the limitations of traditional family relationship models that require "static input and manual maintenance." It ensures that the family relationship model always covers all driving and passenger-related personnel and their relationships, adapting to the high-frequency changes in passenger structure during family trips (such as weekend family outings, weekday single-person commutes, and temporary passenger pick-ups), ensuring accurate and efficient passenger identification for every trip, and providing accurate and comprehensive relationship data support for subsequent personalized services.

[0009] In conjunction with the first aspect, in some implementations of the first aspect, the methods for identifying the identities of each crew member include: At the start of this driving operation, facial images of each occupant inside the vehicle are acquired; Match each occupant's facial image with pre-stored facial images; When a passenger's facial image is matched with a pre-stored facial image, the passenger's identity identifier associated with the pre-stored facial image is obtained; Passenger identities are determined by using a pre-stored database of family member relationships of vehicle owners.

[0010] In this embodiment, facial image matching automatically collects facial information and compares it with pre-stored facial images after an occupant enters the cabin and takes their seat, achieving seamless and automated identity verification. Upon successful matching, the identity information is directly associated with the occupant's identification identifier, improving the stability and accuracy of identity recognition. Furthermore, facial image matching can cover the identity recognition needs of all occupants in the vehicle, resolving the issue of identity confusion in multi-occupant scenarios.

[0011] In conjunction with the first aspect, in some implementations of the first aspect, the methods for confirming the identity of each crew member also include: If a passenger's facial image does not match a pre-stored facial image, that passenger will be added as a new passenger. The information for newly added passengers includes passenger identification, passenger name, and the relationship between the passenger and the vehicle owner and other family members of the vehicle owner. The identity of the new passenger is obtained based on the information set for the new passenger.

[0012] In this embodiment, when an unmatched occupant face image is detected, the occupant is treated as a new occupant, and new occupant information (identity identifier, name, and kinship information) is set. The identity of the newly added occupant is obtained based on the set new occupant information. For new occupants not in the pre-stored face database, a complete occupant identity profile is constructed by supplementing information such as identity identifier, name, and kinship. The constructed occupant identity profile can effectively identify the new user. This solves the problem that traditional face recognition can only match existing users and cannot identify new users. It can realize the identification of all drivers and passengers. In subsequent voice interaction, when a new occupant acts as the issuer of a command, personalized services can be provided based on the voice commands issued by the new occupant.

[0013] In conjunction with the first aspect, in some implementations of the first aspect, the method for identifying the identities of each occupant also includes: When a passenger's facial image does not match a pre-stored facial image, the new passenger's information is stored in the vehicle owner's family member relationship database, and the new passenger's facial image is also stored. The newly added passenger facial images are associated with the newly added passenger identification.

[0014] In this embodiment, when an unmatched occupant facial image is detected, on the one hand, the newly added occupant information is stored in the vehicle owner's family member relationship database, updating the database. This approach overcomes the limitations of traditional identity recognition's "pre-entered fixed list," dynamically adapting to changes in family occupant structure (e.g., adding family members, temporary passengers, etc.), expanding the identity database without relying on professional personnel or complex operations, thus improving practicality and adaptability. On the other hand, after associating and storing the newly added occupant's facial image with their identity identifier, in any subsequent ride scenario, the corresponding occupant's identity information can be retrieved directly through facial image matching, achieving seamless identity verification. There is no need to repeatedly enter occupant names, kinship information, etc., significantly reducing recognition time and improving the user experience for drivers and passengers.

[0015] In conjunction with the first aspect, in some implementations of the first aspect, the method for identifying the identities of each occupant also includes: When a passenger's facial image does not match a pre-stored facial image, retrieve the voice of the newly added passenger; Extract the voiceprint features from the voice of the newly added passenger and bind the voiceprint features to the identity of the newly added passenger.

[0016] In this embodiment, for new passengers outside the pre-stored face database, the voice of the new passenger is obtained, and the voiceprint features of the new passenger are bound to the passenger identity. When the new passenger acts as the issuer of voice commands, the identity of the new passenger can be quickly located based on the voiceprint features to adapt to the voice command response of the new passenger, thereby realizing intelligent and personalized voice interaction services for the new passenger.

[0017] In conjunction with the first aspect, in some implementations of the first aspect, the method for generating a family relationship model based on the identities of each passenger in the vehicle is as follows: extract the relationships between the identities of each passenger from a pre-stored database of family member relationships of the vehicle owner, and generate a family relationship model based on the extracted relationships between the identities of each passenger.

[0018] In this embodiment, the family relationship model is generated by directly extracting the kinship relationships between passengers from the pre-stored vehicle owner family member relationship database, which can ensure the accuracy and authority of the identity association information in the graph.

[0019] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: if, within a set time period, a member in the family relationship model takes fewer rides than a set number of times, then that member and the relationship between them and other members are removed from the family relationship model.

[0020] In this embodiment, by setting time and ride frequency thresholds, graph members and their relationship links that have not been used for a long time are automatically filtered and deleted. On the one hand, this avoids the graph from becoming bloated due to the accumulation of redundant member information, reducing the computing load and storage resource consumption of the in-vehicle terminal, and ensuring the response speed of voice interaction commands. On the other hand, it ensures that the graph always maintains a strong correlation with the current and recent passengers, ensuring the matching accuracy between the associated subject and the target passenger during voice command parsing, and avoiding response errors caused by invalid data interference.

[0021] Secondly, a voice interaction system is provided for use in vehicles. This system includes: The binding module is configured to bind each occupant's voiceprint characteristics to their occupant identity. The model building module is configured to generate a family relationship model based on the identities of each passenger in the vehicle when the vehicle is started. The interaction module is configured to: upon receiving a voice command, confirm the identity of the voice command issuer based on the voiceprint characteristics of the voice command; identify the associated subject in the voice command; and, based on the identities of the associated subject and the voice command issuer, determine the target occupant referred to by the associated subject according to a family relationship model. The control module is configured to control the execution of operations corresponding to voice commands and the target occupant indicated by the associated subject.

[0022] Thirdly, a computer program product is provided, the computer program product comprising: computer program code, which, when run on a computer, causes the computer to execute the voice interaction method of the first aspect described above.

[0023] Fourthly, a computer-readable storage medium is provided that stores computer program code, which is executed by one or more processors, such that when the computer program code is run on a processor, a device including the one or more processors performs the voice interaction method of the first aspect described above.

[0024] Fifthly, embodiments of this application provide a chip system, the chip system including a processor for calling a computer program or computer instructions stored in a memory, so that the processor executes the voice interaction method of the first aspect described above.

[0025] In a sixth aspect, an electronic device according to an embodiment of this application includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the voice interaction method of the first aspect.

[0026] In a seventh aspect, a vehicle is provided. The vehicle includes the voice interaction system described in the second aspect above, or the computer-readable storage medium described in the fourth aspect above, or the chip system described in the fifth aspect above, or the electronic device described in the sixth aspect above.

[0027] The beneficial effects of the technical solutions provided in this application include at least the following: The voice interaction method, system, and vehicle provided in this application are applied to vehicles. Upon vehicle startup, the voiceprint characteristics of each occupant are bound to their identity. A family relationship model is generated based on the identities of each occupant. The source of the voice command is accurately located based on the voiceprint characteristics, and the target object referred to by the associated subject in the voice command is determined by combining the family relationship model. Then, the operation corresponding to the voice command and the target occupant referred to by the voice command is controlled and executed. On the one hand, this eliminates the reliance on a single main user for voice interaction; any occupant with a confirmed identity can issue voice commands. Intelligent and personalized services can be provided based on the voice command issuer, achieving a decentralized interaction effect of "whoever speaks, serves them." On the other hand, it reduces the burden of manual operation and explicit target designation for the driver, making voice interaction more in line with natural conversation in a family setting, improving convenience and driving safety.

[0028] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application, it can be implemented according to the contents of the specification. In order to make the above and other objects, features and advantages of this application more obvious, specific embodiments of this application are given below. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a schematic diagram of the voice interaction method described in the embodiments of this application.

[0031] Figure 2 This is a schematic diagram of the method for binding the voiceprint characteristics of each occupant in the vehicle to their occupant identity in an embodiment of this application.

[0032] Figure 3 This is a family relationship model generated when the passengers in this application are Zhang San, Li Si, and Zhang Xiaowu.

[0033] Figure 4 This is a schematic diagram of a method for confirming the identity of each passenger according to an embodiment of this application.

[0034] Figure 5 This is a schematic diagram of a method for confirming the identity of each occupant according to another embodiment of this application.

[0035] Figure 6 This is a schematic diagram of a method for confirming the identity of each passenger according to another embodiment of this application.

[0036] Figure 7 This is a family relationship model generated when the passengers in this application are Zhang San and Grandma Wang.

[0037] Figure 8 This is a family relationship model generated when the passengers in this application are Zhang San, Li Si, Zhang Xiaowu, and Grandma Wang.

[0038] Figure 9 This is a schematic diagram of the voice interaction method described in Example 1 of this application.

[0039] Figure 10 This is a schematic diagram of the voice interaction method described in Example 2 of this application.

[0040] Figure 11 This is a schematic diagram of the voice interaction method described in Example 3 of this application.

[0041] Figure 12 This is a schematic diagram of the voice interaction method described in Example 4 of this application.

[0042] Figure 13 This is a schematic diagram of the architecture of the voice interaction system according to an embodiment of this application.

[0043] Figure 14 This is a schematic diagram of the vehicle architecture according to an embodiment of this application. Detailed Implementation

[0044] To make the technical problems, solutions, and beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0045] The prefixes such as "first" and "second" used in this application embodiment are merely for distinguishing different descriptive objects and do not limit the position, order, priority, quantity, or content of the described objects. The use of ordinal numbers and other prefixes used to distinguish descriptive objects in this application embodiment does not constitute a limitation on the described objects. The description of the described objects is given in the claims or the context of the embodiments, and should not constitute unnecessary restrictions due to the use of such prefixes. Furthermore, in the description of this embodiment, unless otherwise stated, "multiple" means two or more.

[0046] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " signifies "or," for example, A / B can represent A or B.

[0047] With the rapid development of automotive intelligence and connectivity technologies, in-vehicle voice systems, with their contactless interaction characteristics, have become a core standard feature of intelligent connected vehicles, greatly improving the ease of operation and driving safety during the driving process.

[0048] Current in-vehicle voice systems generally adopt a centralized interaction model, with the interaction logic designed around the driver, prioritizing responses to voice commands issued by the driver. Commands from other occupants often require specific wake words, physical button presses, or manual switching of the interaction target to be recognized by the system. While this interaction model can meet basic usage needs in single-occupant scenarios, it has significant limitations in complex multi-occupant driving scenarios: the in-vehicle voice system lacks the ability to accurately locate the source of voice commands and cannot intelligently distinguish the intended recipient of the command, easily leading to mis-triggering or inconsistent responses. For example, when a child in the back seat issues the command "Open my window," the system cannot accurately determine that "I" refers to the child in the back seat, potentially incorrectly executing the operation of opening the driver's side window. This not only fails to meet the actual needs of the occupants but may also pose a driving safety hazard.

[0049] Meanwhile, the semantic understanding capabilities of existing in-vehicle voice systems are limited to the recognition and execution of isolated commands, lacking a deep understanding of the social relationships and interactive contexts among in-vehicle occupants. In actual driving scenarios, occupant voice commands often contain information related to interpersonal relationships and scenario needs. For example, a command like "Turn up the air conditioning temperature, Grandma is cold" contains not only the core operational need of "turning up the air conditioning temperature," but also implicitly refers to "Grandma" and the reason for the action of "being cold." However, existing in-vehicle voice systems can only parse and execute the basic command of "turning up the air conditioning temperature," unable to identify the specific occupant corresponding to "Grandma," nor can they optimize the adjustment strategy based on the occupant's physiological characteristics (such as the elderly's sensitivity to temperature). Ultimately, they can only provide standardized, mechanical responses, making it difficult to achieve truly intelligent and humanized services.

[0050] In summary, current in-vehicle voice systems are insufficient in terms of command recognition accuracy and semantic understanding depth in multi-occupant interaction scenarios, and can no longer meet users' advanced needs for intelligent and personalized in-vehicle voice interaction.

[0051] Based on the above application scenarios, this application proposes a voice interaction method.

[0052] Figure 1 This is a schematic flowchart illustrating a voice interaction method provided in an embodiment of this application. The method is applicable to vehicles, specifically to vehicle voice systems. The method includes the following steps.

[0053] S1. When the vehicle starts, the voiceprint characteristics of each passenger in the vehicle are bound to their passenger identity, and a family relationship model is generated based on the identities of each passenger in the vehicle.

[0054] For details, see Figure 2 Methods for linking the voiceprint characteristics of each occupant in the vehicle to their occupant identity include: S11. When the vehicle starts, obtain the voices of each passenger in the vehicle; S12. Extract the voiceprint features from the speech of each passenger and associate the voiceprint features of each passenger with their passenger identity.

[0055] For example, the voice of passenger Zhang San is obtained, Zhang San's voiceprint features are extracted, and Zhang San's voiceprint features are associated with Zhang San's identity information. For example, Zhang San's voiceprint features are associated with Zhang San's identification (hereinafter referred to as ID). Through the ID associated with the voiceprint features, it can be known that the passenger's identity is Zhang San.

[0056] In this embodiment, by binding voiceprint features with occupant identities, the in-vehicle voice system can automatically match and identify occupants based on voiceprint features the instant they issue a voice command, without relying on additional operations such as wake-up word switching or physical button triggering. When receiving a voice command containing occupant relationships, the system can quickly locate the command issuer based on voiceprint features. This binding relationship technically breaks the dependence of the traditional centralized interaction mode on the driver user, effectively solving the problem of ambiguous command sources in multi-occupant scenarios. For example, when a child in the back seat issues the voice command "Open my window," the voiceprint features in the voice command can be used to directly determine that the voice command issuer is a child in the back seat, accurately executing the corresponding window opening operation, completely avoiding the risk of misidentification and misresponse, and improving the accuracy and efficiency of command execution in multi-occupant interaction scenarios.

[0057] In this embodiment, by binding voiceprint features to occupant identity, the steps of manually specifying identity and switching interaction objects are eliminated, achieving seamless interaction of "identity confirmation upon command issuance." This lowers the operational threshold for occupants, making it particularly suitable for groups with lower acceptance of complex operations, such as children and the elderly. Simultaneously, occupants do not need to be distracted by identity switching operations, allowing them to focus on the driving process and reducing driving safety hazards caused by operating the in-vehicle voice system, thus balancing the convenience of voice interaction with driving safety.

[0058] Specifically, the family relationship model is a graph structure that includes personnel nodes and relationship edges. Each personnel node corresponds to a passenger, and the relationship edges are directed edges (e.g., relationship edges with arrows) used to represent the family relationships between passengers, such as spousal relationships, father-son relationships, etc.

[0059] For example, the passengers in this vehicle are: Zhang San (the vehicle owner), Li Si (Zhang San's wife), and Zhang Xiaowu (Zhang San's son). A family relationship model is generated based on each passenger's identity as follows: Figure 3 As shown. Figure 3 The family relationship model clearly illustrates the pairwise family relationships among the three occupants in the car. This model helps to understand the relationships between the "father" and "son," the "mother" and "son," and the "husband" and "wife." It should be noted that this family relationship model is a strongly connected graph, meaning that a path exists between any two occupants, indicating that this is a core family unit.

[0060] Specifically, the family relationship model is in list form. The list data includes member identity identifiers (such as USR_001), member names (such as Zhang San), the identity identifier of the associated object (such as Li Si) (USR_002), and the identity label corresponding to the associated object (such as husband).

[0061] For example, the passengers in this vehicle are: Zhang San (the vehicle owner), Li Si (Zhang San's wife), and Zhang Xiaowu (Zhang San's son). The family relationship model generated based on each passenger's identity is shown in Table 1. Table 1 clarifies the two-way relationship between each member and other members. For example, from Zhang San's (USR_001) perspective, he is "husband" or "father" to Li Si (USR_002), and "father" to Zhang Xiaowu (USR_003).

[0062] Table 1

[0063] In this embodiment, a family relationship model for the current trip is generated based on the identities and kinship relationships of each passenger in the vehicle. This model presents the identity attributes and kinship connections of the passengers in a structured form. This family relationship model provides the in-vehicle voice system with an intuitive basis for relationship recognition, enabling it to accurately parse complex commands containing multiple layers of kinship references. For example, when a passenger issues the command "Turn on the seat heater for Grandma," the system can quickly locate the kinship connections between "Grandma" and other passengers through the family relationship model, matching the corresponding passenger identity and avoiding erroneous responses due to ambiguous kinship terms, thus significantly improving the accuracy of command parsing.

[0064] In one embodiment of this application, see Figure 4 The methods for identifying each passenger include: S11. At the start of this driving operation, acquire facial images of each passenger inside the vehicle; S12. Match each occupant's face image with the pre-stored face images; S13. When the occupant's face image is matched with the pre-stored face image, the member identity identifier associated with the pre-stored face image is obtained. S14. Based on member identity identifiers, determine the identity of passengers through a pre-stored database of family member relationships of vehicle owners.

[0065] In this embodiment, facial images of all occupants in the vehicle are acquired all at once upon vehicle startup. Through facial image matching, facial information is automatically collected and compared with pre-stored facial images after occupants enter the cabin and take their seats, eliminating the need for manual input or wake-up triggering by occupants. This achieves seamless and automated batch identity recognition. After successful matching, occupant identity information is directly associated with their identification, avoiding the shortcomings of traditional voiceprint recognition which is susceptible to interference from environmental noise, changes in voice due to colds, etc., significantly improving the stability and accuracy of identity recognition. Furthermore, this method moves the identity recognition process forward to the vehicle startup stage, covering the identity recognition needs of all occupants in the vehicle. This avoids the delay in identity verification when occupants issue commands during driving, achieving seamless interaction of "accurate identity recognition upon command issuance" throughout the entire driving cycle, and solving the problem of identity confusion in multi-occupant scenarios.

[0066] Specifically, in this embodiment of the application, facial images of each passenger are captured by a camera installed inside the vehicle.

[0067] Specifically, in this embodiment, the data in the vehicle owner's family member relationship database includes member identity identifiers (e.g., USR_001), member names (e.g., Zhang San), the identity identifier of the associated object (e.g., Li Si) (USR_002), and the identity tag corresponding to the associated object (e.g., husband). It should be noted that the vehicle owner's family member relationship database is pre-entered by the vehicle owner or stored after system verification, and contains explicit kinship attributes between members (e.g., father and son, husband and wife, siblings, etc.), which can avoid errors that may arise from the algorithm's autonomous inference of relationships.

[0068] It should be noted that the member identification is a number or / and character identifier assigned to each member entered into the vehicle owner's family member relationship database, used to distinguish different individual members in the database.

[0069] In a specific embodiment of this application, the database of family member relationships of the vehicle owner is represented in the form of a list.

[0070] For example, the Zhang family, whose owner is Zhang San, has a family member relationship database as shown in Table 2.

[0071] Table 2

[0072] The representation in Table 2 clarifies the two-way relationship between each member and other members. For example, from Zhang San's (USR_001) perspective, he is "husband" or "father" to Li Si (USR_002), and "father" to Zhang Xiaowu (USR_003).

[0073] For example, the scenario is: Zhang San takes his mother-in-law (Grandma Wang) back to her home alone.

[0074] The passengers are: Zhang San and Granny Wang. When the vehicle starts, two occupant facial images are acquired, namely the first occupant facial image and the second occupant facial image.

[0075] The face image of the first occupant is matched with the pre-stored image. When the face image of the first occupant is matched with the pre-stored face image, the member identity identifier USR_001 corresponding to the pre-stored face image is obtained. The member identity identifier USR_001 is matched with the member identity identifier in the pre-stored vehicle owner family member relationship database (as shown in Table 2). The member identity identifier USR_001 matches the member identity identifier USR_001 in the pre-stored vehicle owner family member relationship database, and the identity of the first occupant is identified as Zhang San.

[0076] The second occupant's face image is matched with a pre-stored image. During the matching process, the member identity identifier USR_004 corresponding to the pre-stored face image is obtained. The member identity identifier USR_004 is then matched with the member identity identifiers in the pre-stored vehicle owner family member relationship database (as shown in Table 2). The occupant identity identifier USR_004 matches the member identity identifier USR_004 in the pre-stored vehicle owner family member relationship database, identifying the second occupant as Granny Wang. Granny Wang is the mother-in-law of the first occupant Zhang San, and Zhang San is the son-in-law of Granny Wang.

[0077] In one embodiment of this application, see Figure 5 The methods for identifying each crew member include: S11. At the start of this driving operation, acquire facial images of each passenger inside the vehicle; S12. Match each occupant's face image with the pre-stored face images; S13. If a passenger's face image does not match a pre-stored face image, the passenger shall be added as a new passenger. S14. Set up new passenger information, which includes new passenger identification, new passenger name, and the relationship between the new passenger and the vehicle owner and other family members of the vehicle owner; S15. Obtain the identity of the newly added crew member based on the set information.

[0078] In this embodiment, when a passenger's face image that does not match a pre-stored face image is detected, that passenger is treated as a new passenger, and their identity identifier, name, and kinship information are set. The new passenger's identity is then determined based on this information. This approach overcomes the limitations of traditional identity recognition's "pre-entered fixed list." For new passengers outside the pre-stored face database, a complete passenger identity profile is constructed by supplementing their identity identifier, name, and kinship information. The constructed passenger identity profile effectively identifies new users. This solves the problem that traditional face recognition can only match existing users and cannot identify new users. Furthermore, after identifying a new passenger, based on their entered identity information and kinship, personalized services can be provided based on their voice commands when they act as the issuer. This allows the in-vehicle voice system to provide personalized services to every family member and temporary passenger, moving beyond the limitation of "only supporting pre-entered personnel" and upgrading services from "partial passenger-specific" to "all-passenger-specific."

[0079] In this embodiment, the newly added passenger information includes kinship attributes with the vehicle owner and other family members. This information, once entered, can be directly added to the family member relationship database, further enriching the cognitive dimension of social relationships among passengers. For example, after the identity information of a temporary passenger, such as an "aunt," is entered, when the vehicle issues the command "Aunt, open the window a little smaller," the system can quickly locate the corresponding passenger and execute the operation. This feature makes the in-vehicle voice system's understanding of voice commands containing kinship references more comprehensive and accurate, providing data support for continuously optimizing user-friendly interactive services.

[0080] For example, the scenario is: Zhang San, his wife (Li Si), and son (Zhang Xiaowu) return to their hometown to visit relatives, and Zhang San's sister (Zhang Si) also returns to their hometown to visit relatives with Zhang San's family.

[0081] The passengers are: Zhang San, Li Si, Zhang Xiaowu, and Zhang Si. When the vehicle starts, facial images of four occupants are acquired, namely the facial image of the first occupant, the facial image of the second occupant, the facial image of the third occupant, and the facial image of the fourth occupant.

[0082] The face image of the first occupant is matched with the pre-stored image. When the face image of the first occupant is matched with the pre-stored face image, the member identity identifier USR_001 corresponding to the pre-stored face image is obtained. The member identity identifier USR_001 is matched with the member identity identifier in the pre-stored vehicle owner family member relationship database (as shown in Table 2). The member identity identifier USR_001 matches the member identity identifier USR_001 in the pre-stored vehicle owner family member relationship database, and the identity of the first occupant is identified as Zhang San.

[0083] The second occupant's facial image is matched with a pre-stored image. During this matching process, the corresponding member identification identifier USR_002 is obtained. This identifier USR_002 is then matched with member identification identifiers in a pre-stored vehicle owner family member relationship database (as shown in Table 2). This match confirms the second occupant's identity as Li Si, who is the wife of the first occupant Zhang San, and the husband of the second occupant Zhang Xiaowu.

[0084] The face image of the third passenger is matched with the pre-stored image. When the face image of the passenger is matched with the pre-stored face image, the member identity identifier USR_003 corresponding to the pre-stored face image is obtained. The member identity identifier USR_003 is matched with the member identity identifier in the pre-stored vehicle owner family member relationship database (as shown in Table 2). The member identity identifier USR_003 matches the member identity identifier USR_003 in the pre-stored vehicle owner family member relationship database, and the identity of the third passenger is identified as Zhang Xiaowu. The third passenger Zhang Xiaowu is the son of the first passenger Zhang San and the second passenger Li Si. The first passenger Zhang San is the father of the third passenger Zhang Xiaowu, and the second passenger Li Si is the mother of the third passenger Zhang Xiaowu.

[0085] The fourth occupant's facial image is matched with a pre-stored image. If the fourth occupant's facial image does not match the pre-stored image, the fourth occupant's facial image is used as a new occupant. The fourth occupant's identity information is set, including the identity tag USR_005, the name Zhang Si, and the relationship with the vehicle owner and other occupants of the vehicle owner's family (e.g., the identity tag with Zhang San is sister, and the identity tag with Zhang Xiaowu is aunt, etc.). According to the set fourth occupant information, the fourth occupant's identity is Zhang Si. The fourth occupant Zhang Si is the sister of the first occupant Zhang San, the sister-in-law (or sister-in-law of the husband) of the second occupant Li Si, and the aunt of the third occupant Zhang Xiaowu. The first occupant Zhang San is the brother of the fourth occupant Zhang Si, the second occupant Li Si is the sister-in-law of the fourth occupant Zhang Si, and the third occupant Zhang Xiaowu is the nephew of the fourth occupant Zhang Si.

[0086] In one embodiment of this application, see Figure 6The methods for identifying each passenger include: S11. At the start of this driving operation, acquire facial images of each passenger inside the vehicle; S12. Match each occupant's face image with the pre-stored face images; S13. When a passenger's facial image does not match a pre-stored facial image, the passenger is added as a new passenger; the information of the new passenger includes the new passenger's identity identifier, the new passenger's name, and the relationship between the new passenger and the vehicle owner and other members of the vehicle owner's family; S14. Set up new passenger information and store it in the vehicle owner's family member relationship database, and store the new passenger's facial image; S15. Associate the newly added occupant's facial image with the newly added occupant's identity identifier; S16. Obtain the identity of the newly added crew member based on the set information.

[0087] In this embodiment, when an unmatched occupant facial image is detected, on the one hand, for newly added occupants outside the pre-stored facial database, a complete occupant identity profile is constructed by supplementing information such as identity identifiers, names, and kinship relationships. The newly constructed occupant identity profile can effectively identify the new user. On the other hand, the information of the newly added occupant is stored in the vehicle owner's family member relationship database, updating the database. This approach overcomes the limitations of traditional identity recognition's "pre-entered fixed list," dynamically adapting to changes in family occupant structure (e.g., new family members, temporary passengers, etc.), expanding the identity database without relying on professional personnel or complex operations, improving practicality and adaptability. Furthermore, after associating and storing the facial image of a newly added occupant with their identity identifier, in any subsequent ride scenario, the corresponding occupant identity information can be retrieved directly through facial image matching, achieving seamless identity verification without repeatedly entering occupant names, kinship relationships, etc., significantly shortening recognition time and improving the user experience for drivers and passengers.

[0088] For example, the scenario is: Zhang San, his wife (Li Si), and son (Zhang Xiaowu) return to their hometown to visit relatives, and Zhang San's sister (Zhang Si) also returns to their hometown to visit relatives with Zhang San's family.

[0089] The passengers are: Zhang San, Li Si, Zhang Xiaowu, and Zhang Si. The facial images were matched against pre-stored facial images. The facial images of Zhang San, Li Si, and Zhang Xiaowu all matched the pre-stored facial images, and their identities were determined using a pre-stored database of vehicle owner family member relationships (as shown in Table 2). Zhang Si's facial image did not match the pre-stored facial image; therefore, Zhang Si's information was stored in the vehicle owner family member relationship database (as shown in Table 2), resulting in an updated database (as shown in Table 3). Zhang Si's facial image was then associated with his identity identifier.

[0090] Table 3

[0091] In one embodiment of this application, the method for identifying each occupant further includes: When a passenger's facial image does not match a pre-stored facial image, retrieve the voice of the newly added passenger; Extract the voiceprint features from the voice of the newly added passenger and bind the voiceprint features to the identity of the newly added passenger.

[0092] In this embodiment, for newly added passengers outside the pre-stored face database, the voice of the newly added passenger is acquired, and the voiceprint features of the newly added passenger are bound to their passenger identity. On the one hand, when the newly added passenger acts as the issuer of voice commands, the identity of the newly added passenger can be quickly located based on the voiceprint features in the voice commands, so as to adapt to the voice command response of the newly added passenger and realize intelligent and personalized voice interaction services for the newly added passenger. On the other hand, the identity of the passenger can also be identified through voiceprint features, effectively avoiding the problem of misjudgment caused by similar facial features, realizing high-precision and high-uniqueness authentication of the identity of the newly added passenger, and improving the reliability of the identity recognition results.

[0093] Specifically, the method for identifying occupants using voiceprint features involves matching each occupant's voiceprint features with pre-stored voiceprint features to obtain a member identity identifier associated with the pre-stored voiceprint features. Based on this member identity identifier, the occupant's identity is determined through a pre-stored database of family member relationships of the vehicle owner. Voiceprint feature-based occupant identification supports rapid identity verification in multiple scenarios. For example, during driving, occupants do not need to actively cooperate with facial recognition; they can trigger voiceprint feature comparison simply by using voice commands (such as "turn on the air conditioning" or "play music"), quickly completing identity verification and adapting to personalized services. In scenarios where the driver's or passenger's hands are occupied and they cannot operate the in-vehicle terminal, identity verification can be completed through voice interaction, achieving contactless identity recognition and improving operational convenience and user experience during driving.

[0094] In one embodiment of this application, the method for generating a family relationship model based on the identity of each passenger is as follows: extract the relationship between the identities of each passenger from a pre-stored database of relationships between family members of the vehicle owner (such as the database of relationships between family members of the vehicle owner shown in Table 2 or Table 3), and generate a family relationship model based on the extracted relationship between the identities of each passenger.

[0095] In this embodiment, the kinship relationship graph between passengers is generated directly from a pre-stored database of vehicle owner family member relationships, ensuring the accuracy and authority of the identity association information in the graph. For example, after generating a family relationship model by extracting the relationship link of "vehicle owner-wife-mother-in-law", the target member can be directly and accurately located when parsing the instruction "help mother-in-law turn down the air conditioning", eliminating erroneous instruction responses caused by relationship inference errors.

[0096] For example, the passengers in this ride are: Zhang San (the car owner), Li Si (Zhang San's wife), and Zhang Xiaowu (Zhang San's son). A family relationship diagram is generated using the method described above, as follows: Figure 3 As shown. Figure 3 The family relationship model clearly illustrates the bidirectional family relationships between the three occupants in the car. This model helps to understand the relationships between the "father" and "son," the "mother" and "son," and the "husband" and "wife." It should be noted that this family relationship model is a strongly connected graph, meaning that a path exists between any two occupants, indicating that this is a core family unit.

[0097] For example, the passengers in this ride are: Zhang San (the car owner) and Granny Wang (Zhang San's mother-in-law). A family relationship model is generated based on the above method as follows: Figure 7 As shown. Figure 7 The family relationship model shown clearly illustrates the two-way family relationship between the two occupants in the car. Through this family relationship model, the relationship between the "son-in-law" and the "mother-in-law" in the car can be understood.

[0098] S2. Upon receiving a voice command, the identity of the voice command sender is confirmed based on the voiceprint characteristics of the voice command.

[0099] In this embodiment, by binding the occupant's voiceprint characteristics to their identity, the voiceprint characteristics of the voice command are quickly matched with pre-stored occupant voiceprint characteristic-identity binding data. This instantly confirms the identity of the command issuer without requiring additional wake words or manual switching of the interaction object. On one hand, this fundamentally breaks the limitations of the traditional centralized interaction mode, effectively solving the pain points of "unclear command source and ambiguous intent" in multi-occupant scenarios. For example, when multiple commands such as "open my window" are issued from inside the vehicle, the voiceprint characteristics can accurately distinguish between different command issuers such as the driver, rear-seat children, and front passenger, and execute the corresponding window opening operation accordingly, eliminating the risk of accidental triggering or response. On the other hand, this identity verification method requires no additional operation from the occupant; complete identity verification is achieved solely through natural language commands, simplifying the interaction steps in multi-occupant scenarios. Occupants do not need to be distracted by actions such as "identity switching" or "wake word repetition," allowing them to focus on the driving process and reducing driving safety hazards caused by operating the in-vehicle voice system. At the same time, groups with lower comprehension levels for complex operations, such as the elderly and children, can easily achieve efficient interaction with the in-vehicle voice system, significantly improving the ease of use and universality of the in-vehicle voice system.

[0100] S3. Identify the associated subject in the voice command, and determine the target occupant referred to by the associated subject based on the identity of the associated subject and the voice command issuer, according to the family relationship model.

[0101] In this embodiment, by identifying the associated subject in the voice command (such as "father," "wife," "son," "grandfather," "brother," etc.), combined with the voiceprint feature to confirm the identity of the voice command issuer, and relying on the kinship link between passengers in the family relationship model, the target member pointed to by the associated subject in the voice command can be accurately located. This technical feature solves the core pain point of traditional in-vehicle voice systems being unable to parse the "referring object of kinship titles." For example, when the voice command issuer (child) gives the command "Turn down the air conditioning for mom," the in-vehicle voice system can quickly match the issuer's mother as the target member through the associated subject "mom" + issuer identity + family relationship model, and then execute the air conditioning adjustment operation for the corresponding seat, eliminating false responses caused by ambiguous referencing.

[0102] For example, the passengers in this ride are: Zhang San (the car owner), Li Si (Zhang San's wife), Zhang Xiaowu (Zhang San's son), and Granny Wang (Zhang San's mother-in-law). A family relationship model is generated based on each passenger as follows: Figure 8 As shown.

[0103] When Zhang Xiaowu issues the command "Turn down the air conditioner for Mom", the system searches the family relationship model for a path that starts with Zhang Xiaowu and has the relationship edge labeled "Mom". The endpoint of this path is Li Si. Therefore, "Mom" refers to passenger Li Si, and the system can then perform the air conditioner adjustment operation (e.g., turn down the air conditioner fan speed) for passenger Li Si's seat.

[0104] S4. Control the execution of operations corresponding to voice commands and the target occupants indicated by the voice commands.

[0105] In this embodiment, the voice interaction method described above binds the voiceprint features of each occupant to their identity upon vehicle startup. A family relationship model is generated based on the identities of each occupant. The source of the voice command is accurately located based on the voiceprint features, and the target object indicated by the associated subject (e.g., husband, wife, son) in the voice command is determined using the family relationship model. The operation corresponding to the voice command and the target occupant indicated by the associated subject in the voice command is then controlled. On one hand, this eliminates reliance on a single main user for voice interaction; any occupant with a confirmed identity can issue voice commands. Intelligent and personalized services can be provided based on the voice command issuer, achieving a decentralized interaction effect of "whoever speaks, serves them." On the other hand, it reduces the burden of manual operation and explicit target designation for the driver, making voice interaction more natural and suitable for family scenarios, thus improving convenience and driving safety.

[0106] In one embodiment of this application, the method further includes: If a new passenger boards the vehicle during this trip, or if a new passenger boards the vehicle at the next start of the trip, the identity of the new passenger shall be identified. The relationship between the new passenger and other passengers is incorporated into the family relationship model, and the family relationship model is updated.

[0107] In this embodiment, new passengers are identified during the journey (when boarding midway) or at the next start of the journey, no longer limited to a fixed identification node for a single journey, and can cover all scenarios of passengers getting on and off the vehicle. Whether it is a family member traveling temporarily (such as the owner's parents traveling temporarily) or a newly added long-term family member (such as a newborn traveling for the first time after growing up), the initial identification and filing of the new passenger's identity can be completed through facial image matching. This solves the problem that traditional solutions can only identify pre-stored passengers and cannot adapt to changes in family members, allowing the identity recognition system to dynamically expand as the composition of family members changes.

[0108] In this embodiment, after identifying a new passenger, the relationship between the new passenger and existing passengers in the vehicle (such as "owner-owner's siblings," "parents-children," etc.) is entered and updated in the family relationship model, eliminating the need for the vehicle owner to manually maintain the database. This breaks the limitations of traditional family relationship models that require "static entry and manual maintenance," effectively adapting to dynamic driving scenarios such as the addition or removal of family members, picking up passengers mid-journey, and temporary rides. It ensures that the family relationship model always covers all driving-related personnel and their relationships, guaranteeing the real-time nature and completeness of family relationship data. Compared to static graphs, dynamically updated graphs can better reflect changes in passenger relationships in real time, significantly improving the accuracy of semantic understanding in multi-member, dynamic scenarios. This enables the in-vehicle voice system to accurately understand complex kinship referencing commands involving newly added passengers, achieving precise command execution.

[0109] Based on the updated family relationship model, complex kinship referencing instructions, including those involving newly added members, can be accurately parsed during the current trip (after boarding midway) or the next trip. For example, the identity information of "aunt" entered during the first trip can be directly located through the family relationship model (the target object referred to by the associated subject "aunt") when a passenger issues the instruction "turn on the seat ventilation for aunt" during the next trip, thus achieving accurate execution of the instruction. This feature allows the relationship recognition capability of the family relationship model to continuously improve with the frequency of use, gradually building an intelligent interaction model adapted to the specific needs of families.

[0110] For example, scenario: Zhang San, his wife (Li Si), and son (Zhang Xiaowu) are going back to their hometown to visit relatives. On the way, Zhang Si (Zhang San's sister) gets on the bus and goes back to their hometown with Zhang San's family.

[0111] At the start of this trip, the occupants were: Zhang San, Li Si, and Zhang Xiaowu. A family relationship model was generated based on the identities of each occupant (Zhang San, Li Si, Zhang Xiaowu) (e.g., Figure 3 During this journey, Zhang Si boards the bus midway, changing the passengers to: Zhang San, Li Si, Zhang Xiaowu, and Zhang Si. A family relationship model is generated based on the identities of each passenger (Zhang San, Li Si, Zhang Xiaowu, and Zhang Si). Figure 8 Based on the updated family relationship model, after Zhang Si gets on the bus, Zhang San, Li Si, Zhang Xiaowu, and Zhang Si can all accurately execute voice commands through voice interaction, achieving a decentralized interaction effect of "whoever speaks serves whom".

[0112] In one embodiment of this application, the method further includes: if, within a set time period, a member in the family relationship model takes fewer rides than a set number of times, then that member and the relationship between them and other members are removed from the family relationship model.

[0113] Information about members who haven't ridden in the car for a long time loses its practical value over time. If it's continuously retained in the family relationship model, it will make the model bloated and may lead to invalid matches when the system parses commands (e.g., recognizing a relative's title but finding no corresponding passenger). In this embodiment, by setting time and ride frequency thresholds, long-term inactive members and their relationship links are automatically filtered and deleted. This avoids the family relationship model becoming bloated due to the accumulation of redundant member information, reducing the computing load and storage resource consumption of the in-vehicle terminal and ensuring the response speed of voice interaction commands. It also ensures that the family relationship model maintains a strong correlation with current and recent ride passengers, ensuring the matching accuracy between the associated subject and the target passenger during voice command parsing and avoiding response errors caused by invalid data interference. Furthermore, through an automated frequency filtering and deletion mechanism, the graph data is autonomously iterated and updated, completing redundant data cleanup without user intervention, significantly reducing the user's manual maintenance costs and improving the intelligence and ease of use of the in-vehicle voice system.

[0114] For example, within 30 days, such as Figure 6 In the family relationship model shown, Grandma Wang only took the bus once, which is less than the set number (2 times). Therefore, the relationship between Grandma Wang and Zhang San, Li Si, and Zhang Xiaowu will be removed from the family relationship model, and the family relationship model will become as follows: Figure 4 As shown.

[0115] Example 1, see Figure 9 This example provides a voice interaction method, the steps of which include: S1. When the vehicle starts, acquire facial images and voice recordings of each passenger inside the vehicle.

[0116] S2. Match each passenger's face image with the pre-stored face images.

[0117] S3. When the occupant's face image is matched with the pre-stored face image, the member identity identifier associated with the pre-stored face image is obtained.

[0118] S4. Match the member identity identifier with the member identity identifier in the pre-stored vehicle owner family member relationship database to determine the passenger identity.

[0119] S5. Extract the voiceprint features from the speech of each passenger and bind the voiceprint features of each passenger to the passenger's identity.

[0120] S6. Extract the relationships between the identities of each passenger from the pre-stored database of family member relationships of vehicle owners, and generate a family relationship model based on the extracted relationships between the identities of each passenger.

[0121] S7. Upon receiving a voice command, confirm the identity of the voice command sender based on the voiceprint characteristics of the voice command.

[0122] S8. Identify the associated subject in the voice command, and determine the target occupant referred to by the associated subject based on the identity of the associated subject and the voice command issuer, according to the family relationship model.

[0123] S9. Control the execution of operations corresponding to the target occupant indicated by the voice command and the associated subject.

[0124] In this example, at the start of the trip, facial images and voice information of each occupant are collected. A family relationship model is generated based on the identities of each occupant determined from the facial images. Voiceprint features from the occupants' voices are then linked to their identities. The source of the voice command is accurately located based on these voiceprint features, and the target object indicated by the associated subject is determined using the family relationship model. The operation corresponding to the voice command and the target occupant indicated by the associated subject is then executed. On one hand, this eliminates the reliance on a single main user for voice interaction; any occupant with a confirmed identity can issue voice commands. Intelligent and personalized services can be provided based on the voice command issuer, achieving a decentralized interaction effect of "whoever speaks, serves them." On the other hand, it reduces the burden of manual operation and explicit target designation for the driver, making voice interaction more natural and suitable for family-oriented conversations, thus improving convenience and driving safety.

[0125] Example 2, see Figure 10 This example provides a voice interaction method, the steps of which include: S1. When the vehicle starts, acquire facial images and voice recordings of each passenger inside the vehicle.

[0126] S2. Match each passenger's face image with the pre-stored face images.

[0127] S3. When the occupant's face image is matched with the pre-stored face image, the member identity identifier associated with the pre-stored face image is obtained.

[0128] S4. Match the member identity identifier with the member identity identifier in the pre-stored vehicle owner family member relationship database to determine the passenger identity.

[0129] S5. If a passenger's face image does not match a pre-stored face image, store the passenger's face image as a new passenger face image.

[0130] S6. Set up new passenger information and store it in the vehicle owner's family member relationship database. The new passenger information includes the new passenger's identity identifier, the new passenger's name, and the relationship between the new passenger and the vehicle owner and other family members.

[0131] S7. Associate the newly added passenger's facial image with the newly added passenger's identity identifier to confirm the identity of the newly added passenger.

[0132] S8. Extract the voiceprint features from the speech of each passenger and bind the voiceprint features of each passenger to the passenger's identity.

[0133] S9. Extract the relationships between the identities of each passenger from the pre-stored database of family member relationships of vehicle owners, and generate a family relationship model based on the extracted relationships between the identities of each passenger.

[0134] S10. Upon receiving a voice command, confirm the identity of the voice command sender based on the voiceprint characteristics of the voice command.

[0135] S11. Identify the associated subject in the voice command, and determine the target occupant referred to by the associated subject based on the identity of the associated subject and the voice command issuer, according to the family relationship model.

[0136] S12, Control the execution of operations corresponding to the voice commands and the target occupants indicated by the associated subject.

[0137] In this example, when a passenger's face image that does not match a pre-stored face image is detected, it is stored as a new passenger image. The new passenger's identity identifier, name, and kinship information are then entered and simultaneously updated in the vehicle owner's family member relationship database. This approach overcomes the limitations of traditional identity recognition methods that rely on a "pre-entered fixed list," dynamically adapting to changes in the family passenger structure (e.g., adding new family members, temporary passengers, etc.). It expands the identity database without relying on professional personnel or complex operations, improving practicality and adaptability. Furthermore, after confirming the identity of a new passenger, personalized services can be provided based on their entered identity information and kinship. When the new passenger is the one issuing commands, personalized services can be implemented according to their voice commands, extending the personalized services of the in-vehicle voice system to every family member and temporary passenger. This overcomes the limitation of "only supporting pre-entered personnel services" and upgrades the service from "partial passenger-specific" to "all-passenger-specific."

[0138] Example 3, see Figure 11 This example provides a voice interaction method, the steps of which include: S1. When the vehicle starts, acquire facial images and voice recordings of each passenger inside the vehicle.

[0139] S2. Match each passenger's face image with the pre-stored face images.

[0140] S3. When the occupant's face image is matched with the pre-stored face image, the member identity identifier associated with the pre-stored face image is obtained.

[0141] S4. Match the member identity identifier with the member identity identifier in the pre-stored vehicle owner family member relationship database to determine the passenger identity.

[0142] S5. Extract the voiceprint features from the speech of each passenger and bind the voiceprint features of each passenger to the passenger's identity.

[0143] S6. Extract the relationships between the identities of each passenger from the pre-stored database of family member relationships of vehicle owners, and generate a family relationship model based on the extracted relationships between the identities of each passenger.

[0144] S7. When a new passenger boards the vehicle, obtain the new passenger's facial image and voice. S8. Compare the new passenger's facial image with the pre-stored database of family member relationships of the vehicle owner to confirm the new passenger's identity.

[0145] S9. Extract the voiceprint features from the new passenger's speech and bind the new passenger's voiceprint features to the new passenger's identity.

[0146] S10. Incorporate the relationship between the new passenger and other passengers into the family relationship model to obtain a new family relationship model.

[0147] S11. Upon receiving a voice command, the identity of the voice command sender is confirmed based on the voiceprint characteristics of the voice command.

[0148] S12. Identify the associated subject in the voice command, and determine the target member referred to by the associated subject based on the identity of the associated subject and the voice command issuer, according to the family relationship model.

[0149] S13. Control the execution of operations corresponding to the voice commands and the target occupants indicated by the associated subject.

[0150] In this example, when a new passenger boards the vehicle during the journey, the system proactively collects the new passenger's facial image and voice information, completing identity verification without interrupting driving operations. This process overcomes the limitations of traditional methods that only support identity entry at the initial stage of driving, effectively adapting to dynamic driving scenarios such as picking up passengers mid-journey or temporary rides. It avoids issues like misrecognition of voice commands and chaotic responses caused by the addition of new passengers, ensuring the continuity of interactive services throughout the entire driving cycle. Simultaneously, the kinship relationship between the new passenger and existing passengers is added to the family relationship model in real time, generating an updated graph structure. This operation enables the in-vehicle voice system to accurately understand complex kinship reference commands involving new passengers, achieving precise command execution. Compared to a static graph, a dynamically updated graph can better reflect changes in passenger relationships in real time, significantly improving the accuracy of semantic understanding in multi-member, dynamic scenarios.

[0151] Example 4, see Figure 10This example provides a voice interaction method, the steps of which include: S1. Acquire facial images of each passenger upon the next vehicle start-up.

[0152] S2. Match each passenger's face image with the pre-stored face images.

[0153] S3. If a passenger's face image does not match a pre-stored face image, store the passenger's face image as a new passenger face image.

[0154] S4. Set up new passenger information and store it in the vehicle owner's family member relationship database. The new passenger information includes the new passenger's identity identifier, the new passenger's name, and the relationship between the new passenger and the vehicle owner and other passengers in the vehicle owner's family.

[0155] S5. Associate the newly added passenger's facial image with the newly added passenger's identity identifier, and obtain the newly added passenger's identity based on the set newly added passenger information.

[0156] S6. Obtain the voice of the new passenger, extract the voiceprint features from the new passenger's voice, and bind the voiceprint features of the new passenger to the passenger's identity.

[0157] S7. Add the relationship between the new passenger and other passengers to the family relationship model to obtain a new family relationship model.

[0158] S8. Upon receiving a voice command, confirm the identity of the voice command sender based on the voiceprint characteristics of the voice command.

[0159] S9. Identify the associated subject in the voice command, and determine the target member referred to by the associated subject based on the identity of the associated subject and the voice command issuer, according to the family relationship model.

[0160] S10, Control the execution of operations corresponding to the voice commands and the target occupants indicated by the associated subject.

[0161] In this example, the batch acquisition and matching of occupant facial images are completed at the point where the current trip ends and the next trip begins. This maintains the continuity of historical occupant identification and supports seamless registration of new occupants. For occupants not matched in the pre-stored facial database, the entire process, including facial image storage, identity information entry, and voiceprint feature binding, is automatically completed and updated simultaneously to the vehicle owner's family member relationship database. This mechanism overcomes the limitations of identity management within a single driving cycle and can adapt to the high-frequency changes in occupant structure during family travel (such as weekend family trips, weekday single commutes, and temporary rides with relatives), ensuring accurate and efficient occupant identification for every trip.

[0162] This application provides a voice interaction system suitable for vehicles. Figure 13 The diagram shown is a structural schematic of the voice interaction system.

[0163] The voice interaction system 100 includes: Binding module 101 is configured to bind the voiceprint features of each occupant to their occupant identity; The model building module 102 is configured to generate a family relationship model based on the identities of each passenger in the vehicle when the vehicle is started. The interaction module 103 is configured to: upon receiving a voice command, confirm the identity of the voice command issuer based on the voiceprint characteristics of the voice command; identify the associated subject in the voice command; and, based on the identities of the associated subject and the voice command issuer, determine the target occupant referred to by the associated subject according to a family relationship model. The control module 104 is configured to control the execution of operations corresponding to voice commands and the target occupant indicated by the associated subject.

[0164] In one embodiment of this application, the model building module 102 is further configured to: add the relationship between the new passenger and other passengers to the family relationship model and update the family relationship model.

[0165] In one embodiment of this application, the model building module 102 is further configured to: if a member in the family relationship model takes the bus less than a set number of times within a set time period, then delete that member and the relationship between them and other members from the family relationship model.

[0166] See also Figure 13 In one embodiment of this application, the voice interaction system 100 further includes: Storage module 105 stores a database of family member relationships of the vehicle owner and pre-stored facial images; The information collection module 106 is installed inside the vehicle and is used to acquire facial images and voices of each passenger inside the vehicle. The identity recognition module 107 is configured to: match each occupant's face image with a pre-stored face image; when the occupant's face image matches the pre-stored face image, obtain the member identity identifier associated with the pre-stored face image; and determine the occupant's identity based on the member identity identifier through a pre-stored vehicle owner's family member relationship database.

[0167] In one embodiment of this application, the information acquisition module 106 includes: Camera 1061 is used to acquire facial images of each passenger in the vehicle and send the acquired facial images to the identity recognition module 107; Microphone array 1062 is used to acquire the voice of each occupant in the vehicle and send the acquired voice to the binding module 101.

[0168] See also Figure 13 In one embodiment of this application, the voice interaction system 100 further includes a setting module 108, which is configured to set new member information.

[0169] In one embodiment of this application, the identity recognition module 107 is further configured to: when a passenger's face image does not match a pre-stored face image, treat the passenger as a new passenger and obtain the identity of the new passenger based on the set new member information.

[0170] In one embodiment of this application, the identity recognition module 107 is further configured to: store the newly added passenger information in the vehicle owner's family member relationship database, and store the newly added passenger's facial image; and associate the newly added passenger's facial image with the newly added passenger's identity identifier.

[0171] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0172] This application also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to execute the voice interaction method described in the above embodiments. This computer program can be installed in a vehicle system.

[0173] This application also provides a computer-readable storage medium storing program code that is executed by one or more processors. When the program code runs on the processor, it causes a device including one or more processors to perform the voice interaction method described in the above embodiments. The processor running this computer-readable storage medium can be installed in a vehicle system.

[0174] It should be understood that when the modules or units described herein are implemented using software, they can be implemented in whole or in part as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0175] This application provides a chip system including a processor, or the chip system including a memory and a processor, for calling computer programs or computer instructions stored in the memory to cause the processor to execute the voice interaction method involved in the above embodiments. The chip system can be a single chip or a chip module composed of multiple chips. This chip system can be installed in a vehicle system.

[0176] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the voice interaction method described in the above embodiments. This electronic device can be installed in a vehicle system.

[0177] This application provides a vehicle.

[0178] For example, see Figure 14 The vehicle 200 includes a memory 201, a processor 202, and a computer program 203 stored in the memory 201 and executable on the processor 202. When the processor 202 executes the computer program 203, it enables the processor 202 to implement the voice interaction method involved in the above embodiments.

[0179] Those skilled in the art will recognize that the modules, units, and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0180] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be covered. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A voice interaction method applied to a vehicle, characterized in that, The method includes: When the vehicle starts, the voiceprint characteristics of each passenger in the vehicle are bound to their identity, and a family relationship model is generated based on the identities of each passenger in the vehicle. Upon receiving a voice command, the identity of the voice command sender is confirmed based on the voiceprint characteristics of the voice command. Identify the associated subject in the voice command, and determine the target occupant referred to by the associated subject based on the identity of the associated subject and the voice command issuer, according to the family relationship model; Controls the execution of operations corresponding to voice commands and the target occupant indicated by the associated entity.

2. The method according to claim 1, characterized in that, The method further includes: If a new passenger boards the vehicle during this trip, or if a new passenger boards the vehicle at the next start of the trip, the identity of the new passenger shall be identified. The relationship between the new passenger and other passengers is incorporated into the family relationship model, and the family relationship model is updated.

3. The method according to claim 1, characterized in that, The methods for identifying each passenger include: At the start of this driving operation, facial images of each occupant inside the vehicle are acquired; Match each occupant's facial image with pre-stored facial images; When a passenger's facial image is matched with a pre-stored facial image, the passenger's identity identifier associated with the pre-stored facial image is obtained; Passenger identities are determined by using a pre-stored database of family member relationships of vehicle owners.

4. The method according to claim 3, characterized in that, The methods for identifying each passenger also include: If a passenger's facial image does not match a pre-stored facial image, that passenger will be added as a new passenger. The information for newly added passengers includes passenger identification, passenger name, and the relationship between the passenger and the vehicle owner and other family members of the vehicle owner. The identity of the new passenger is obtained based on the information set for the new passenger.

5. The method according to claim 4, characterized in that, The methods for identifying each passenger also include: When a passenger's facial image does not match a pre-stored facial image, the new passenger's information is stored in the vehicle owner's family member relationship database, and the new passenger's facial image is also stored. The newly added passenger facial images are associated with the newly added passenger identification.

6. The method according to claim 4, characterized in that, The methods for identifying each passenger also include: When a passenger's facial image does not match a pre-stored facial image, retrieve the voice of the newly added passenger; Extract the voiceprint features from the voice of the newly added passenger and bind the voiceprint features to the identity of the newly added passenger.

7. The method according to claim 3, characterized in that, The method for generating a family relationship model based on the identities of each passenger in the vehicle is as follows: extract the relationships between the identities of each passenger from a pre-stored database of family member relationships of the vehicle owner, and generate a family relationship model based on the extracted relationships between the identities of each passenger.

8. The method according to claim 1, characterized in that, The method further includes: if a member in the family relationship model takes the bus less than a set number of times within a set time period, then that member and the relationship between them and other members are removed from the family relationship model.

9. A voice interaction system applied to a vehicle, characterized in that, include: The binding module is configured to bind each occupant's voiceprint characteristics to their occupant identity. The model building module is configured to generate a family relationship model based on the identities of each passenger in the vehicle when the vehicle is started. The interaction module is configured to: upon receiving a voice command, confirm the identity of the voice command issuer based on the voiceprint characteristics of the voice command; identify the associated subject in the voice command; and, based on the identities of the associated subject and the voice command issuer, determine the target occupant referred to by the associated subject according to a family relationship model. The control module is configured to control the execution of operations corresponding to voice commands and the target occupant indicated by the associated subject.

10. A vehicle, characterized in that, The vehicle includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it causes the processor to implement the voice interaction method according to any one of claims 1 to 8; the vehicle includes the voice interaction system according to claim 9.