A vehicle-mounted multi-person quick-answer method, device, vehicle, medium and product

By combining visual and voice detection with in-car cameras and microphones, the position and movements of the users who are answering the questions are identified, solving the hardware dependency and rough judgment problems of in-car multi-person answering solutions, and achieving accurate answering order judgment and a user-friendly interactive experience.

CN120388359BActive Publication Date: 2025-09-12CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510886789.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-12
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Existing in-vehicle multi-person quick-answer solutions have problems such as strong hardware dependence, poor scene adaptability, and rough priority determination, making it difficult to accurately determine the order of quick answers in complex in-vehicle environments.

Method used

By collecting images inside the cabin, identifying the position and movements of users who are trying to answer, using the in-car camera and microphone for visual detection, and combining the timing correlation of lip movements and voice information, the order of answering is determined. A three-level answer priority determination rule is provided, including frame timing, hand-raising height and lip movement timing matching, reducing hardware costs and implementation difficulty.

Benefits of technology

It achieves accurate determination of the order of quick responses in complex in-vehicle environments, reduces hardware costs, improves the accuracy and fairness of quick response recognition, adapts to the participation of users of different age groups, and reduces noise interference and confusion errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388359B_ABST
    Figure CN120388359B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of in-vehicle human-computer interaction, and discloses an in-vehicle multi-person quick-answer method, device, vehicle, medium, and product. The method comprises: determining the position information of users participating in the quick-answer based on an image in the cabin; outputting a question; identifying the quick-answer action within each position information; determining a first quick-answer order based on the order in which the quick-answer actions appear in a picture frame; prompting users of different ranks in the first quick-answer order to answer the question in sequence; prompting users of the same rank to answer the question simultaneously, and collecting lip movement information and voice information during the simultaneous answering process; matching the lip movement information and voice information; determining the starting time frame of each user's lip movement information, and determining the second quick-answer order of users of the same rank based on the starting time frame of each user; and determining the quick-answer result based on the first quick-answer order, the second quick-answer order, and the voice information. The present invention reduces hardware dependence and improves the accuracy of priority determination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle-mounted human-computer interaction, and in particular to a vehicle-mounted multi-person quick-answer method, device, vehicle, medium and product. Background Art

[0002] With the development of smart cockpit technology, in-car multiplayer interactive games (such as knowledge quiz and song identification) have become increasingly important for enhancing the driving experience. Related technologies for implementing multiplayer quiz interactions in vehicles primarily rely on acoustic positioning solutions, biometric recognition solutions, visual recognition solutions, and hardware recognition solutions.

[0003] Acoustic localization solutions use microphone arrays to locate sound sources and determine the order of responses based on the order of the sound sources. However, the noisy environment inside a vehicle can easily cause sound source aliasing, resulting in poor localization accuracy and a high rate of misjudgment of the response order. Biometric recognition solutions typically use voiceprint recognition technology to bind user identities and then combine voice timestamps to determine the response order. This limitation is that it requires users to pre-register their voiceprint information, resulting in a complex identity binding process, making it less user-friendly for casual passengers and poorly adaptable to different scenarios. Visual recognition solutions currently disassemble camera frames, capture frames showing user responses, and then determine the response order based on the order of the frames. However, this solution cannot prevent multiple users from responding simultaneously in the same frame, making it more likely that the response order will be unclear. Hardware recognition solutions typically install multiple sensors on the vehicle that act as response buttons. Users must touch the sensors next to their seats to respond. While this solution makes it easier to determine the response order, the additional sensors increase vehicle cost and are unlikely to have other uses, resulting in a waste of money.

[0004] In summary, related technologies generally have three major constraints: strong hardware dependence, poor scene adaptability, and rough priority determination. Therefore, there is an urgent need for a quick-answer interaction solution with low hardware cost, high robustness, and adaptability to complex in-vehicle environments. Summary of the Invention

[0005] In view of this, the present invention provides a vehicle-mounted multi-person quick-answer method, device, vehicle, medium and product to solve the problems of strong hardware dependence, poor scene adaptability and rough priority determination.

[0006] In a first aspect, the present invention provides a method for multiple people in a vehicle to answer questions quickly, the method comprising: collecting images in the cockpit, and determining position information of users participating in the quick answer based on the images in the cockpit; outputting questions; identifying quick answer actions in each position information; determining a first quick answer order for each user participating in the quick answer based on the order in which the quick answer actions appear in the picture frame; for users with different ranks in the first quick answer order, prompting users to answer questions in sequence and collecting voice information of each user; for multiple users with the same rank in the first quick answer order, prompting users to answer questions at the same time, and collecting lip movement information and voice information of each user in the process of users with the same rank answering questions at the same time; matching the lip movement information with the voice information collected in the process of users with the same rank answering questions at the same time to obtain answer data of users with the same rank; determining the starting timing frame of the lip movement information of each user based on the answer data, and judging the second quick answer order of users with the same rank based on the starting timing frame of each user; determining the quick answer result based on the first quick answer order, the second quick answer order and the voice information of each user.

[0007] According to the above technical means, the present invention does not require the vehicle model to have a multi-sound zone configuration or additional quick-answer sensing device to locate the sound source or determine the timing of the quick answer, nor does it require the user to access additional equipment for interactive experience. It can directly use the existing cameras and microphones in the car to achieve multi-person quick-answer interaction, reducing hardware costs and implementation difficulties. The present invention has no restrictions on the identity of the participants in the car, does not need to record user information such as voiceprints or faces, and does not require users to learn complex and cumbersome postures or action instructions. Only a simple hand-raising action can participate in the quick answer, allowing drivers and passengers of different ages to participate easily. The present invention determines the order and position of the quick answer through real-time visual detection, is not affected by noise, and combines the temporal correlation of lip movement information and audio information to avoid confusion errors in the voice system and improve the accuracy of audio positioning and answer content recognition. This solves the problems of existing in-vehicle quick-answer solutions such as strong hardware dependence, poor scene adaptability, and rough priority determination.

[0008] In some optional embodiments, identifying the quick-answer action in each position information includes: identifying the arm angle and hand-raising height of the user in each position information; when the arm angle falls within a preset angle range and the hand-raising height is greater than a preset height threshold, determining that the quick-answer action occurs in the corresponding position.

[0009] According to the above technical means, the quick-answer action is determined based on raising the hand, and the dual constraints of the arm angle and the hand-raising height are used to judge whether the user has performed the specified hand-raising action, thereby improving the accuracy of action judgment.

[0010] In some optional embodiments, the method further includes: dividing the location information into areas, and visualizing the divided areas on the vehicle's central control screen so that participants can confirm their own answering areas.

[0011] According to the above-mentioned technical means, during the process of answering questions through quick-answering actions, the areas of each position are divided through visual detection and visualized on the central control screen for participants to confirm their own answering areas. On the one hand, it avoids the misjudgment of limb crossing during the answering process (the algorithm itself is based on individual matching of torso and limb detection to avoid cross-matching). The introduction of position area division can further guide users and enhance the accuracy of quick-answering action detection and distinction; on the other hand, through visualization, the user's gestures are guided to appear within the visible area, avoiding the failure of action judgment due to occlusion, thereby improving participation.

[0012] In some optional embodiments, the identifying of the arm angle and hand-raising height of the user in each position information includes: identifying the wrist joint point and elbow joint point of the user in the current position information; drawing the arm line connecting the wrist joint point and the elbow joint point; obtaining the arm angle in the current position information by calculating the angle between the arm line and the horizontal line of the cockpit screen; identifying the arm key points and head and torso key points of the user in the current position information, the arm key points being predefined landmark points on the arm, and the head and torso key points being predefined landmark points on the head or torso; and calculating the hand-raising height in the current position information based on the vertical distance between the arm key points and the head and torso key points.

[0013] Based on the above technical means, by identifying specific key points on the arms and torso, it is calculated whether the arm movement meets the constraints of arm angle and hand-raising height. The solution principle is simple and the calculation is accurate, providing a method for judging quick-response actions that takes into account both efficiency and accuracy.

[0014] In some optional embodiments, the calculating the hand-raising height in the current position information based on the vertical distance between the arm key point and the head and torso key point includes: calculating the vertical distance between the arm key point and the head and torso key point; determining the relative distance between the current position information and the camera, and determining a correction weight based on the relative distance; and determining the hand-raising height using the product of the correction weight and the vertical distance.

[0015] According to the above technical means, the measured arm height is corrected by calculating the corresponding correction weight based on the distance between the user and the camera, which solves the problem of distortion in the two-dimensional image measurement of the vehicle interior space for the first time. This method can eliminate the visual differences caused by the different distances between the user and the camera, making the calculation of the hand-raising height more accurate.

[0016] In some optional embodiments, when the camera is at the front position of the vehicle, determining the relative distance between the current position information and the camera, and determining the correction weight based on the relative distance, includes: obtaining the front row distance from the front row of the vehicle seat to the camera, and obtaining the back row distance from the back row of the vehicle seat to the camera; calculating the sum of the front row distance and the back row distance to obtain the total distance; judging whether the current position information is located in the front row or the back row; if the current position information is located in the front row, calculating the ratio of the front row distance to the total distance to obtain the correction weight; if the current position information is located in the back row, calculating the ratio of the back row distance to the total distance to obtain the correction weight.

[0017] According to the above technical means, when the camera is at the front of the vehicle, the front-seat users are closer to the camera, so the same physical height occupies a larger proportion of the frame, while the back-seat users are farther away from the camera, so the same physical height occupies a smaller proportion of the frame. The front-seat perspective magnification effect is eliminated by calculating the ratio of the front-seat distance to the sum of the distances, and the back-seat perspective reduction effect is compensated by calculating the ratio of the back-seat distance to the sum of the distances. By establishing a distance-weight mapping function, the spatial perspective height correction model converts the two-dimensional visual measurement values ​​in the vehicle into comparable height data in real physical space for the first time, fundamentally solving the fairness problem of quick-answering caused by seating position.

[0018] In some optional embodiments, the method of determining the first rush-answer order of each participating user based on the order in which the rush-answer actions appear in the picture frames includes: extracting the picture frames when each participating user appears to make a rush-answer action; determining the time sequence of each picture frame, and sorting each picture frame in chronological order from early to late to obtain a first sort; judging whether there is a target picture frame, wherein the target picture frame includes the rush-answer actions of multiple users; sorting the hand-raising heights of the rush-answer actions of different users in the target picture frame from high to low to obtain a second sort, wherein if the hand-raising heights of multiple target users in the target picture frame are also the same, the target users are defined as users of the same order; generating the first rush-answer order based on the first sort and the second sort.

[0019] Based on the above technical means, a three-level priority determination rule for quick responses is provided. The first-level priority determination rule includes the first and second levels, and the second-level priority determination rule includes the third level. The first level prioritizes frame timing and identifies the frame number where the hand-raising gesture first appears. The lower the frame number, the higher the priority, addressing 90% of common quick response scenarios. The second level prioritizes height. For the same frame, a spatial perspective correction model is used to calculate the actual hand-raising height. The one with the higher height value wins. Furthermore, combined with the third level based on lip movement timing matching, for the same frame and height, lip movement detection is used to locate the onset of pronunciation. The consistency of lip movement and speech timing is verified by audio segmentation. The first valid speaker is given priority for answering. This three-level progressive determination covers all dimensions, from millisecond-level timing (frame sequence) to behavioral characteristics (height) and finally to biometric characteristics (lip movement), addressing the failure of traditional acoustic solutions when multiple people are simultaneously responding. The first level directly determines 95% of scenarios (calculation time < 2ms), while 5% of extreme scenarios trigger the third level analysis, meeting the real-time requirements of in-vehicle systems (total latency ≤ 50ms). The three-level collaboration reduces the priority misjudgment rate to almost zero, significantly improving the accuracy of quick responses.

[0020] In some optional embodiments, identifying the quick-answer action in each location information includes: identifying whether the user's palm touches the designated area in each location information; when the user's palm touches the designated area, determining that the quick-answer action occurs at the corresponding position.

[0021] According to the above technical means, the system can also set up multiple designated answering areas in the car, such as setting touch areas on the door armrests near each seat, on the seat backs or on the center console. The system detects the position of the user's palm through image recognition technology and determines whether the palm is in contact with the designated area. When the system detects that the user's palm touches the designated area, it is considered that the user has completed the answering action. This way of answering is simpler and more direct than raising the hand to answer. The user only needs to touch the designated area to complete the answer without making an obvious hand-raising action. This method is more suitable for situations where space is limited in the car. Under the condition of limited space in the car, the answering method of touching the designated area can also reduce the system's misjudgment of the user's non-standard actions and improve the accuracy of answer recognition.

[0022] In some optional implementations, collecting the image inside the cabin includes: collecting the image inside the cabin using an RGB-IR type camera.

[0023] Using these technologies, the system captures in-cabin images using an RGB-IR camera. This camera, capable of both visible light (RGB) and infrared (IR) imaging, can clearly capture images of the cabin in a variety of lighting conditions, including low-light conditions at night. This ensures the system can accurately identify the user's quick-response gestures and lip movements.

[0024] In the second aspect, the present invention provides a vehicle-mounted multi-person quick-answer device, which includes: a position binding module for collecting images in the cockpit and determining the position information of users participating in the quick-answer based on the images in the cockpit; a question-generating module for outputting questions; an action recognition module for identifying the quick-answer actions in each position information; a first quick-answer order determination module for determining the first quick-answer order of each participating user according to the order in which the quick-answer actions appear in the picture frame; a first answer module for prompting users in different orders in the first quick-answer order to answer questions in turn and collecting voice information of each user; a second answer module for determining the first quick-answer order of each participating user according to the order in which the quick-answer actions appear in the picture frame; Multiple users with the same order in the sequence are prompted to answer questions at the same time, and the lip movement information and voice information of each user are collected during the process of users with the same order answering questions at the same time; a matching module is used to match the lip movement information with the voice information collected during the process of users with the same order answering questions at the same time to obtain the answer data of users with the same order; a second quick-answer order determination module is used to determine the starting timing frame of the lip movement information of each user according to the answer data, and judge the second quick-answer order of users with the same order according to the starting timing frame of each user; a result module is used to determine the quick-answer result based on the first quick-answer order, the second quick-answer order and the voice information of each user.

[0025] In the third aspect, the present invention provides a vehicle, comprising: a memory, a cockpit domain controller, an in-vehicle camera, an in-vehicle microphone, an in-vehicle central control screen and an in-vehicle speaker; the memory, the in-vehicle camera, the in-vehicle microphone, the in-vehicle central control screen and the in-vehicle speaker are all communicatively connected to the cockpit domain controller; the in-vehicle camera is used to capture images in the cockpit, the in-vehicle microphone is used to pick up voice information of users participating in a quick-answer session, the in-vehicle central control screen is used to display the overall application interface and interactive content, and the in-vehicle speaker is used to broadcast the question content and prompt information of the quick-answer interaction stage. The memory stores computer instructions, and the cockpit domain controller executes the computer instructions to execute the method of the above-mentioned first aspect or any corresponding embodiment thereof.

[0026] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method of the first aspect or any corresponding embodiment thereof.

[0027] In a fifth aspect, the present invention provides a computer program product comprising computer instructions for causing a computer to execute the method of the first aspect or any corresponding embodiment thereof.

[0028] The beneficial effects of the present invention are:

[0029] (1) According to the above technical means, the order of answering is determined by a pure visual detection mechanism. The present invention does not require the vehicle model to have a multi-sound zone configuration or additional answering sensing device to locate the sound source or determine the timing of answering, nor does it require the user to access additional equipment for interactive experience. It can directly use the existing camera and microphone in the car to realize multi-person answering interaction, reducing hardware costs and implementation difficulties. The present invention also has no restrictions on the identity of the participants in the car, does not require the recording of user information such as voiceprints or faces, and does not require users to learn complex and cumbersome gestures or action instructions. Only a simple hand-raising action can participate in the answering, allowing drivers and passengers of different ages to participate easily. The present invention determines the order and position of the answering through real-time visual detection, is not affected by noise, and combines the temporal correlation of lip movement information and audio information to avoid confusion errors in the voice system and improve the accuracy of audio positioning and answer content recognition. This solves the problems of strong hardware dependence, poor scene adaptability, and rough priority determination in existing in-vehicle answering solutions.

[0030] (2) According to the above technical means, the quick-answer action is determined based on the hand-raising action, and the dual constraints of the arm angle and the hand-raising height are used to judge whether the user has performed the specified hand-raising action, thereby improving the accuracy of action judgment.

[0031] (3) Based on the above technical means, by identifying specific key points on the arms and torso, it is calculated whether the arm movement meets the constraints of the arm angle and the hand-raising height. The scheme is simple in principle and accurate in calculation, providing a method for judging the quick-response action that takes into account both efficiency and accuracy.

[0032] (4) Based on the above technical means, the distance between the user and the camera is used to calculate the corresponding correction weight to correct the arm height measurement, which solves the problem of distortion in the two-dimensional image measurement of the vehicle interior space for the first time. This method can eliminate the visual differences caused by the different distances between the user and the camera, making the calculation of the hand-raising height more accurate.

[0033] (5) According to the above technical means, when the camera is at the front of the car, considering that the front row users are close to the camera, the same physical height occupies a larger proportion in the picture, while the back row users are far from the camera, the same physical height occupies a smaller proportion in the picture. By calculating the ratio of the front row distance to the total distance, the front row perspective magnification effect is eliminated, and by calculating the ratio of the back row distance to the total distance, the back row perspective reduction effect is compensated. The spatial perspective height correction model converts the two-dimensional visual measurement values ​​in the car into comparable height data in the real physical space for the first time by establishing a distance-weight mapping function, fundamentally solving the fairness problem of answering questions caused by seat position.

[0034] (6) Based on the above technical means, a three-level priority judgment rule for answering is provided. The first answering order includes the first and second levels, and the second answering order includes the third level. The first level is based on the principle of frame timing priority, identifying the frame number where the hand-raising action first appears. The smaller the frame number, the higher the priority, solving 90% of the conventional answering scenarios. The second level is based on the principle of height priority. When the frame is the same, the spatial perspective correction model is called to calculate the actual hand-raising height. The one with the larger height value wins. Furthermore, combined with the subsequent third level based on the principle of lip movement timing matching, when the frame is the same and the height is the same, the pronunciation start time point is locked through lip movement detection, and the lip movement-speech timing consistency is verified by audio segmentation. The first effective speaker gets the priority to answer. The triple progressive judgment covers the full dimension from millisecond timing (frame sequence) to behavioral characteristics (height) and then to biometric characteristics (lip movement), solving the problem of failure of traditional acoustic solutions when multiple people answer at the same time. 95% of the scenarios are directly judged by the first level (calculation time < 2ms), and 5% of the extreme scenarios trigger the third level analysis, meeting the real-time requirements of the vehicle (total delay ≤ 50ms). The three-level collaboration reduces the priority misjudgment rate to almost zero, significantly improving the accuracy of quick responses.

[0035] (7) Based on the above technical means, the system can also set up multiple designated answering areas in the car, such as setting up touch areas on the door armrests, seat backs or center consoles near each seat. The system detects the position of the user's palm through image recognition technology and determines whether the palm is in contact with the answering area. When the system detects that the user's palm touches the answering area, it is considered that the user has completed the answering action. This answering method is simpler and more direct than raising the hand to answer. The user only needs to touch the answering area to complete the answering without making an obvious hand-raising action. This method is more suitable for situations where the space in the car is limited. Under the condition of limited space in the car, the answering method of touching the answering area can also reduce the system's misjudgment of the user's non-standard actions and improve the accuracy of answer recognition.

[0036] (8) Based on the above technical means, when collecting images inside the cockpit, the system uses an RGB-IR camera to collect images inside the cockpit. The RGB-IR camera has both visible light (RGB) and infrared (IR) imaging capabilities, and can clearly capture images inside the cockpit under various lighting conditions (including low-light environments at night), ensuring that the system can accurately identify the user's quick response actions and lip movement information.

[0037] (9) Through visual detection, the area of ​​each position is divided and visualized on the central control screen so that the participants can confirm their own answering area. On the one hand, it avoids the misjudgment of limb cross-judgment during the answering process (the algorithm itself is based on individual matching of torso and limb detection to avoid cross-matching). Introducing position area division can further guide users and enhance the accuracy of answering action detection and distinction; on the other hand, through visualization, the user's gesture action is guided to appear in the visible area, avoiding the failure of action judgment due to occlusion, thereby improving participation. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0039] Figure 1 1 is a flow chart of a vehicle-mounted multi-person quick-answer method according to an embodiment of the present invention;

[0040] Figure 2 This is another flowchart of a vehicle-mounted multi-person quick-answer method according to an embodiment of the present invention;

[0041] Figure 3 This is another flowchart of a vehicle-mounted multi-person quick-answer method according to an embodiment of the present invention;

[0042] Figure 4 2 is a schematic diagram of a process for determining a quick-answer action according to an embodiment of the present invention;

[0043] Figure 5 This is a schematic diagram of collecting key points of a human body according to an embodiment of the present invention;

[0044] Figure 6 2 is a schematic structural diagram of a vehicle-mounted multi-person quick-answer device according to an embodiment of the present invention;

[0045] Figure 7 4 is a schematic diagram of the hardware structure of a vehicle according to an embodiment of the present invention. DETAILED DESCRIPTION

[0046] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0047] According to an embodiment of the present invention, an embodiment of an in-vehicle multi-person quick-answering method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0048] In this embodiment, a vehicle-mounted multi-person quick answering method is provided, which can be used in vehicles. Figure 1 This is a flow chart of a vehicle-mounted multi-person quick-answer method according to an embodiment of the present invention, the process comprising the following steps:

[0049] Step S101: Capture an image inside the cockpit, and determine the location information of users participating in the quick answer based on the image inside the cockpit.

[0050] Specifically, the embodiment of the present invention is equipped with a camera in the car. When the user plays the quick-answer game in the car, the camera in the car captures the image in the car in real time, and then analyzes the user's quick-answer action, providing a purely visual quick-answer judgment method. In the embodiment of the present invention, the number and position of the camera in the car are not specifically limited. The camera can be located at the front of the cabin, the rear of the cabin, and the door of the cabin, as long as the camera in the cabin can clearly capture the cabin image from all angles. In a specific embodiment, the embodiment of the present invention can be reused for the in-car monitoring high-definition camera arranged on the front ceiling, the front windshield area or the instrument panel. The camera itself is part of the cabin OMS (Occupancy Monitoring System) system and is used for conventional functions in the cabin such as in-car photography / video, occupant attribute detection, and object detection. There is no need to add additional camera equipment and sensing equipment, thereby reducing the production cost of the vehicle.

[0051] In one optional embodiment, the in-vehicle camera utilizes an RGB-IR camera. RGB-IR (Red, Green, Blue, Infrared) is an image sensor technology that combines visible light (RGB) and infrared (IR) imaging. Its core lies in a specialized filter array. The RGB-IR camera possesses both visible light (RGB) and infrared (IR) imaging capabilities, enabling clear capture of the cabin interior in a variety of lighting conditions, including low-light nighttime environments. This ensures that the system can accurately identify the user's quick-response actions and maintains image acquisition quality in dimly lit environments. This enables the in-vehicle multi-person quick-response interactive method of this embodiment to support a full range of interactive scenarios, both daytime and nighttime.

[0052] Afterwards, since the image inside the cabin contains the position distribution of the passengers in the car, the location information of the users participating in the quick answer can be determined based on the image inside the cabin. The system analyzes the image inside the cabin through image processing technology, identifies the users in various positions in the car, and assigns a unique location identifier to each user, such as "driver's seat", "front passenger seat", "back row left", "back row middle", "back row right", etc. This location information will be used for subsequent quick answer action recognition and quick answer order determination. The identity of each user does not need to be bound through facial recognition, voiceprint recognition, etc. Even strangers who get on the car for the first time can be assigned a unique label based on their seating position, so that the identity of the user participating in the quick answer is uniquely represented by location information. This not only ensures the accuracy of identity representation, but also reduces the complexity of the identity binding algorithm in the early stage of the quick answer game.

[0053] Step S102: output the question.

[0054] Step S103: identifying the quick-answer action in each position information.

[0055] Specifically, after a question is given, the cockpit domain controller uses the divided location areas to identify the quick-answer actions within each location information from the collected images inside the cockpit. In an embodiment of the present invention, the quick-answer action can be a hand-raising action, a head-shaking action, a clapping action, or an action of touching a designated area. This embodiment is only an example and is not limited to this. As long as the user's action can be identified as having the intention to answer, it is within the scope of protection of the embodiment of the present invention. For example, the vehicle prompts the user through voice, asking him to make a quick-answer action within 5 seconds.

[0056] Step S104: determining the first answering order of each participating user according to the order in which the answering actions appear in the picture frame.

[0057] Specifically, the system extracts the picture frames when each participating user makes a quick answering action, determines the time sequence of each picture frame, and sorts the picture frames in order from early to late, so that the picture frames with an earlier sequence can indicate that the user in the corresponding position has the right to answer first, thereby determining the first answering order of the users participating in the quick answering.

[0058] Step S105 , for users in different order in the first response sequence, prompt the users to answer questions in sequence and collect voice information of each user.

[0059] Step S106: prompt multiple users with the same priority in the first rush-to-answer sequence to answer questions at the same time, and collect lip movement information and voice information of each user during the process of users with the same priority answering questions at the same time.

[0060] Step S107: Match the lip movement information with the voice information collected when the users in the same order answer questions at the same time to obtain the answer data of the users in the same order.

[0061] Step S108, determining the starting time frame of each user's lip movement information based on the answer data, and determining the second answer order of users in the same order based on the starting time frame of each user.

[0062] Step S109, determining the quick answer result based on the first quick answer order, the second quick answer order and the voice information of each user.

[0063] Specifically, due to the limitations of the camera's shooting frequency, it is sometimes inevitable that the same ranking will occur. Although the probability of this is small, it will also affect the accuracy and fairness of the quick answer. Therefore, the first quick answer order still includes two situations. The first situation is that only one user makes a quick answer action in each frame. This situation reflects the scenario where no one grabs the question at the same time. The second situation is that in some frames, multiple users make a quick answer action. This situation reflects the scenario where multiple users grab the question at the same time. For example, two users raise their hands at the same time and are captured in the same frame.

[0064] For the first case, the embodiment of the present invention can prompt users to answer questions in sequence through the vehicle and collect voice information of each user. For example, assuming the first rush-to-answer order is: 1. Driver's seat, 2. Front passenger seat, 3. Left side of the rear row, 4. Middle of the rear row, 5. Right side of the rear row, the order of answering is: Driver's seat user answers, Front passenger seat user answers, Left side of the rear row user answers, Middle side of the rear row user answers, Right side of the rear row user answers. According to the first rush-to-answer order, the system sends answer prompts to users in different positions through the in-car speakers or central control screen, such as "Please ask the first responder to answer the question." When the user starts to answer the question, the system collects the user's voice information through the in-car microphone, and associates the voice information with the user's location information and stores it.

[0065] For the second scenario, assume the first order of response is: 1. Driver's seat, 2. Front passenger seat, 3. Left and middle rear seats, 4. Right rear seat. The user on the left rear seat and the user in the middle rear seat are both captured raising their hands simultaneously in the third order, thus including the user on the left rear seat and the user in the middle rear seat in the same order. The system prompts multiple users in the same order to answer the question simultaneously. For example, the system prompts the driver to answer the question first, then prompts the user in the front passenger seat, and then prompts through the in-car speaker or central control screen, "Please answer the question simultaneously for the user on the left rear seat and the user in the middle rear seat." As the users answer the questions simultaneously, the system collects lip movement information from each user through the in-car camera and mixed voice information through the in-car microphone.

[0066] When multiple users answer a question simultaneously, some speak at different speeds. Therefore, embodiments of the present invention propose further determining which user in the same order has priority in answering based on the starting timing frame of lip movement information. However, considering that lip movement does not necessarily represent a user's actual answer (for example, if a first user does not speak an answer during a simultaneous answering session, but simply opens their mouth without making a sound, while a second user speaks after a moment's deliberation, it is clear that the second user has actually answered successfully. However, determining the first user's success based solely on the speed of their lip movement is inaccurate), embodiments of the present invention first match lip movement information with voice information.

[0067] Specifically, if Figure 2 As shown, the system analyzes the lip movement characteristics of each user, such as lip shape changes, opening and closing frequency, etc., and compares them with the acoustic characteristics of the voice signal, thereby separating the mixed voice information and matching it to the corresponding user, and obtaining independent answer data for each user.

[0068] After obtaining the response data, the system analyzes each user's lip movement information, including its duration and starting point, to determine the actual starting time frame for each user's speech. By comparing the temporal order of the starting time frames of different users' lip movement information, the system determines the second response order for users with the same priority. For example, if user A's lip movement starts at frame 100 and user B's starts at frame 105, user A ranks before user B in the second response order.

[0069] At the same time, by comparing the voice information included in the answer data of users in the same order with the standard answer, it can be determined whether the corresponding user answer content is correct.

[0070] Finally, the answer result is determined based on the first answer order, the second answer order and the voice information of each user.

[0071] Specifically, in the embodiment of the present invention, the process of determining the answer result can also adopt two specific solutions:

[0072] The first option is to have each user speak out their answer, and then determine who is the fastest and most accurate winner based on the order in which each user answers and the correctness of their answers, and then proceed to the next question.

[0073] The second option is to have the winner who answers the question correctly first and then proceed to the next question. Figure 3As shown, based on the order of answers in the first rush-answer sequence, users are prompted to answer in sequence, and the voice information of the first responder is compared with the preset standard answer to determine whether the answer is correct or not. If it is correct, the next question will be advanced. If it is incorrect, the next responder will answer and the correctness of the answer content will be judged until someone answers correctly or all responders are incorrect and the next question is advanced. In the above process, if there are multiple users of the same order who successfully answer at the same time, when it is the turn of multiple participants who answer at the same time, the prompts are that most of the participants answer the question at the same time, and based on the temporal correlation between the lip movement information and the voice information, the speed and audio content of each participant's simultaneous answer are distinguished and matched, the second rush-answer sequence of each respondent is determined, and based on their respective voice recognition results, the correctness of their respective answer content is confirmed. Similarly, if the voice result of the user in the first order in the second rush-answer sequence is correct, the next question will be advanced. If it is incorrect, the correctness of the voice result of the user in the second rush-answer sequence will be judged, and if it is correct, the next question will be advanced. If all answers in the same order are incorrect, the next individual respondent will proceed until someone answers correctly or all responders are incorrect, moving on to the next question. Finally, based on the responses of the corresponding users, participation statistics are compiled, and interaction information is output on the in-vehicle terminal to show each user the results of the competition.

[0074] The above two schemes for determining the results of the quick answer can be flexibly selected and enabled according to the user's preferences. The embodiment of the present invention does not limit which scheme for determining the results of the quick answer must be used.

[0075] The technical solution provided by the embodiments of the present invention determines the order of responses based on a visual detection mechanism. Users who simultaneously respond simultaneously answer the question, and each user's lip movements are captured. The order of responses is further determined based on the order of their lip movements. For users who respond simultaneously, each user's true response is determined based on the matching of lip movement information and voice information, thus avoiding language confusion. This invention eliminates the need for the vehicle to have multiple sound zones or additional response sensors for sound source localization or response timing determination, nor does it require users to connect to additional devices for interactive experience. It can directly utilize existing cameras and microphones within the vehicle to enable multi-person response interaction, reducing hardware costs and implementation difficulty. This invention has no restrictions on the identity of participants within the vehicle, eliminates the need to record user information such as voiceprints or facial features, and eliminates the need for users to learn complex gestures or movement commands. Participation in the response can be achieved through a simple hand-raising gesture, making it easy for drivers of all ages to participate. By using real-time visual detection to determine the order and location of responses, the invention is unaffected by noise. By combining the temporal correlation of lip movement information with audio information, this invention avoids confusion errors in the voice system and improves the accuracy of audio localization and response content recognition. This solves the problems of existing in-vehicle quick-answer solutions, such as strong hardware dependence, poor scene adaptability, and rough priority determination.

[0076] In some optional implementations, step S103 includes:

[0077] Step a1: Identify the arm angle and hand-raising height of the user in each position information;

[0078] Step a2: When the arm angle falls within the preset angle range and the hand-raising height is greater than the preset height threshold, it is determined that a quick response action occurs at the corresponding position.

[0079] Specifically, for the recognition scheme where the quick-response action is a hand-raising gesture, the system identifies the user's arm angle and hand-raising height. When the arm angle falls within a preset angle range (e.g., 60 to 120 degrees) and the hand-raising height exceeds a preset height threshold (e.g., above 10 cm), the system determines that a quick-response action has occurred at the corresponding location. Using the dual constraints of arm angle and hand-raising height to determine whether the user has performed the specified hand-raising gesture improves the accuracy of action determination.

[0080] In some optional embodiments, step a1 includes:

[0081] Step b1, identifying the wrist joint and elbow joint of the user in the current position information;

[0082] Step b2, draw the arm line connecting the wrist joint point and the elbow joint point;

[0083] Step b3, calculating the angle of the arm in the current position information by using the angle between the arm connecting line and the horizontal line of the cockpit screen;

[0084] Step b4, identifying the user's arm key points and head and torso key points in the current position information, where the arm key points are predefined landmarks on the arms, and the head and torso key points are predefined landmarks on the head or torso;

[0085] Step b5, calculate the hand-raising height within the current position information based on the vertical distance between the arm key point and the head and torso key point.

[0086] Specifically, if Figure 4 As shown, the system first identifies the wrist and elbow joints of the user in the current position information. The system uses the human body key point detection algorithm to locate the coordinates of the user's wrist and elbow joints in the image inside the cockpit. Then, the system draws the arm line connecting the wrist and elbow joints to form a straight line indicating the direction of the arm. Next, the system calculates the angle of the arm in the current position information through the angle between the arm line and the horizontal line of the cockpit image. If the angle between the arm line and the horizontal line is 90 degrees, it means that the user's arm is raised vertically. For example Figure 5 As shown, point a is the user's wrist joint, point b is the user's elbow joint, and the angle between the line connecting point a and point b and the horizontal line of the cockpit image is the arm angle .

[0087] For the calculation method of hand-raising height, the system first extracts the key points of the arm and the key points of the head and torso. The key points of the arm refer to the predefined landmark points on the arm. There is no fixed restriction as long as the point is on the arm, for example Figure 5 The center point c of the line connecting the elbow and wrist is in the middle. The head-torso key point is a predefined landmark point on the head or torso. In some optional implementations, the head-torso key point can be the center of the eyebrows of the user's face or the horizontal center of the shoulder. Figure 5 The eyebrow center point d of the face is used. Then, the hand-raising height within the current position information is calculated based on the vertical distance between the arm key point and the head and torso key point, for example Figure 5 The h in.

[0088] Traditional gesture detection is prone to misjudgment during bumpy rides. However, the preset arm angle range can filter out over 85% of unintentional gestures (such as raising a hand to wipe sweat). Vertical distance detection allows for equal recognition of hand-raising gestures from passengers of varying body sizes (children and adults), with spatial adaptability. Furthermore, the two detection conditions can be calculated in parallel, enabling judgment within 10ms, ensuring algorithm efficiency.

[0089] In particular, the vertical distance between the arm keypoint and the head-torso keypoint is used to calculate the hand-raised height within the current position information. This seemingly simple yet ingenious method avoids the adaptability issues associated with fixed thresholds (for example, for passengers with varying heights, arm lengths, and weight levels, this solution adapts the hand-raised height to their body type, ensuring that all parameters are consistent and that the individual body types do not affect the hand-raised height evaluation criteria). Furthermore, it provides an interface for subsequent spatial perspective correction.

[0090] Based on the above technical means, by identifying specific key points on the arms and torso, it is calculated whether the arm movement meets the constraints of arm angle and hand-raising height. The solution principle is simple and the calculation is accurate, providing a method for judging quick-response actions that takes into account both efficiency and accuracy.

[0091] In some optional embodiments, step b5 includes:

[0092] Step c1, calculate the vertical distance between the arm key point and the head and torso key point;

[0093] Step c2, determining the relative distance between the current position information and the camera, and determining the correction weight according to the relative distance;

[0094] Step c3, determine the hand-raising height using the product of the correction weight and the vertical distance.

[0095] Specifically, the on-board monocular camera has an inherent defect due to the principle of perspective projection - the farther the object is from the lens, the smaller the image size. Assuming that the camera is in the front row of the car, it is typically shown that the actual hand-raising height of the user in the back row is 35cm, which only occupies 20 pixels in the picture, while the action of the same height in the front row occupies 60 pixels. If the pixel distance is compared directly, it may be mistakenly judged that the hand-raising height of the user in the back row is 67% lower than that of the front row, resulting in serious distortion in the priority determination of the quick response. Based on this, the embodiment of the present invention also calculates the relative distance between the position of each user who issues a quick response action and the camera, and then determines the corresponding correction weight based on the distance, thereby correcting the numerical value of the hand-raising height. This solution solves the unfair problem of quick response caused by the distortion of the two-dimensional picture measurement in the car space for the first time. This method can eliminate the visual differences caused by the different distances between the user and the camera, making the calculation of the hand-raising height more accurate.

[0096] In some optional implementations, when the camera is at the front of the vehicle, step c2 includes:

[0097] Step d1, obtaining the distance between the front row of the vehicle seat and the front row of the camera, and obtaining the distance between the rear row of the vehicle seat and the rear row of the camera;

[0098] Step d2, calculating the sum of the front row distance and the back row distance to obtain the total distance;

[0099] Step d3, determining whether the current position information is in the front row or the back row;

[0100] Step d4: If the current position information is in the front row, calculate the ratio of the front row distance to the total distance to obtain the correction weight;

[0101] Step d5: If the current position information is in the back row, the ratio of the back row distance to the total distance is calculated to obtain a correction weight.

[0102] Specifically, the embodiment of the present invention is based on the vertical distance measured by the camera h The relative height is calculated based on the distance between the front and rear rows and the camera, and correction weights are used to perform the correction. In this embodiment of the present invention, considering that the distance between each row of seats and the camera is almost the same, this embodiment uses the specific scenario of two rows of seats as an example and defines the correction weights to include front row correction weights and rear row correction weights.

[0103] Therefore, the front row hand height , back row hand raising height ,in, Indicates the front row correction weight when the user sits in the front row, Indicates the rear seat correction weight when the user sits in the rear seat.

[0104] The system then obtains the front-row distance from the front seat to the camera and the rear-row distance from the rear seat to the camera. For example, the front-row distance might be 0.8 meters, and the rear-row distance might be 1.5 meters. The system calculates the sum of the front-row and rear-row distances to obtain the total distance, which in this example is 2.3 meters.

[0105] In the embodiment of the present invention, the size of the object in the picture is completely dependent on the inverse relationship between distance and size. At the same physical height, an increase in distance will cause the picture size to decrease proportionally, resulting in a systematic deviation in the measurement.

[0106] Therefore, the modified weight calculation method provided by the embodiment of the present invention is as follows:

[0107] , , where Indicates the distance between the front passenger and the camera. Indicates the distance between the rear passengers and the camera. In the above example, the calculated front-seat correction weight is 0.8 / 2.3≈0.35, and the calculated rear-seat correction weight is 1.5 / 2.3≈0.65. This correction weight calculation method takes into account the difference in distance between users and the camera, allowing the system to fairly compare the hand-raising heights of users in different positions.

[0108] According to the above technical approach, when the camera is at the front of the vehicle, considering that front-seat users are closer to the camera, their physical height accounts for a larger portion of the frame, while back-seat users are farther from the camera, their physical height accounts for a smaller portion of the frame, the calculated correction weight needs to be inversely proportional to the distance between the user and the camera. Therefore, this embodiment of the present invention eliminates the front-row perspective magnification effect by calculating the ratio of the front-row distance to the total distance, and compensates for the back-row perspective reduction effect by calculating the ratio of the back-row distance to the total distance. By establishing a distance-weight mapping function, the spatial perspective height correction model, for the first time, converts in-vehicle two-dimensional visual measurements into comparable height data in real physical space, fundamentally resolving the fairness issue of quick-answering due to seating position.

[0109] In some optional implementations, step S104 includes:

[0110] Step e1, extracting the picture frames when each participating user performs the quick answer action;

[0111] Step e2, determining the time sequence of each picture frame, and sorting each picture frame in order from earliest to latest time to obtain a first sorting;

[0112] Step e3, determining whether there is a target frame, wherein the target frame includes the response actions of multiple users;

[0113] Step e4: sorting the hand-raising heights of different users' quick-answer actions in the target frame from high to low to obtain a second ranking, wherein if the hand-raising heights of multiple target users in the target frame are the same, the target users are defined as users with the same ranking;

[0114] Step e5: Generate a first response order based on the first sorting and the second sorting.

[0115] Specifically, based on the above technical means, a three-level rush-answer priority determination rule is provided. The first rush-answer order includes the first and second level determinations, and the second rush-answer order includes the third level determination. The first level determination is based on the principle of frame timing priority. The frame number in which each user's hand-raising action first appears is identified. The smaller the frame number, the higher the priority. This results in the first ranking, which can solve 90% of common rush-answer scenarios.

[0116] The second-level priority determination is based on the principle of height priority. When multiple users raise their hands in the same frame, the embodiment of the present invention uses the spatial perspective correction model based on the aforementioned steps to calculate the actual height of each user's hand raise. The hand raise heights of the users are then compared. The user with the higher hand raise height value wins and is given higher priority, resulting in the second ranking. The first and second rankings are combined to form a complete first-level response order.

[0117] Combined with the principle of lip movement timing matching in the subsequent third level, only when the user's hand-raising action appears in the same frame and at the same height, the pronunciation starting time point is locked through lip movement detection. Combined with audio segmentation to verify the lip movement-speech timing consistency, the person who pronounces the pronunciation effectively the earliest will get priority to answer and get the second answer order.

[0118] The triple-level progressive decision-making process provided by this embodiment of the present invention covers all dimensions, from millisecond-level timing (frame sequence) to behavioral characteristics (height) and finally to biometric characteristics (lip movement), addressing the inefficiency of traditional acoustic solutions when multiple people are simultaneously responding. 95% of scenarios are directly determined by the first level (computational time <2ms), while 5% of extreme scenarios trigger the third level of analysis, meeting the real-time requirements of in-vehicle performance (total latency ≤50ms). This three-level collaboration reduces the priority misjudgment rate to almost zero, significantly improving the accuracy of the response.

[0119] In some optional implementations, the above step S103 further includes:

[0120] Step f1, identifying whether the user's palm touches a designated area in each position information;

[0121] Step f2: When the user's palm touches the designated area, a quick answer action is determined to occur at the corresponding position.

[0122] Specifically, in an embodiment of the present invention, the system can also set up multiple designated answering areas in the car, such as setting touch areas on the door armrests near each seat, seat backs, ceilings, etc., or on the center console. The system detects the position of the user's palm through image recognition technology and determines whether the palm is in contact with the designated area. When the system detects that the user's palm touches the designated area, it is considered that the user has completed the answering action. This way of answering is simpler and more direct than answering by raising the hand. The user only needs to touch the designated area to complete the answer without making an obvious hand-raising action. This method is more suitable for situations where space is limited in the car. Under the condition of limited space in the car, the answering method of touching the designated area can also reduce the system's misjudgment of the user's non-standard actions and improve the accuracy of answer recognition.

[0123] This embodiment also provides a vehicle-mounted multi-person quick-answer device, which is used to implement the above-mentioned embodiments and preferred embodiments. Details that have already been described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0124] In some optional embodiments, the in-vehicle multi-person quick-answer method provided by the present invention further includes: dividing the location information into regions, and visualizing the divided regions on the vehicle's central control screen so that the participants can confirm their own quick-answer regions.

[0125] According to the above-mentioned technical means, during the process of answering questions through quick-answering actions, the areas of each position are divided through visual detection and visualized on the central control screen for participants to confirm their own answering areas. On the one hand, it avoids the misjudgment of limb crossing during the answering process (the algorithm itself is based on individual matching of torso and limb detection to avoid cross-matching). The introduction of position area division can further guide users and enhance the accuracy of quick-answering action detection and distinction; on the other hand, through visualization, the user's gestures are guided to appear within the visible area, avoiding the failure of action judgment due to occlusion, thereby improving participation.

[0126] This embodiment provides a vehicle-mounted multi-person quick answering device, such as Figure 6 Shown, including:

[0127] The location binding module 601 is used to collect images inside the cockpit and determine the location information of the users participating in the quick answer based on the images inside the cockpit;

[0128] The question generating module 602 outputs the question;

[0129] Action recognition module 603, used to identify the quick-answer action in each position information;

[0130] The first rush-answer order determination module 604 is used to determine the first rush-answer order of each participating user according to the order in which the rush-answer actions appear in the picture frame;

[0131] The first answering module 605 is used to prompt users in different order in the first answering sequence to answer questions in sequence and collect voice information of each user;

[0132] The second answering module 606 is configured to prompt multiple users who are in the same order in the first answering sequence to answer the question simultaneously, and collect lip movement information and voice information of each user during the process of the users in the same order answering the question simultaneously;

[0133] Matching module 607, used to match the lip movement information with the voice information collected during the simultaneous answering process of the user in the same order, to obtain the answer data of the user in the same order;

[0134] The second rush-to-answer order determination module 608 is used to determine the starting time frame of each user's lip movement information based on the answer data, and determine the second rush-to-answer order of users with the same priority based on the starting time frame of each user;

[0135] The result module 609 is used to determine the result of the quick answer based on the first quick answer order, the second quick answer order and the voice information of each user.

[0136] In some optional implementations, the action recognition module 603 includes:

[0137] A user identification unit is used to identify the arm angle and hand-raising height of the user in each position information;

[0138] The action determination unit is used to determine that a quick-answer action occurs at the corresponding position when the arm angle falls within a preset angle range and the hand-raising height is greater than a preset height threshold.

[0139] In some optional implementations, the user identification unit includes:

[0140] The arm joint point recognition subunit is used to identify the wrist and elbow joint points of the user in the current position information;

[0141] The arm direction drawing subunit is used to draw the arm connection line between the wrist joint point and the elbow joint point;

[0142] An angle calculation subunit, used to calculate the angle of the arm in the current position information through the angle between the arm connection line and the horizontal line of the cockpit screen;

[0143] The body key point recognition subunit is used to identify the user's arm key points and head and torso key points in the current position information. The arm key points are predefined landmarks on the arms, and the head and torso key points are predefined landmarks on the head or torso.

[0144] The height calculation subunit is used to calculate the hand-raising height within the current position information based on the vertical distance between the arm key point and the head and torso key point.

[0145] In some optional embodiments, the height calculation subunit includes:

[0146] A vertical distance calculation subunit is used to calculate the vertical distance between the arm key points and the head and torso key points;

[0147] A correction weight calculation subunit is used to determine the relative distance between the current position information and the camera, and determine the correction weight according to the relative distance;

[0148] The correction subunit is used to determine the hand-raising height by using the product of the correction weight and the vertical distance.

[0149] In some optional implementations, the first answering order determination module 604 includes:

[0150] A picture extraction unit is used to extract the picture frames when each participating user makes a quick answer action;

[0151] A first sorting unit is used to determine the time sequence of each picture frame and sort the picture frames in order from earliest to latest time to obtain a first sorting sequence;

[0152] A target frame determination unit, configured to determine whether a target frame exists, wherein the target frame includes the response actions of multiple users;

[0153] a second sorting unit, configured to sort the hand-raising heights of different users' quick-answer actions in the target frame from high to low to obtain a second sorting, wherein if the hand-raising heights of multiple target users in the target frame are the same, the target users are defined as users of the same ranking;

[0154] The first rush-to-answer sequence determining unit is used to generate a first rush-to-answer sequence according to the first sorting and the second sorting.

[0155] In some optional implementations, the action recognition module 603 further includes:

[0156] An area determination unit, configured to identify whether the user's palm touches a designated area within each position information;

[0157] The action determination unit is used to determine whether a quick-answer action occurs at the corresponding position when the user's palm touches the designated area.

[0158] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0159] The embodiment of the present invention further provides a vehicle, such as Figure 7 As shown, the system includes: memory, a cockpit domain controller, an in-car camera, an in-car microphone, an in-car central control screen, and an in-car speaker. These memory, in-car camera, in-car microphone, in-car central control screen, and in-car speakers are all connected to the cockpit domain controller. The in-car camera is used to capture images from the cockpit, the in-car microphone is used to pick up voice information from users participating in the quick-answer session, the in-car central control screen is used to display the overall application interface and interactive content, and the in-car speaker is used to announce the content of the questions and prompts during the quick-answer interaction phase. The in-car camera, in-car microphone, in-car central control screen, and in-car speakers belong to the hardware layer.

[0160] The memory stores the device model provided by the aforementioned device embodiment, corresponding to the computer instructions provided by the method embodiment. The cockpit domain controller executes the computer instructions, thereby executing the method provided by the aforementioned method embodiment. This randomly selects preset quick-answer questions based on the type of quick-answer game and displays them. The order of responses is determined based on real-time detection of participants' quick-answer behavior. After multiple participants successfully answer simultaneously, their speed and content matching are distinguished. The correctness of the answers is determined based on voice content analysis and comparison. Statistics are then displayed for each stage of the quick-answer interaction.

[0161] During the multi-person quick-answer interaction stage, the hardware layer in the vehicle obtains visual image information, which is processed by the cockpit domain controller and calls the software module in the memory to make interactive logic judgments. The hardware layer then performs visual display and language prompts. The overall interactive system is mainly based on visual detection in terms of software and hardware to achieve interactions such as quick-answer priority, position matching, and answer speed differentiation. Speech recognition technology only determines the accuracy of the answer.

[0162] The cockpit domain controller can be a central processing unit (CPU), a network processor (NPU), or a combination thereof. The cockpit domain controller can further include a hardware chip. The hardware chip can be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device can be a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a general purpose array logic (GAL), or any combination thereof.

[0163] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0164] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0165] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A vehicle-mounted multi-person quick answering method, characterized in that: The method comprises: Collecting images of the interior of the cockpit and determining the location information of users participating in the quick answer based on the images of the interior of the cockpit; Output the title; Identify the quick-answer actions within each location information; Determine the first answer order of each participating user according to the order in which the answer actions appear in the picture frame; For users in different order in the first quick-answer sequence, prompt the users to answer questions in sequence and collect voice information from each user; For multiple users in the same order in the first quick-answer sequence, prompt the users to answer the question simultaneously, and collect lip movement information and voice information of each user during the process of the users in the same order answering the question simultaneously; Matching the lip movement information with voice information collected during the simultaneous answering of questions by users in the same order to obtain answer data of users in the same order; matching the lip movement information with voice information collected during the simultaneous answering of questions by users in the same order to obtain answer data of users in the same order, including: analyzing the lip movement characteristics of each user based on the lip movement information of each user; and comparing the lip movement characteristics of each user with the acoustic characteristics of the corresponding voice signal, separating and matching the mixed voice information to the corresponding user, and obtaining independent answer data for each user; Determining the starting time frame of each user's lip movement information based on the answer data, and judging the second rush-to-answer order of users in the same order based on the starting time frame of each user; determining the starting time frame of each user's lip movement information based on the answer data, and judging the second rush-to-answer order of users in the same order based on the starting time frame of each user, including: analyzing the lip movement duration and the starting point of each user's lip movement based on the independent answer data of each user, and determining the starting time frame at which each user actually starts speaking based on the lip movement duration and the starting point of the lip movement; judging the second rush-to-answer order among users in the same order by comparing the time sequence of the starting time frames of the lip movement information of different users; The quick answer result is determined based on the first quick answer order, the second quick answer order and the voice information of each user.

2. The method according to claim 1, characterized in that The identifying of the quick-answer action in each position information includes: Identify the user's arm angle and hand-raising height in each location information; When the arm angle falls within the preset angle range and the hand-raising height is greater than a preset height threshold, it is determined that the quick-answer action occurs at the corresponding position.

3. The method according to claim 2, characterized in that The identifying of the arm angle and hand-raising height of the user in each position information includes: Identify the user's wrist and elbow joints within the current location information; Draw an arm line connecting the wrist joint point and the elbow joint point; The arm angle in the current position information is calculated by calculating the angle between the arm connection line and the horizontal line of the cockpit screen; Identify arm key points and head and torso key points of the user in the current position information, wherein the arm key points are predefined landmark points on the arm, and the head and torso key points are predefined landmark points on the head or torso; The hand-raising height in the current position information is calculated based on the vertical distance between the arm key point and the head and torso key point.

4. The method according to claim 3, characterized in that The calculating the hand-raising height in the current position information according to the vertical distance between the arm key point and the head and torso key point includes: Calculating the vertical distance between the arm key point and the head and torso key point; Determining the relative distance between the current position information and the camera, and determining a correction weight based on the relative distance; The hand-raising height is determined by multiplying the correction weight by the vertical distance.

5. The method according to claim 4, characterized in that When the camera is at the front of the vehicle, determining the relative distance between the current position information and the camera, and determining the correction weight according to the relative distance, includes: Obtaining the distance from the front row of the vehicle seat to the front row of the camera, and obtaining the distance from the back row of the vehicle seat to the back row of the camera; Calculating the sum of the front row distance and the back row distance to obtain a total distance; Determining whether the current position information is in the front row or the back row; If the current position information is in the front row, then calculating the ratio of the front row distance to the total distance to obtain the correction weight; If the current position information is in the back row, the ratio of the back row distance to the total distance is calculated to obtain the correction weight.

6. The method according to claim 2, characterized in that The determining of the first answering order of each participating user according to the order in which the answering actions appear in the picture frame includes: Extract the picture frames when each participating user makes a quick answer action; Determine the time sequence of each picture frame, and sort the picture frames in order from earliest to latest time to obtain a first sorting; Determining whether there is a target picture frame, wherein the target picture frame includes the quick-answer actions of multiple users; sorting the hand-raising heights of different users' quick-answer actions in the target screen frame from high to low to obtain a second ranking, wherein if the hand-raising heights of multiple target users in the target screen frame are also the same, the target users are defined as users with the same ranking; The first quick-answer order is generated according to the first sorting and the second sorting.

7. The method according to claim 1, characterized in that The identifying of the quick-answer action in each position information further includes: Identify whether the user's palm touches the designated area in each location information; When the user's palm touches the designated area, it is determined that the quick answer action occurs at the corresponding position.

8. The method according to claim 1, characterized in that The collecting of images in the cockpit includes: The images inside the cabin are collected through an RGB-IR camera.

9. The method according to claim 1, characterized in that The method further comprises: The location information is divided into regions, and the divided regions are visualized on the vehicle's central control screen so that participants can confirm their own answering areas.

10. A vehicle-mounted multi-person quick-answer device, characterized in that: The device comprises: A location binding module is used to collect images of the cabin and determine the location information of users participating in the quick answer based on the images of the cabin; Question module, outputs questions; Action recognition module, used to identify the quick-response actions in each location information; The first rush-answer order determination module is used to determine the first rush-answer order of each participating user according to the order in which the rush-answer actions appear in the picture frame; A first answering module is configured to prompt users at different positions in the first answering sequence to answer questions in sequence and collect voice information from each user; a second answering module, configured to prompt multiple users who are in the same order in the first answering sequence to answer questions simultaneously, and to collect lip movement information and voice information of each user during the process of the users in the same order answering questions simultaneously; a matching module for matching the lip movement information with voice information collected during the simultaneous answering of questions by users in the same order to obtain answer data of the users in the same order; matching the lip movement information with voice information collected during the simultaneous answering of questions by users in the same order to obtain answer data of the users in the same order, including: analyzing the lip movement characteristics of each user; and comparing the lip movement characteristics of each user with the acoustic characteristics of the corresponding voice signal, separating the mixed voice information and matching it to the corresponding user to obtain independent answer data for each user; A second rush-to-answer order determination module is configured to determine the starting timing frame of each user's lip movement information based on the answer data, and to determine the second rush-to-answer order of users in the same order based on the starting timing frame of each user; determine the starting timing frame of each user's lip movement information based on the answer data, and to determine the second rush-to-answer order of users in the same order based on the starting timing frame of each user, including: analyzing the lip movement duration and the starting point of each user's lip movement based on the independent answer data of each user, and determining the starting timing frame at which each user actually starts speaking based on the lip movement duration and the starting point of the lip movement; and determining the second rush-to-answer order between users in the same order by comparing the time sequence of the starting timing frames of the lip movement information of different users; A result module is used to determine a quick answer result based on the first quick answer order, the second quick answer order and the voice information of each user.

11. A vehicle, characterized in that: include: Memory, cockpit domain controller, in-car camera, in-car microphone, in-car central control screen and in-car speakers; The memory, the in-car camera, the in-car microphone, the in-car central control screen and the in-car speaker are all communicatively connected to the cockpit domain controller; the in-car camera is used to capture images in the cockpit, the in-car microphone is used to pick up voice information of users participating in the quick-answer session, the in-car central control screen is used to display the overall application interface and interactive content, and the in-car speaker is used to broadcast the question content and prompt information of the quick-answer interaction stage. Computer instructions are stored in the memory, and the cockpit domain controller executes the method described in any one of claims 1 to 9 by executing the computer instructions.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 9.

13. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Voice responding method, device and system

    CN110838211A

  • Vehicle-mounted interactive voice game method and system and computer readable storage medium

    CN118059466A