Vehicle-mounted multi-person preemptive answering method and device, vehicle, medium and product

By using the camera and microphone to combine visual detection and lip-moving voice information in the on-board environment, the hardware dependence and rough judgment problems of the on-board multi-person quick-response solution are solved, and low-cost and accurate multi-person quick-response interaction is achieved, which is suitable for users of different ages.

CN120388359AActive Publication Date: 2025-07-29CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510886789.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-07-29
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The existing multi-person quick-answer solution in the vehicle has problems such as strong hardware dependence, poor scenario adaptability, and rough priority determination, making it difficult to achieve accurate multi-person quick-answer interaction in complex vehicle environments.

Method used

By collecting image recognition of the user's position information in the cockpit, using existing cameras and microphones for visual detection, combining the timing correlation of lip movement and voice information, the order of answers is determined, including the dual constraints of arm angle and hand height, and the three-level answer priority judgment rules are adopted to reduce hardware costs and improve accuracy.

Benefits of technology

It realizes low-cost and accurate multi-person quick answer interaction in complex vehicle environments, reduces hardware dependence, improves the accuracy of judging the order and scenario adaptability of the quick answer order, and is suitable for driving and riding users of different ages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388359A_ABST
    Figure CN120388359A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vehicle-mounted man-machine interaction, and discloses a vehicle-mounted multi-person preemptive answering method and device, a vehicle, a medium and a product, and the method comprises the steps: determining the position information of a user participating in preemptive answering according to an image in a cabin; outputting the questions; identifying preemptive answering actions in each piece of position information; judging a first preemptive answering sequence according to the appearing sequence of the preemptive answering actions in the picture frame; prompting users in different sequences in the first preemptive answering sequence to answer the questions in sequence; prompting the users of the same syn-position to answer the questions at the same time, and collecting lip movement information and voice information in the process of answering the questions at the same time; matching the lip movement information with the voice information; determining an initial time sequence frame of lip movement information of each user, and judging a second preemptive answering sequence of the same syn-position users according to the initial time sequence frame of each user; and determining a preemptive answer result based on the first preemptive answer sequence, the second preemptive answer sequence and the voice information. According to the invention, the hardware dependence is reduced, and the priority judgment accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of in-vehicle human-computer interaction, and particularly to an in-vehicle multi-person quick-answer method, device, vehicle, medium and product. Background Art

[0002] With the development of intelligent cockpit technology, in-vehicle multi-person interactive games (such as knowledge quick-answer, song identification by listening, etc.) have gradually become an important scenario for enhancing the driving and riding experience. In related technologies, realizing multi-person quick-answer interaction in the vehicle mainly relies on acoustic positioning solutions, biometric recognition solutions, visual recognition solutions, and hardware recognition solutions.

[0003] Among them, the acoustic positioning solution realizes sound source positioning through a microphone array and judges the quick-answer order according to the order of the sound sources. However, the in-vehicle noise environment is likely to cause sound source aliasing, resulting in problems such as poor positioning accuracy and a high misjudgment rate of the quick-answer order. The biometric recognition solution generally uses voiceprint recognition technology to bind the user identity and then combines the voice timestamp to determine the quick-answer order. Its limitation is that it requires users to pre-register voiceprint information in advance, and the process of binding the identity is complex, which is not friendly to the participation of temporary passengers and has poor scene adaptability. The visual recognition solution currently first disassembles the camera frame, and then captures the frame with the user's quick-answer action, and determines the user's quick-answer order according to the order of the frames. However, this solution also makes it difficult to avoid multiple users acting simultaneously in the same frame, so it is also relatively easy to have problems where the quick-answer order cannot be distinguished. The hardware recognition solution generally installs multiple sensors as quick-answer buttons on the vehicle, and users need to touch the sensor next to their seats when answering quickly. Although this solution is relatively easy to determine the quick-answer order, the additional installed sensors will increase the vehicle cost, and the installed sensors are also difficult to have other functions, resulting in a waste of costs.

[0004] In summary, the related technologies generally have three major constraints: strong hardware dependence, poor scene adaptability, and rough priority determination. Therefore, there is an urgent need for a quick-answer interaction solution with low hardware cost, high robustness, and suitable for complex in-vehicle environments. Summary of the Invention

[0005] In view of this, the present invention provides an in-vehicle multi-person quick-answer method, device, vehicle, medium and product to solve the problems of strong hardware dependence, poor scene adaptability, and rough priority determination.

[0006] In a first aspect, the present invention provides a method for multiple-person quick answering in a vehicle. The method includes: collecting an image inside the cockpit and determining the position information of the users participating in the quick answering based on the image inside the cockpit; outputting a question; identifying the quick answering actions within each position information; determining the first quick answering order of each participating user according to the order in which the quick answering actions appear in the frame; for users with different ranks in the first quick answering order, prompting the users to answer the questions in sequence and collecting the voice information of each user; for multiple users with the same rank in the first quick answering order, prompting the users to answer the questions simultaneously, and collecting the lip movement information and voice information of each user during the simultaneous answering of the users with the same rank; matching the lip movement information with the voice information collected during the simultaneous answering of the users with the same rank to obtain the answering data of the users with the same rank; determining the starting time sequence frame of the lip movement information of each user according to the answering data, and judging the second quick answering order of the users with the same rank according to the starting time sequence frame of each user; and determining the quick answering result based on the first quick answering order, the second quick answering order, and the voice information of each user.

[0007] According to the above technical means, the present invention does not require the vehicle model to have a multi-zone configuration or an additional quick answering sensing device for sound source localization or quick answering time sequence determination, nor does it require users to additionally connect devices for an interactive experience. It can directly utilize the existing cameras and microphones in the vehicle to achieve multiple-person quick answering interaction, reducing the hardware cost and implementation difficulty. The present invention has no restrictions on the identities of the participants in the vehicle, does not require user information such as voiceprints or faces to be entered, and does not require users to learn complex and cumbersome postures or action instructions. Users can participate in the quick answering simply by raising their hands, enabling driving users of different ages to easily participate. The present invention determines the quick answering order and position through real-time visual detection, is not affected by noise interference, and combines the time sequence correlation between lip movement information and audio information to avoid confusion errors in the voice system, improving the accuracy of audio localization and answer content recognition. Thus, it solves the problems of strong hardware dependence, poor scene adaptability, and rough priority determination existing in the existing vehicle quick answering solutions.

[0008] In some optional embodiments, identifying the quick answering actions within each position information includes: identifying the arm angle and the raising height of the hand of the user within each position information; when the arm angle falls within a preset angle range and the raising height is greater than a preset height threshold, determining that the quick answering action appears at the corresponding position.

[0009] According to the above technical means, the quick answering action is determined based on raising the hand, and a double constraint of the arm angle and the raising height is used to judge whether the user makes the specified raising hand action, improving the accuracy of action determination.

[0010] In some optional embodiments, the method further includes: dividing the position information into regions and visualizing the divided regions on the vehicle center control screen so that the answering participants can confirm their own quick answering regions.

[0011] According to the above technical means, during the rush-answer process through rush-answer actions, regional division of each position is performed through visual detection and visualized on the central control screen for the participants to confirm their own rush-answer areas. On the one hand, it avoids misjudgment caused by limb crossing during the rush-answer process (the algorithm itself is based on individual matching of the torso and limbs to avoid cross-matching). Introducing the regional division of positions can further guide users and enhance the accuracy of differentiating rush-answer actions. On the other hand, through visualization, the gesture actions of users are guided to appear within the visible area, avoiding the failure of action determination caused by occlusion and improving the participation rate.

[0012] In some alternative embodiments, identifying the arm angle and raising height of the user within each position information includes: identifying the wrist joint point and elbow joint point of the user within the current position information; drawing the arm connection line between the wrist joint point and the elbow joint point; calculating the arm angle within the current position information through the angle between the arm connection line and the horizontal line of the cockpit screen; identifying the arm key point and head-torso key point of the user within the current position information, where the arm key point is a predefined landmark on the arm and the head-torso key point is a predefined landmark on the head or torso; calculating the raising height within the current position information according to the vertical distance between the arm key point and the head-torso key point.

[0013] According to the above technical means, by identifying specific key points on the arm and torso to calculate whether the arm movement meets the constraint conditions of the arm angle and raising height, the scheme principle is simple and the calculation is accurate, providing a rush-answer action determination method that takes into account both efficiency and accuracy.

[0014] In some alternative embodiments, calculating the raising height within the current position information according to the vertical distance between the arm key point and the head-torso key point includes: calculating the vertical distance between the arm key point and the head-torso key point; determining the relative distance between the current position information and the camera and determining the correction weight according to the relative distance; using the product of the correction weight and the vertical distance to determine the raising height.

[0015] According to the above technical means, the corresponding correction weight is calculated through the distance between the user and the camera to correct the measured height of the arm, which solves the problem of measurement distortion of the two-dimensional image in the vehicle interior space for the first time. This method can eliminate the visual difference caused by the different distances between the user and the camera, making the calculation of the raising height more accurate.

[0016] In some alternative embodiments, when the camera is at the front of the vehicle, determining the relative distance between the current position information and the camera and determining the correction weight according to the relative distance includes: obtaining the front-row distance from the front row of the vehicle seat to the camera and obtaining the rear-row distance from the rear row of the vehicle seat to the camera; calculating the sum of the front-row distance and the rear-row distance to obtain the total distance; determining whether the current position information is in the front row or the rear row; if the current position information is in the front row, calculating the ratio of the front-row distance to the total distance to obtain the correction weight; if the current position information is in the rear row, calculating the ratio of the rear-row distance to the total distance to obtain the correction weight.

[0017] According to the above technical means, when the camera is at the front of the vehicle, considering that the front-row users are closer to the camera and the same physical height occupies a larger proportion in the picture, while the rear-row users are farther from the camera and the same physical height occupies a smaller proportion in the picture. The front perspective magnification effect is eliminated by calculating the ratio of the front-row distance to the total distance, and the rear perspective reduction effect is compensated by calculating the ratio of the rear-row distance to the total distance. The spatial perspective height correction model fundamentally solves the problem of the fairness of the rush answer due to the seat position by establishing a distance-weight mapping function and converting the in-vehicle two-dimensional visual measurement value into the height-comparable data of the real physical space for the first time.

[0018] In some alternative embodiments, determining the first rush-answer order of each participating rush-answer user according to the order in which the rush-answer actions appear in the picture frames includes: extracting the picture frames when each participating rush-answer user makes a rush-answer action; determining the time order of each picture frame and sorting each picture frame in the order from the earliest to the latest to obtain the first sorting; determining whether there is a target picture frame that includes the rush-answer actions of multiple users; sorting the raising heights of the rush-answer actions of different users in the target picture frame from high to low to obtain the second sorting, where if the raising heights of multiple target users in the target picture frame are also the same, the target users are defined as users with the same rank; generating the first rush-answer order according to the first sorting and the second sorting.

[0019] According to the above technical means, a three-level priority determination rule for a quick answer is provided. The first quick answer order includes a first level and a second level, and the second quick answer order includes a third level. The first level is based on the principle of frame timing priority, identifying the frame number of the picture frame where the raising hand action first appears. The smaller the frame number, the higher the priority, which solves 90% of the conventional quick answer scenarios. The second level is based on the principle of height priority. When the frames are the same, the spatial perspective correction model is called to calculate the real raising hand height, and the one with a larger height value wins. Furthermore, combined with the subsequent third level based on the principle of lip movement timing matching, when the frames and heights are the same, the starting time point of pronunciation is locked through lip movement detection, and the consistency of lip movement - speech timing is verified by audio segmentation. The earliest valid speaker obtains the priority to answer quickly. The triple progressive determination covers all dimensions from millisecond-level timing (frame order) to behavioral characteristics (height) and then to biometric characteristics (lip movement), solving the problem of the failure of traditional acoustic solutions when multiple people answer quickly at the same time. 95% of the scenarios are directly determined by the first level (computing time < 2ms), and 5% of the extreme scenarios trigger the third-level analysis, meeting the in-vehicle real-time requirement (total delay ≤ 50ms). The three-level collaboration reduces the priority misjudgment rate to almost zero, significantly improving the accuracy of quick answers.

[0020] In some alternative embodiments, identifying the quick answer actions within each position information includes: identifying whether the palm of the user within each position information touches a specified area; when the palm of the user touches the specified area, determining that the quick answer action appears at the corresponding position.

[0021] According to the above technical means, the system can also set multiple specified quick answer areas inside the vehicle. For example, touch areas are set on the door armrests near each seat, on the seat backs, or on the center console. The system detects the palm position of the user through image recognition technology and determines whether the palm contacts the specified area. When the system detects that the palm of the user touches the specified area, it is considered that the user has completed the quick answer action. This quick answer method is simpler and more direct than raising the hand to answer. The user only needs to gently touch the specified area to complete the quick answer without making an obvious raising hand action. This method is more suitable for the limited space inside the vehicle. Under the condition of limited space inside the vehicle, the quick answer method of touching the specified area can also reduce the misjudgment of the system for the user's non-standard actions and improve the accuracy of quick answer recognition.

[0022] In some alternative embodiments, collecting the images inside the cockpit includes: collecting the images inside the cockpit through a RGB-IR type camera.

[0023] According to the above technical means, when collecting the images inside the cockpit, the system collects the images inside the cockpit through a RGB-IR type camera. The RGB-IR camera has both visible light (RGB) and infrared (IR) imaging capabilities, and can clearly capture the images inside the cockpit under various lighting conditions (including low-light environments at night), ensuring that the system can accurately identify the quick answer actions and lip movement information of the user.

[0024] In a second aspect, the present invention provides a vehicle-mounted multi-person quick-answer device, which includes: a position binding module for collecting an image inside the cockpit and determining the position information of the users participating in the quick answer according to the image inside the cockpit; a question-setting module for outputting questions; an action recognition module for recognizing the quick-answer actions within each position information; a first quick-answer order determination module for determining the first quick-answer order of each user participating in the quick answer according to the order in which the quick-answer actions appear in the frame of the picture; a first answering module for prompting the users in different positions in the first quick-answer order to answer questions in turn and collecting the voice information of each user; a second answering module for prompting the multiple users in the same position in the first quick-answer order to answer questions simultaneously, and collecting the lip movement information and voice information of each user during the simultaneous answering of the users in the same position; a matching module for matching the lip movement information with the voice information collected during the simultaneous answering of the users in the same position to obtain the answer data of the users in the same position; a second quick-answer order determination module for determining the starting time sequence frame of the lip movement information of each user according to the answer data, and judging the second quick-answer order of the users in the same position according to the starting time sequence frame of each user; and a result module for determining the quick-answer result based on the first quick-answer order, the second quick-answer order, and the voice information of each user.

[0025] In a third aspect, the present invention provides a vehicle, which includes: a memory, a cockpit domain controller, an in-vehicle camera, an in-vehicle microphone, an in-vehicle central control screen, and an in-vehicle speaker; the memory, the in-vehicle camera, the in-vehicle microphone, the in-vehicle central control screen, and the in-vehicle speaker are all communicatively connected to the cockpit domain controller; the in-vehicle camera is used for collecting an image inside the cockpit, the in-vehicle microphone is used for picking up the voice information of the users participating in the quick answer, the in-vehicle central control screen is used for displaying the overall application interface and interaction content, the in-vehicle speaker is used for broadcasting the question content and prompt information during the quick-answer interaction stage, the memory stores computer instructions, and the cockpit domain controller executes the above-mentioned method of the first aspect or any corresponding implementation manner by executing the computer instructions.

[0026] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the method of the first aspect or any corresponding implementation manner thereof.

[0027] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the method of the first aspect or any corresponding implementation manner thereof.

[0028] The beneficial effects of the present invention are as follows: (1) According to the above technical means, the order of answering the question by rush is judged through a pure vision detection mechanism. The present invention does not require the vehicle model to have a multi-zone configuration or an additional rush answering induction device for sound source localization or rush answering timing determination, nor does it require the user to additionally access a device for an interactive experience. It can directly utilize the existing cameras and microphones in the vehicle to achieve multi-person rush answering interaction, reducing the hardware cost and implementation difficulty. The present invention also has no restrictions on the identities of the participants in the vehicle, does not require the entry of user information such as voiceprints or faces, and does not require the user to learn complex and cumbersome postures or action instructions. The user can participate in the rush answering only through a simple raising hand action, enabling driving users of different age groups to easily participate. The present invention determines the rush answering order and position through real-time vision detection, is not affected by noise interference, and combines the timing correlation of lip movement information and audio information to avoid confusion errors in the voice system and improve the accuracy of audio localization and answer content recognition. Thus, the problems of strong hardware dependence, poor scene adaptability, and rough priority determination existing in the existing in-vehicle rush answering solutions are solved.

[0029] (2) According to the above technical means, the rush answering action is determined based on the raising hand, and whether the user makes the specified raising hand action is judged through the dual constraints of the arm angle and the raising hand height, improving the accuracy of action determination.

[0030] (3) According to the above technical means, by identifying specific key points on the arm and torso to calculate whether the arm movement conforms to the constraint conditions of the arm angle and the raising hand height, the scheme principle is simple and the calculation is accurate, providing a rush answering action determination method that takes into account both efficiency and accuracy.

[0031] (4) According to the above technical means, the corresponding correction weight is calculated based on the distance between the user and the camera to correct the measured height of the arm, solving for the first time the problem of measurement distortion of the two-dimensional image in the vehicle interior space. This method can eliminate the visual differences caused by different distances between the user and the camera, making the calculation of the raising hand height more accurate.

[0032] (5) According to the above technical means, when the camera is at the front of the vehicle, considering that the front-row users are closer to the camera and the same physical height occupies a larger proportion in the image, while the rear-row users are farther from the camera and the same physical height occupies a smaller proportion in the image. The front-row perspective magnification effect is eliminated by calculating the ratio of the front-row distance to the total distance, and the rear-row perspective reduction effect is compensated by calculating the ratio of the rear-row distance to the total distance. The spatial perspective height correction model converts the two-dimensional visual measurement value in the vehicle interior into height comparable data in the real physical space for the first time by establishing a distance-weight mapping function, fundamentally solving the problem of rush answering fairness caused by the seat position.

[0033] (6)Based on the above technical means, a three-level rush-answer priority determination rule is provided. The first rush-answer order includes the first level and the second level, and the second rush-answer order includes the third level. The first level is based on the frame timing priority principle to identify the frame number of the picture where the raising hand action first appears. The smaller the frame number, the higher the priority, which solves 90% of the conventional rush-answer scenarios. The second level is based on the height priority principle. When the frames are the same, the spatial perspective correction model is called to calculate the real raising hand height, and the one with a larger height value wins. Furthermore, combined with the subsequent third level based on the principle of lip movement timing matching, when the frames and heights are the same, the starting time point of pronunciation is locked through lip movement detection, and the consistency of lip movement-voice timing is verified by audio segmentation. The earliest effective speaker obtains the priority to rush-answer. The triple progressive determination covers the full dimension from millisecond-level timing (frame order) to behavioral characteristics (height), and then to biometric characteristics (lip movement), solving the problem of the failure of traditional acoustic solutions in the case of multiple people rushing to answer at the same time. 95% of the scenarios are directly determined by the first level (the calculation time consumption < 2ms), and 5% of the extreme scenarios trigger the third-level analysis, meeting the in-vehicle real-time requirement (the total delay ≤ 50ms). The three-level cooperation reduces the priority misjudgment rate to almost zero, significantly improving the accuracy of rush-answer.

[0034] (7)Based on the above technical means, the system can also set multiple designated rush-answer areas in the vehicle, such as setting touch areas on the door armrests near each seat, on the seat backs, or on the center console. The system detects the position of the user's palm through image recognition technology and determines whether the palm is in contact with the rush-answer area. When the system detects that the user's palm touches the rush-answer area, it is considered that the user has completed the rush-answer action. This rush-answer method is more simple and direct than the raising hand rush-answer. The user only needs to gently touch the rush-answer area to complete the rush-answer without making an obvious raising hand action. This method is more suitable for the situation where the in-vehicle space is limited. Under the condition of limited in-vehicle space, the rush-answer method of touching the rush-answer area can also reduce the misjudgment of the system for the user's non-standard actions and improve the accuracy of rush-answer recognition.

[0035] (8)Based on the above technical means, when collecting the images in the cockpit, the system collects the images in the cockpit through an RGB-IR type camera. The RGB-IR camera has both visible light (RGB) and infrared (IR) imaging capabilities, and can clearly capture the images in the cockpit under various lighting conditions (including low-light environments at night), ensuring that the system can accurately identify the user's rush-answer actions and lip movement information.

[0036] (9) Divide the regions of each position through visual detection and visualize them on the central control screen for the participants to confirm their own rush-answer regions. On the one hand, this can avoid misjudgment due to physical crossing during the rush-answer process (the algorithm itself is based on individual matching of the torso and limbs to avoid cross-matching). Introducing the position region division can further guide users and enhance the accuracy of differentiating rush-answer actions. On the other hand, through visualization, it guides the user's gesture actions to appear within the visible area, avoiding the failure of action determination caused by occlusion and improving the participation rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0038] Figure 1 It is a schematic flowchart of a vehicle-mounted multi-person rush-answer method according to an embodiment of the present invention; Figure 2 It is another schematic flowchart of a vehicle-mounted multi-person rush-answer method according to an embodiment of the present invention; Figure 3 It is another schematic flowchart of a vehicle-mounted multi-person rush-answer method according to an embodiment of the present invention; Figure 4 It is a schematic flowchart of rush-answer action determination according to an embodiment of the present invention; Figure 5 It is a schematic diagram of human key point acquisition according to an embodiment of the present invention; Figure 6 It is a schematic structural diagram of a vehicle-mounted multi-person rush-answer device according to an embodiment of the present invention; Figure 7 It is a schematic hardware structure diagram of a vehicle according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0040] According to an embodiment of the present invention, an embodiment of a vehicle-mounted multi-person quick-answer method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0041] In this embodiment, a vehicle-mounted multi-person quick-answer method is provided, which can be used in a vehicle. Figure 1 It is a flowchart of a vehicle-mounted multi-person quick-answer method according to an embodiment of the present invention. The process includes the following steps: Step S101, collect the image inside the cockpit and determine the position information of the users participating in the quick answer according to the image inside the cockpit.

[0042] Specifically, in the embodiment of the present invention, a camera is equipped inside the vehicle. When the user conducts an in-vehicle quick-answer game, the in-vehicle camera collects the in-vehicle picture in real time, and then analyzes the user's quick-answer actions, providing a pure-vision quick-answer determination method. In the embodiment of the present invention, the number and position of the in-vehicle cameras are not specifically limited. They can be at the front of the cockpit, the rear of the cockpit, and the door of the cockpit, as long as the cameras inside the cockpit can clearly collect the in-vehicle images at all angles inside the cockpit. In a specific implementation, the embodiment of the present invention can be reused for the in-vehicle monitoring high-definition camera arranged on the front-row ceiling, the front windshield area, or the instrument panel. It itself is a part of the cockpit OMS (Occupancy Monitoring System), used for in-vehicle photographing / video, occupant attribute detection, and detection of left-behind items and other conventional functions inside the cabin, without the need to add additional camera devices and sensing devices, reducing the production cost of the vehicle.

[0043] In an alternative embodiment, the in-vehicle camera uses an RGB-IR type camera. RGB-IR (Red Green Blue-Infrared) is an image sensor technology that combines visible light (RGB) and infrared light (IR) imaging. Its core lies in a special filter array. The RGB-IR camera has both visible light (RGB) and infrared (IR) imaging capabilities, can clearly capture the image inside the cockpit under various lighting conditions (including low-light environments at night), ensuring that the system can accurately identify the user's quick-answer actions, to ensure the image acquisition effect in a dim environment, and enable the vehicle-mounted multi-person quick-answer interaction method of the embodiment of the present invention to support the full interaction scenarios such as day and night.

[0044] After that, since the in-cabin image contains the position distribution of the passengers in the vehicle, the position information of the users participating in the quick answer is determined based on the in-cabin image. The system analyzes the in-cabin image through image processing technology, identifies the users at each position in the vehicle, and assigns a unique position identifier to each user, such as "driver's seat", "passenger seat", "left rear row", "middle rear row", "right rear row", etc. These position information will be used for subsequent quick answer action recognition and quick answer order determination. The identity of each user does not need to be bound through face recognition, voiceprint recognition, etc. Even for a strange user who gets in the car for the first time, a unique label can be assigned to them according to their seating position, so that the identity of the users participating in the quick answer can be uniquely represented by the position information, which not only ensures the accuracy of identity representation, but also reduces the complexity of the identity binding algorithm in the early stage of the quick answer game.

[0045] Step S102, output the question.

[0046] Step S103, recognize the quick answer actions within each position information.

[0047] Specifically, after the question is given, through the divided position areas, the in-cabin domain controller recognizes the quick answer actions within each position information from the collected in-cabin image. In the embodiment of the present invention, the quick answer action can be a raising hand action, a shaking head action, a clapping action or an action of touching a specified area. This embodiment only takes this as an example and is not limited thereto. As long as an action that can identify the user's intention to quickly answer can be recognized, it is within the protection scope of the embodiment of the present invention. For example, the vehicle prompts the user through voice to make a quick answer action within 5s.

[0048] Step S104, determine the first quick answer order of each user participating in the quick answer according to the order in which the quick answer actions appear in the frame of the picture.

[0049] Specifically, the system extracts the frames of the pictures when each user participating in the quick answer makes a quick answer action, determines the time order of each frame of the picture, and sorts each frame of the picture in the order from early to late. Thus, the order of the earlier frames of the picture can represent that the user at the corresponding position grabs the right to answer first, so as to determine the first quick answer order of the users participating in the quick answer.

[0050] Step S105, for the users in different positions in the first quick answer order, prompt the users to answer the questions in turn and collect the voice information of each user.

[0051] Step S106, for multiple users in the same position in the first quick answer order, prompt the users to answer the questions at the same time, and collect the lip movement information and voice information of each user during the process of the users in the same position answering the questions at the same time.

[0052] Step S107: Match the lip movement information with the voice information collected during the simultaneous answering process of users in the same order to obtain the answering data of users in the same order.

[0053] Step S108: Determine the starting time sequence frame of the lip movement information of each user according to the answering data, and judge the second answering order of users in the same order according to the starting time sequence frame of each user.

[0054] Step S109: Determine the answering result based on the first answering order, the second answering order and the voice information of each user.

[0055] Specifically, due to the limitation of the camera shooting frequency, the situation of the same order may also occur inevitably sometimes. Although this probability is very small, it will also affect the accuracy and fairness of the answering. Therefore, for the first answering order, there are still two situations. The first situation is that only one user has an answering action in each frame of the picture, which reflects the scenario where questions are not grabbed simultaneously. The second situation is that there are multiple users with answering actions in some of the picture frames, which reflects the scenario where multiple users grab questions simultaneously. For example, two users raise their hands at the same time and are captured in the same picture frame.

[0056] For the first situation, in the embodiment of the present invention, the vehicle can prompt users to answer questions in turn and collect the voice information of each user. For example, assume that the first answering order is: 1. Driver's seat, 2. Passenger seat, 3. Left rear row, 4. Middle rear row, 5. Right rear row, and the answering order is: the user in the driver's seat answers, the user in the passenger seat answers, the user in the left rear row answers, the user in the middle rear row answers, and the user in the right rear row answers. The system issues answering prompts to users in different orders through the in-vehicle speaker or the central control screen according to the first answering order, such as "Please ask the first user who grabbed the question to answer the question". When the user starts to answer the question, the system collects the voice information of the user through the in-vehicle microphone and stores the voice information in association with the position information of the user.

[0057] For the second situation, assume that the first answering order is: 1. Driver's seat, 2. Passenger seat, 3. Left rear row and middle rear row are tied, 4. Right rear row. Among them, the users in the left rear row and the middle rear row are both captured raising their hands at the same time in the 3rd order, so there are users in the same order in the left rear row and the middle rear row. The system prompts multiple users in the same order to answer questions simultaneously. For example, the system first prompts the driver's seat to answer, then prompts the passenger seat to answer, and then the system prompts through the in-vehicle speaker or the central control screen "Please ask the users in the left rear row and the middle rear row to answer the question simultaneously". During the process of users answering questions simultaneously, the system collects the lip movement information of each user through the in-vehicle camera, and at the same time collects the mixed voice information through the in-vehicle microphone.

[0058] Among them, during the process of multiple users answering questions simultaneously, since users speak at different speeds, the embodiments of the present invention propose to further determine which user among those with the same rank grabs the right to answer first according to the starting time sequence frame of lip movement information. However, considering that lip movement does not necessarily represent a user's true answer (for example, the first user does not speak the answer during the simultaneous answering process, but only opens the mouth without making a sound at the first moment, while the second user speaks the answer after thinking for a while. Obviously, in fact, the second user successfully grabs the answer, but it is inaccurate to determine that the first user grabs the answer successfully only according to the speed of lip movement), the embodiments of the present invention first perform the matching of lip movement information and voice information.

[0059] Specifically, as Figure 2 shown, the system analyzes the lip movement characteristics of each user, such as lip shape changes, opening and closing frequencies, etc., and compares them with the acoustic characteristics of the voice signal, so as to separate the mixed voice information and match it to the corresponding user, and obtain the independent answer data of each user.

[0060] After obtaining the answer data, analyze the lip movement information of each user according to the answer data, including the lip movement duration and the starting point of lip movement, to determine the starting time sequence frame when each user actually starts speaking. By comparing the time sequence of the starting time sequence frames of the lip movement information of different users, the system judges the second answering order among users with the same rank. For example, if the starting time sequence frame of user A's lip movement is the 100th frame and the starting time sequence frame of user B's lip movement is the 105th frame, then user A ranks before user B in the second answering order.

[0061] At the same time, compare the voice information included in the answer data of users with the same rank with the standard answer, and then it can be judged whether the answer content of the corresponding user is correct.

[0062] Finally, determine the answering result according to the above first answering order, second answering order and the voice information of each user.

[0063] Specifically, in the embodiments of the present invention, the determination process of the answering result can also adopt two specific schemes: The first scheme is to let each user speak out the answer content, and finally judge who is the fastest and most accurate winner according to the answering order and the correctness of the answer content of each user, and then proceed to the next question.

[0064] The second scheme is that if the winner who answers the question correctly first appears, then proceed to the next question. For example Figure 3As shown in the figure, the users are prompted to answer in turn based on the answering order of the first rush-answer order. The voice information of the first rush-answerer is compared with the preset standard answer to judge whether the answer is correct. If it is correct, the next question will be entered. If it is wrong, the next rush-answerer will answer, and the correctness of the answer content will be judged until the next question is entered because someone answers correctly or all rush-answerers are wrong. During the above process, if there are multiple users with the same ranking who succeed in rushing to answer at the same time, when it is the turn of the multiple participants who rush to answer at the same time, it is prompted that multiple participants answer the question at the same time, and based on the timing correlation of lip movement information and voice information, the speed of answering and the audio content of each participant are distinguished and matched to determine the second rush-answer order of each answerer, and based on their respective voice recognition results, the correctness of their respective answer content is confirmed. Similarly, if the voice result of the user with the earlier order in the second rush-answer order is correct, the next question will be entered. If it is wrong, the correctness of the voice result of the user with the later order in the second rush-answer order will be judged. If it is correct, the next question will be entered. If the answers of users with the same ranking are all incorrect, the next single rush-answerer will continue to answer until the next question is entered because someone answers correctly or all rush-answerers are wrong. Finally, according to the judgment of the corresponding user's rush-answer situation, the participation situation is counted, and the interaction information is output on the in-vehicle terminal to show the victory or defeat situation to each user.

[0065] The above two schemes for determining the rush-answer result can be flexibly selected and enabled according to the user's preference. The embodiments of the present invention do not limit which scheme for determining the rush-answer result must be used.

[0066] Through the technical solution provided by the embodiments of the present invention, the rush-answer order is judged based on the mechanism of visual detection. For users who simultaneously make rush-answer actions, they answer the question at the same time, and the lip movement pictures of each user are taken, and the rush-answer order is further determined according to the order of lip movement. For users who rush to answer at the same time, the true answers of each user are also obtained based on the matching result of lip movement information and voice information, avoiding language confusion. The present invention does not require the vehicle model to have a multi-zone configuration or an additional rush-answer induction device for sound source localization or rush-answer timing determination, nor does it require users to additionally access devices for interactive experience. It can directly use the existing cameras and microphones in the vehicle to realize multi-person rush-answer interaction, reducing the hardware cost and implementation difficulty. The present invention has no restrictions on the identities of the participants in the vehicle, does not require user information such as voiceprint or face to be entered, and does not require users to learn complex and cumbersome postures or action instructions. Users can participate in the rush-answer only through a simple raising hand action, enabling driving users of different ages to easily participate. The present invention determines the rush-answer order and position through real-time visual detection, is not affected by noise interference, and combines the timing correlation of lip movement information and audio information to avoid confusion errors in the voice system, improving the accuracy of audio localization and answer content recognition. Thus, the problems of strong hardware dependence, poor scene adaptability, and rough priority determination existing in the existing in-vehicle rush-answer schemes are solved.

[0067] In some alternative embodiments, step S103 includes: Step a1, identifying the included angle between the user's arms and the raising height of the hands within each position information; Step a2, when the included angle between the arms falls within a preset included angle range and the raising height of the hands is greater than a preset height threshold, determining that a rush-answer action occurs at the corresponding position.

[0068] Specifically, for the recognition scheme where the rush-answer action is a hand-raising action, the system recognizes the included angle between the user's arms and the raising height of the hands. When the included angle between the arms falls within a preset included angle range (for example, 60 degrees to 120 degrees) and the raising height of the hands is greater than a preset height threshold (for example, higher than 10 cm), it is determined that a rush-answer action occurs at the corresponding position. By using the dual constraints of the included angle between the arms and the raising height of the hands to judge whether the user makes the specified hand-raising action, the accuracy of action determination is improved.

[0069] In some alternative embodiments, step a1 includes: Step b1, identifying the wrist joint point and the elbow joint point of the user within the current position information; Step b2, drawing a line connecting the wrist joint point and the elbow joint point of the arm; Step b3, calculating the included angle between the arm within the current position information through the included angle between the arm line and the horizontal line of the cockpit screen; Step b4, identifying the key points of the user's arm and the key points of the head and torso within the current position information. The key points of the arm are predefined landmark points on the arm, and the key points of the head and torso are predefined landmark points on the head or torso; Step b5, calculating the raising height of the hands within the current position information according to the vertical distance between the key points of the arm and the key points of the head and torso.

[0070] Specifically, as Figure 4 shown, the system first identifies the wrist joint point and the elbow joint point of the user within the current position information. The system uses a human key point detection algorithm to locate the coordinates of the wrist joint point and the elbow joint point of the user in the cockpit image. Then, the system draws a line connecting the wrist joint point and the elbow joint point of the arm to form a straight line representing the direction of the arm. Next, the system calculates the included angle between the arm within the current position information through the included angle between the arm line and the horizontal line of the cockpit screen. If the included angle between the arm line and the horizontal line is 90 degrees, it means that the user's arm is raised vertically. For example Figure 5 shown, where point a is the wrist joint point of the user, point b is the elbow joint point of the user, and the included angle between the line connecting point a and point b and the horizontal line of the cockpit screen is the included angle of the arm .

[0071] For the calculation method of the raising height of the hands, the system first extracts the key points of the arm and the key points of the head and torso. The key points of the arm refer to predefined landmark points on the arm. This point is not fixedly restricted as long as it is on the arm. For exampleFigure 5 The center point c is the midpoint of the line connecting the elbow and the wrist. The head and torso key points are predefined landmark points on the head or torso. In some alternative embodiments, the head and torso key points may be the midpoint between the eyebrows of the user's face or the horizontal center point of the shoulders. In Figure 5 the midpoint d between the eyebrows of the face is used. Then, the raising height within the current position information is calculated based on the vertical distance between the arm key points and the head and torso key points. For example, Figure 5 the h in

[0072] Traditional gesture detection is prone to misjudgment when the vehicle jolts, while more than 85% of non-intentional actions (such as raising the hand to wipe sweat) can be filtered out through the preset arm angle range. The vertical distance detection enables the raising actions of passengers of different body types (children / adults) to be fairly recognized, with spatial self-adaptability. In addition, the two detection conditions can be calculated in parallel to complete the judgment within 10 ms, ensuring the algorithm efficiency.

[0073] Specifically, calculating the raising height within the current position information based on the vertical distance between the arm key points and the head and torso key points seems simple but is actually ingenious. It not only avoids the adaptability problems caused by setting a fixed threshold (for example, for passengers with different heights, arm lengths, and body fat percentages, the raising height calculated using this solution adapts to the passenger's body size, and the parameters are all under the same standard, and the different body sizes of passengers will not affect the raising height judgment standard), but also leaves an interface for subsequent spatial perspective correction.

[0074] According to the above technical means, by identifying specific key points on the arm and torso to calculate whether the arm movement meets the constraint conditions of the arm angle and the raising height, the solution principle is simple and the calculation is accurate, providing a method for judging the rush answer action that takes into account both efficiency and accuracy.

[0075] In some alternative embodiments, step b5 includes: Step c1, calculating the vertical distance between the arm key points and the head and torso key points; Step c2, determining the relative distance between the current position information and the camera, and determining the correction weight according to the relative distance; Step c3, determining the raising height by multiplying the correction weight and the vertical distance.

[0076] Specifically, the on-board monocular camera has an inherent defect due to the principle of perspective projection - the farther the object is from the lens, the smaller the image size. Assuming that the camera is in the front row of the car, it is typically shown that the actual hand-raising height of the user in the back row is 35cm, which only occupies 20 pixels in the picture, while the action of the same height in the front row occupies 60 pixels. If the pixel distance is compared directly, it may be mistakenly judged that the hand-raising height of the user in the back row is 67% lower than that of the front row, resulting in serious distortion in the priority determination of the quick response. Based on this, the embodiment of the present invention also calculates the relative distance between the position of each user who issues a quick response action and the camera, and then determines the corresponding correction weight based on the distance, thereby correcting the numerical value of the hand-raising height. This solution solves the unfair problem of quick response caused by the distortion of the two-dimensional picture measurement in the car space for the first time. This method can eliminate the visual differences caused by the different distances between the user and the camera, making the calculation of the hand-raising height more accurate.

[0077] In some optional implementations, when the camera is at the front of the vehicle, step c2 includes: Step d1, obtaining the distance between the front row of the vehicle seat and the front row of the camera, and obtaining the distance between the rear row of the vehicle seat and the rear row of the camera; Step d2, calculating the sum of the front row distance and the back row distance to obtain the total distance; Step d3, determining whether the current position information is in the front row or the back row; Step d4: If the current position information is in the front row, calculate the ratio of the front row distance to the total distance to obtain the correction weight; Step d5: If the current position information is in the back row, the ratio of the back row distance to the total distance is calculated to obtain a correction weight.

[0078] Specifically, the embodiment of the present invention is based on the vertical distance measured by the camera h The relative height is calculated based on the distance between the front and rear rows and the camera, and correction weights are used to perform the correction. In this embodiment of the present invention, considering that the distance between each row of seats and the camera is almost the same, this embodiment uses the specific scenario of two rows of seats as an example and defines the correction weights to include front row correction weights and rear row correction weights.

[0079] Therefore, the front row hand height , back row hand raising height ,in, Indicates the front row correction weight when the user sits in the front row, Indicates the rear seat correction weight when the user sits in the rear seat.

[0080] After that, the system obtains the front row distance from the front row of the vehicle seat to the camera and the rear row distance from the rear row of the vehicle seat to the camera. For example, the front row distance may be 0.8 meters and the rear row distance may be 1.5 meters. The system calculates the sum of the front row distance and the rear row distance to obtain the total distance, which is 2.3 meters in this example.

[0081] In the embodiments of the present invention, the size of an object in the picture completely depends on the distance-size inverse relationship. For the same physical height, an increase in distance will cause the picture size to shrink proportionally, resulting in a systematic measurement deviation.

[0082] Therefore, the correction weight calculation method provided by the embodiments of the present invention is as follows: , , where, represents the distance between the front row passenger and the camera, represents the distance between the rear row passenger and the camera. In the above example, the calculated front row correction weight is 0.8 / 2.3≈0.35, and the calculated rear row correction weight is 1.5 / 2.3≈0.65. This calculation method of correction weight takes into account the distance difference between the user and the camera, enabling the system to fairly compare the raising heights of users in different positions.

[0083] According to the above technical means, when the camera is at the front of the vehicle, considering that the front row user is closer to the camera and the same physical height occupies a larger proportion in the picture, while the rear row user is farther from the camera and the same physical height occupies a smaller proportion in the picture, the calculated correction weight needs to have an inverse relationship with the distance from the user to the camera. Therefore, the embodiments of the present invention eliminate the front row perspective magnification effect by calculating the ratio of the front row distance to the total distance, and compensate for the rear row perspective reduction effect by calculating the ratio of the rear row distance to the total distance. The spatial perspective height correction model fundamentally solves the fairness problem of the rush answer caused by the seat position by establishing a distance-weight mapping function and converting the two-dimensional visual measurement value in the vehicle into height comparable data in the real physical space for the first time.

[0084] In some alternative embodiments, step S104 includes: Step e1, extracting the picture frames when each user participating in the rush answer makes a rush answer action; Step e2, determining the time sequence of each picture frame and sorting each picture frame in ascending order of time to obtain the first sorting; Step e3, determining whether there is a target picture frame, and the target picture frame includes the rush answer actions of multiple users; Step e4, sorting the raising heights of the rush answer actions of different users in the target picture frame from high to low to obtain the second sorting, where if the raising heights of multiple target users in the target picture frame are also the same, the target users are defined as users with the same rank; Step e5, generate the first answering order according to the first sorting and the second sorting.

[0085] Specifically, according to the above technical means, a three-level answering priority determination rule is provided. The first answering order includes the determination of the first level and the second level, and the second answering order includes the determination of the third level. Among them, the first-level determination is based on the principle of priority of the video frame timing. Identify the video frame number when each user's raising hand action first appears. The smaller the frame number, the higher the priority, so as to obtain the first sorting, which can solve 90% of the conventional answering scenarios.

[0086] The second-level priority determination is based on the principle of height priority. When the raising hand actions of multiple users appear in the same frame, the embodiment of the present invention calls the spatial perspective correction model to calculate the real raising hand height of each user based on the foregoing steps, and then compares the raising hand heights of each user. Among them, the one with a larger raising hand height value wins and is determined to be more priority, obtaining the second sorting. Integrate the first sorting and the second sorting together to form a complete first answering order.

[0087] Combined with the subsequent third-level principle based on lip movement timing matching, only when the raising hand actions of users appear in the same frame and at the same height, the starting time point of pronunciation is locked through lip movement detection, and the lip movement-voice timing consistency is verified by combining audio segmentation. The earliest effective pronouncer obtains the priority to answer first, obtaining the second answering order.

[0088] The triple progressive determination provided by the embodiment of the present invention covers the full dimension from millisecond-level timing (frame order) to behavioral characteristics (height) and then to biometric characteristics (lip movement), solving the problem of the failure of traditional acoustic solutions when multiple people answer at the same time. 95% of the scenarios are directly determined at the first level (the calculation time < 2 ms), and 5% of the extreme scenarios trigger the third-level analysis, meeting the in-vehicle real-time requirements (the total delay ≤ 50 ms). The three-level collaboration reduces the priority misjudgment rate to almost zero, significantly improving the answering accuracy.

[0089] In some alternative embodiments, the above step S103 further includes: Step f1, identify whether the palm of the user in each position information touches the specified area; Step f2, when the palm of the user touches the specified area, determine that a answering action appears at the corresponding position.

[0090] Specifically, in the embodiments of the present invention, the system can also set multiple designated rush-answer areas inside the vehicle. For example, touch areas can be set on the door armrests near each seat, the seat backs, the ceiling, or the center console. The system detects the position of the user's palm through image recognition technology and determines whether the palm is in contact with the designated area. When the system detects that the user's palm touches the designated area, it is considered that the user has completed the rush-answer action. This rush-answer method is simpler and more direct than raising a hand to answer. The user only needs to gently touch the designated area to complete the rush-answer, without having to make an obvious hand-raising action. This method is more suitable for the situation where the space inside the vehicle is limited. Under the condition of limited space inside the vehicle, the rush-answer method of touching the designated area can also reduce the misjudgment of non-standard actions by the system and improve the accuracy of rush-answer recognition.

[0091] In this embodiment, a vehicle-mounted multi-person rush-answer device is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0092] In some alternative implementation manners, the vehicle-mounted multi-person rush-answer method provided by the present invention further includes: dividing the position information into regions and visualizing the divided regions on the vehicle center control screen so that the participants can confirm their own rush-answer regions.

[0093] According to the above technical means, during the rush-answer process by rush-answer actions, the regions of each position are divided through visual detection and visualized on the center control screen for the participants to confirm their own rush-answer regions. On the one hand, it avoids misjudgment of limb crossing during the rush-answer process (the algorithm itself is based on individual matching of the torso and limbs to avoid cross-matching). Introducing the division of position regions can further guide the user and enhance the accuracy of rush-answer action detection and differentiation. On the other hand, through visualization, the user's gesture actions are guided to appear within the visible area, avoiding the failure of action determination caused by occlusion and improving the participation rate.

[0094] This embodiment provides a vehicle-mounted multi-person rush-answer device, as Figure 6 shown, including: A position binding module 601, configured to collect images inside the cockpit and determine the position information of the users participating in the rush-answer according to the images inside the cockpit; A question-setting module 602, which outputs questions; An action recognition module 603, configured to recognize the rush-answer actions within each position information; A first rush-answer order determination module 604, configured to determine the first rush-answer order of each user participating in the rush-answer according to the order in which the rush-answer actions appear in the picture frames; The first answering module 605 is used to prompt users in different positions in the first rush-answer order to answer questions in turn and collect the voice information of each user; The second answering module 606 is used to prompt multiple users in the same position in the first rush-answer order to answer questions simultaneously, and collect the lip movement information and voice information of each user during the simultaneous answering of users in the same position; The matching module 607 is used to match the lip movement information and the voice information collected during the simultaneous answering of users in the same position to obtain the answering data of users in the same position; The second rush-answer order determination module 608 is used to determine the starting time sequence frame of the lip movement information of each user according to the answering data, and judge the second rush-answer order of users in the same position according to the starting time sequence frame of each user; The result module 609 is used to determine the rush-answer result based on the first rush-answer order, the second rush-answer order and the voice information of each user.

[0095] In some alternative embodiments, the motion recognition module 603 includes: A user recognition unit for recognizing the arm angle and raising height of the user within each position information; A motion determination unit for determining that a rush-answer motion appears at the corresponding position when the arm angle falls within a preset angle range and the raising height is greater than a preset height threshold.

[0096] In some alternative embodiments, the user recognition unit includes: An arm joint point recognition subunit for recognizing the wrist joint point and elbow joint point of the user within the current position information; An arm direction drawing subunit for drawing the arm connection line between the wrist joint point and the elbow joint point; An angle calculation subunit for calculating the arm angle within the current position information through the angle between the arm connection line and the horizontal line of the cockpit screen; A body key point recognition subunit for recognizing the arm key points and head and body key points of the user within the current position information, where the arm key points are predefined landmark points on the arm, and the head and body key points are predefined landmark points on the head or torso; A height calculation subunit for calculating the raising height within the current position information according to the vertical distance between the arm key points and the head and body key points.

[0097] In some alternative embodiments, the height calculation subunit includes: A vertical distance calculation subunit for calculating the vertical distance between the arm key points and the head and body key points; A correction weight calculation subunit for determining the relative distance between the current position information and the camera, and determining the correction weight according to the relative distance; A correction subunit, configured to determine the raising height of the hand by using the product of the correction weight and the vertical distance.

[0098] In some alternative embodiments, the first rush answering order determining module 604 includes: A screen extraction unit, configured to extract the screen frames when each user participating in the rush answering makes a rush answering action; A first sorting unit, configured to determine the chronological order of each screen frame, and sort each screen frame in ascending order of time to obtain a first sorting; A target frame determining unit, configured to determine whether there is a target screen frame, where the target screen frame includes the rush answering actions of multiple users; A second sorting unit, configured to sort the raising heights of the rush answering actions of different users in the target screen frame from high to low to obtain a second sorting, where if the raising heights of multiple target users in the target screen frame are the same, the target users are defined as users with the same rank; A first rush answering order determining unit, configured to generate a first rush answering order according to the first sorting and the second sorting.

[0099] In some alternative embodiments, the action recognition module 603 further includes: A region determination unit, configured to identify whether the palm of the user within each position information touches a specified region; An action determination unit, configured to determine that a rush answering action appears at the corresponding position when the palm of the user touches the specified region.

[0100] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding above-mentioned embodiments, and will not be elaborated here.

[0101] The embodiment of the present invention further provides a vehicle, as Figure 7 shown, including: a memory, a cockpit domain controller, an in-vehicle camera, an in-vehicle microphone, an in-vehicle central control screen, and an in-vehicle speaker; the memory, the in-vehicle camera, the in-vehicle microphone, the in-vehicle central control screen, and the in-vehicle speaker are all communicatively connected to the cockpit domain controller; the in-vehicle camera is used to collect images inside the cockpit, the in-vehicle microphone is used to pick up the voice information of the users participating in the rush answering, the in-vehicle central control screen is used to display the overall application interface and interaction content, and the in-vehicle speaker is used to broadcast the question content and prompt information in the rush answering interaction stage. Among them, the in-vehicle camera, the in-vehicle microphone, the in-vehicle central control screen, and the in-vehicle speaker belong to the hardware layer.

[0102] The device model provided by the above device embodiment is stored in the memory, corresponding to the computer instructions provided by the method embodiment. The cockpit domain controller executes the computer instructions to execute the method provided by the foregoing method embodiment. Thus, preset rush-answer questions are randomly selected according to the rush-answer game type for display and broadcast, the order of answering is determined based on the real-time detection of the rush-answer behavior of the participants, the answering speed and content matching of multiple people who successfully rush-answer at the same time are distinguished, the correctness of the answering content is determined based on the analysis and comparison of the voice content, and the situations of each stage of the rush-answer interaction are statistically displayed.

[0103] In the multi-person rush-answer interaction stage, the hardware layer in the vehicle acquires visual image information, which is processed by the cockpit domain controller, and the software module of the memory is called for interaction logic determination, and then the hardware layer performs visual display and voice prompt. The overall interaction system is mainly based on visual detection in terms of software and hardware, realizing interactions such as rush-answer priority, position matching, and answering speed differentiation. The voice recognition technology only determines the answering accuracy.

[0104] The cockpit domain controller can be a central processing unit, a network processor, or a combination thereof. Among them, the cockpit domain controller can further include a hardware chip. The above hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.

[0105] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored on such a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware for software processing. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiment is implemented.

[0106] A part of the present invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can call or provide the methods and / or technical solutions according to the present invention through the operations of the computer. Those skilled in the art should understand that the forms of existence of computer program instructions in a computer-readable medium include but are not limited to source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.

[0107] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for multiple-person quick answering in a vehicle, characterized in that, The method includes: Collecting an image inside the cockpit and determining the position information of the users participating in the quick answer based on the image inside the cockpit; Outputting a question; Identifying the quick answer actions within each position information; Determining the first quick answer order of each user participating in the quick answer according to the order in which the quick answer actions appear in the frame of the picture; For users with different ranks in the first quick answer order, prompting the users to answer questions in sequence and collecting the voice information of each user; For multiple users with the same rank in the first quick answer order, prompting the users to answer questions simultaneously, and collecting the lip movement information and voice information of each user during the simultaneous answering of the users with the same rank; Matching the lip movement information and the voice information collected during the simultaneous answering of the users with the same rank to obtain the answer data of the users with the same rank; Determining the starting time sequence frame of the lip movement information of each user according to the answer data, and judging the second quick answer order of the users with the same rank according to the starting time sequence frame of each user; Determining the quick answer result based on the first quick answer order, the second quick answer order, and the voice information of each user.

2. The method according to claim 1, characterized in that The identifying the quick answer actions within each position information includes: Identifying the arm angle and the raising height of the hand of the user within each position information; When the arm angle falls within a preset angle range and the raising height of the hand is greater than a preset height threshold, determining that the quick answer action appears at the corresponding position.

3. The method according to claim 2, characterized in that The identifying the arm angle and the raising height of the hand of the user within each position information includes: Identifying the wrist joint point and the elbow joint point of the user within the current position information; Drawing the arm connection line between the wrist joint point and the elbow joint point; Calculating the arm angle within the current position information through the angle between the arm connection line and the horizontal line of the cockpit picture; Identifying the arm key point and the head and torso key point of the user within the current position information, where the arm key point is a predefined landmark point on the arm, and the head and torso key point is a predefined landmark point on the head or torso; Calculating the raising height of the hand within the current position information according to the vertical distance between the arm key point and the head and torso key point.

4. The method according to claim 3, wherein The calculating the raising height of the hand within the current position information according to the vertical distance between the arm key point and the head and torso key point includes: Calculating the vertical distance between the arm key point and the head and torso key point; Determining the relative distance between the current position information and the camera, and determining the correction weight according to the relative distance; Determining the raising height of the hand by multiplying the correction weight and the vertical distance.

5. The method according to claim 4, wherein When the camera is at the position of the vehicle head, the determining the relative distance between the current position information and the camera, and determining the correction weight according to the relative distance includes: Obtaining the front row distance from the front row of the vehicle seat to the camera, and obtaining the rear row distance from the rear row of the vehicle seat to the camera; Calculating the sum of the front row distance and the rear row distance to obtain the total distance; Judging whether the current position information is in the front row or the rear row; If the current position information is in the front row, calculating the ratio of the front row distance to the total distance to obtain the correction weight; If the current position information is in the rear row, calculating the ratio of the rear row distance to the total distance to obtain the correction weight.

6. The method according to claim 2, wherein Determining the first rush-answer order of each participating user according to the order in which the rush-answer actions appear in the video frames includes: Extracting the video frames when each participating user performs a rush-answer action; Determining the time order of each video frame, and sorting each video frame in ascending order of time to obtain a first sorting; Judging whether there is a target video frame that includes rush-answer actions of multiple users; Sorting the raising heights of the rush-answer actions of different users in the target video frame from high to low to obtain a second sorting, where if the raising heights of multiple target users in the target video frame are the same, the target users are defined as users with the same rank; Generating the first rush-answer order according to the first sorting and the second sorting.

7. The method according to claim 1, wherein Identifying the rush-answer actions within each position information further includes: Identifying whether the palm of the user within each position information touches a specified area; When the palm of the user touches the specified area, determining that the rush-answer action appears at the corresponding position.

8. The method according to claim 1, wherein Collecting the in-cockpit image includes: Collecting the in-cockpit image through a RGB-IR type camera.

9. The method according to claim 1, characterized in that, The method further includes: Dividing the position information into regions, and visualizing the divided regions on the vehicle center control screen so that the answering participants can confirm their own rush-answer regions.

10. A vehicle-mounted multi-person quick-answer device, characterized in that, The device includes: A position binding module, configured to collect the in-cockpit image and determine the position information of the participating users according to the in-cockpit image; A question-setting module, which outputs questions; An action recognition module, configured to recognize the rush-answer actions within each position information; A first rush-answer order determination module, configured to determine the first rush-answer order of each participating user according to the order in which the rush-answer actions appear in the video frames; A first answering module, configured to prompt the users to answer questions in sequence according to the users with different ranks in the first rush-answer order and collect the voice information of each user; A second answering module, configured to prompt the multiple users with the same rank in the first rush-answer order to answer questions simultaneously, and collect the lip movement information and voice information of each user during the simultaneous answering of the users with the same rank; A matching module, configured to match the lip movement information with the voice information collected during the simultaneous answering of the users with the same rank to obtain the answer data of the users with the same rank; A second rush-answer order determination module, configured to determine the starting time sequence frame of the lip movement information of each user according to the answer data, and judge the second rush-answer order of the users with the same rank according to the starting time sequence frame of each user; A result module, configured to determine the rush-answer result based on the first rush-answer order, the second rush-answer order, and the voice information of each user.

11. A vehicle, characterized in that, Includes: A memory, a cockpit domain controller, an in-vehicle camera, an in-vehicle microphone, an in-vehicle center control screen, and an in-vehicle speaker; The memory, the in-vehicle camera, the in-vehicle microphone, the in-vehicle central control screen, and the in-vehicle speaker are all communicatively connected to the cockpit domain controller; the in-vehicle camera is used to collect images inside the cockpit, the in-vehicle microphone is used to pick up the voice information of the users participating in the quick answer, the in-vehicle central control screen is used to display the overall application interface and interaction content, the in-vehicle speaker is used to broadcast the question content and prompt information during the quick answer interaction stage, computer instructions are stored in the memory, and the cockpit domain controller executes the computer instructions to execute the method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

13. A computer program product, characterized in that, It includes computer instructions, and the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • In-vehicle entertainment interaction method and device, vehicle, and machine readable medium

    CN110211585A

  • Voice responding method, device and system

    CN110838211A

  • Vehicle-mounted preemptive answering system and method

    CN113256920A

  • Classroom preemptive answering implementation method and device, computer equipment and storage medium

    CN115689831A

  • Virtual image interaction method, related device, equipment, system and medium

    CN116088675A