Car lamp entertainment interaction method and device and computer readable storage medium

By using infrared sensors and front-view cameras in the headlight system combined with deep learning network to detect user gestures, identify and convert them into operation instructions, and control projected headlight screens, the problem of the headlight system being unable to interact instantly is solved, and the user's entertainment experience and security are improved.

CN120302489APending Publication Date: 2025-07-11ZHEJIANG ZEEKR INTELLIGENT TECH CO LTD +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510351720.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing headlight system cannot realize instant interaction between users and headlights, and lacks the ability to respond to user operations in real time, reducing the fun and sense of participation of the entertainment experience.

Method used

By obtaining the real-life image of the interactive area projected by the projection headlights in front of the vehicle, using infrared sensors and front-view cameras to capture user action and environment images, combining deep learning convolutional neural network to detect human key points, identify user interaction behaviors, and convert them into in-game operation instructions, and control the projection headlights to project corresponding pictures in the interactive area.

Benefits of technology

It realizes that users directly control the car light entertainment system through gestures, provides a more intuitive and natural interaction method, enriches the user's driving experience, and ensures that the entertainment function is activated under appropriate conditions through safety status judgments, avoiding potential safety hazards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302489A_ABST
    Figure CN120302489A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle lamp entertainment interaction method and device and a computer readable storage medium, and the method comprises the steps: obtaining a live-action image of an interaction region projected by a projection headlamp in front of a vehicle, and detecting human body key points in the behavior performance of an interaction object participating in a game indicated by the live-action image in the interaction region, and analyzing and identifying through a visual algorithm, and determining the interaction behavior of the interaction object. The interaction behavior is converted into an operation instruction in a game, and the projection headlamp is controlled to project a picture indicated by the operation instruction in the interaction area, so that a user is allowed to directly control a vehicle lamp entertainment system through gestures, and the driving experience of the user is greatly enriched and improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of vehicle lighting systems, and in particular to a vehicle lighting entertainment interaction method, device, and computer-readable storage medium. Background Art

[0002] With the development of smart car technology and the improvement of users' requirements for driving experience, the lighting system of vehicles is no longer limited to traditional lighting and signal indication functions. Modern car lighting systems are gradually incorporating more intelligent and entertaining elements to meet users' needs for personalized and interactive experience. For example, the existing car lighting system unilaterally displays information or images to the outside world, informs the outside world of established information, lacks the ability to respond to user operations in real time, and cannot achieve instant interaction between users and car lights, reducing the fun and sense of participation in the entertainment experience. Summary of the invention

[0003] In order to overcome the problems existing in the related art, this specification provides a vehicle light entertainment interaction method, device and computer-readable storage medium.

[0004] According to a first aspect of an embodiment of this specification, a vehicle light entertainment interaction method is provided, the method comprising:

[0005] Acquire a real-scene image of the interactive area projected by the projection headlight in front of the vehicle, wherein the real-scene image indicates the behavior performance of the interactive objects participating in the game in the interactive area;

[0006] Detecting key points of a human body in the real scene image to determine the interactive behavior of the interactive object;

[0007] Convert the interactive behavior into in-game operation instructions;

[0008] The projection headlight is controlled to project a picture indicated by the operation instruction in the interactive area.

[0009] According to a vehicle light entertainment interaction method provided by the present application, the detecting of key points of a human body in the real scene image and determining the interactive behavior of the interactive object include:

[0010] Detecting key points of a human body in the real scene image to determine key point information of the human body;

[0011] Based on the human body key point information, the interactive behavior of the interactive object is determined through a deep learning convolutional neural network.

[0012] According to a headlight entertainment interaction method provided by the present application, a human body detection model is used to detect key points of a human body in the real scene image, and the human body detection model includes a basic feature extraction module, a human body area focusing module, a detection frame generation module and a key point detection module;

[0013] The basic feature extraction module is used to extract features from the real scene image and output a basic feature map;

[0014] The human body region focusing module is used to perform convolution and pooling operations on the basic feature map, focus on the features of the human body region where the interaction object is located in the real scene image, and output a human body region feature map;

[0015] The detection box generation module is used to perform a regression operation on the human body region feature map to output a detection box representing the human body boundary, and is also used to analyze the features within the detection box through convolution and fully connected layers to output the human body features within the detection box;

[0016] The key point detection module is used to perform convolution and deconvolution operations on the human body features within the detection box to extract the features of the human body, and convert the feature map obtained after feature extraction into a heat map to determine the human body key point information.

[0017] According to a vehicle headlight entertainment interaction method provided by the present application, the converting the feature map obtained after feature extraction into a heat map to determine the information of the human body key points includes:

[0018] Converting the feature map obtained after feature extraction into a heat map, where each pixel value in the heat map represents the probability that this position is a key point;

[0019] Performing threshold processing and non-maximum suppression operations on the heat map to determine the human body key points;

[0020] Converting the coordinates of the human body key points in the heat map to determine the coordinate information of the human body key points in the image where the detection box is located as the human body key point information.

[0021] According to a vehicle headlight entertainment interaction method provided by the present application, during the process that the human body region focusing module is used to perform convolution and pooling operations on the basic feature map, focus on the features of the human body region in the real scene image, and output a human body region feature map, the receptive field of the convolution kernel is dynamically adjusted according to the scale size and position information of the interaction object in the human body region feature map; as the human body scale becomes larger, the receptive field is increased.

[0022] According to a vehicle headlight entertainment interaction method provided by the present application, the real scene image includes an infrared image and an RGB image.

[0023] According to a vehicle headlight entertainment interaction method provided by the present application, the basic feature extraction module is used to extract features from the real scene image and output a basic feature map, including:

[0024] The basic feature extraction module includes a two-stream network, and the two-stream network includes a residual network. The convolutional layer of the residual network performs a convolutional operation on the RGB image to extract the feature map of the RGB image;

[0025] The convolutional layer of the residual network extracts features from the infrared image to obtain the feature map of the infrared image;

[0026] The feature map of the RGB image and the feature map of the infrared image are superimposed and fused to obtain the basic feature map.

[0027] According to a vehicle headlight entertainment interaction method provided by the present application, the converting the interaction behavior into an operation instruction in the game includes:

[0028] The interaction behavior is matched with the operation behavior information of the user game interaction pre-stored in the game animation library to obtain the target operation behavior;

[0029] According to the target operation behavior, the pictures pre-stored in the game animation library are retrieved; the pictures include game animations and result feedback pictures;

[0030] Based on the target operation behavior, a corresponding operation instruction is generated, and the operation instruction includes the picture corresponding to the target operation behavior.

[0031] According to a vehicle headlight entertainment interaction method provided by the present application, the method further includes:

[0032] Collect the voice information of the interaction object;

[0033] Combined with the voice information and the interaction behavior, an operation instruction in the game is generated.

[0034] According to a vehicle headlight entertainment interaction method provided by the present application, before obtaining the real scene image of the interaction area projected by the projection headlight in front of the vehicle, the method further includes:

[0035] Obtain the status data of the vehicle;

[0036] Combined with a preset game trigger condition, it is judged whether to activate the entertainment function of the vehicle; the game trigger condition includes that the status data meets the set safety threshold;

[0037] After activating the entertainment function of the vehicle, execute obtaining the real scene image of the interaction area projected by the projection headlight in front of the vehicle

[0038] The present application also provides a vehicle headlight entertainment interaction device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the vehicle headlight entertainment interaction method as described in any one of the above.

[0039] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for interactive entertainment of vehicle lights as described in any one of the above is implemented.

[0040] In the embodiments of this specification, the method, device and computer-readable storage medium for interactive entertainment of vehicle lights, compared with the current situation where instant interaction between the user and the vehicle lights cannot be achieved, by obtaining the real-scene image of the interactive area projected by the projection headlight in front of the vehicle, detecting the human key points in the behavior performance of the interaction object indicated by the real-scene image in the interactive area, analyzing and identifying through visual algorithms, and determining the interaction behavior of the interaction object. Converting the interaction behavior into an operation instruction in the game, and controlling the projection headlight to project the picture indicated by the operation instruction in the interactive area, so as to allow the user to directly control the vehicle light entertainment system through gestures, greatly enriching and enhancing the user's driving experience.

[0041] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this specification, and are used together with the specification to explain the principles of this specification.

[0043] Figure 1 is a schematic flowchart of a vehicle light entertainment interaction system shown according to an exemplary embodiment of this specification;

[0044] Figure 2 is a schematic flowchart of a method for interactive entertainment of vehicle lights shown according to an exemplary embodiment of this specification;

[0045] Figure 3 is a schematic structural diagram of a human detection model shown according to an exemplary embodiment of this specification;

[0046] Figures 4(a)-4(c) is an interaction diagram of a method for interactive entertainment of vehicle lights shown according to an exemplary embodiment of this specification;

[0047] Figure 5 is a schematic block diagram of a vehicle light entertainment interaction device shown according to an exemplary embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] Here, the technical solutions in the embodiments (or "embodiment modes") of the present application will be clearly and completely described in conjunction with the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements.

[0049] If there are terms related to directional indication or positional relationship in the embodiments of this application (such as up, down, left, right, front, back, inside, outside, top, bottom, center, vertical, horizontal, longitudinal, transverse, length, width, counterclockwise, clockwise, axial, radial, circumferential, etc.), such terms are only used to explain the relative positional relationship, movement conditions, etc. between components in a specific posture (as shown in the drawings); if the specific posture changes, the directional indication or positional relationship will also change accordingly. In addition, the terms "first", "second", etc. involved in the embodiments of this application are only for the purpose of convenient description and should not be construed as indicating or implying relative importance.

[0050] This application provides a method, device, and computer-readable storage medium for headlight entertainment interaction. The following will describe this application in detail with reference to the drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.

[0051] With the development of intelligent vehicle technology and the improvement of users' requirements for driving experience, the vehicle lighting system is no longer limited to traditional lighting and signal indication functions. Modern vehicle lighting systems are gradually integrating more intelligent and entertainment elements to meet users' needs for personalized and interactive experiences. Against this background, a headlight entertainment solution based on vision algorithms has emerged, aiming to provide a richer and more interactive entertainment experience through intelligent headlights.

[0052] This specification provides a method for headlight entertainment interaction.

[0053] Aiming to achieve real-time recognition and response to the external environment of the vehicle and user gestures through a set vision algorithm in combination with an in-vehicle camera and sensors, thereby providing a brand-new headlight entertainment experience. This solution can not only recognize and track users' gestures but also dynamically adjust the projection content of the headlights according to users' behaviors and preferences, realizing truly personalized and interactive headlight entertainment. In addition, this solution will also consider the safety status of the vehicle to ensure that the entertainment function is activated under appropriate conditions to avoid potential safety hazards.

[0054] Figure 1 It is a schematic flowchart of a headlight entertainment interaction system provided by an embodiment of this specification. The system includes an information acquisition unit, a data processing unit, a projection unit, and an in-vehicle display unit.

[0055] Among them, the information collection unit includes, but is not limited to, an infrared sensor and a front-view camera, which are used to capture the user's actions and environmental images in the interaction area. The infrared sensor is used to capture the shape and actions of the human body and generate an infrared humanoid image. These images reflect the posture and position of the user in the interaction area. The front-view camera is responsible for capturing the environmental images in front of the user or the vehicle and may be used to assist the infrared sensor to provide richer environmental information.

[0056] The data processing unit is responsible for processing all data from sensors and user interactions, including image processing results and game logic calculations, to ensure the smooth operation and response of the game. As an example, an image algorithm processing unit is configured in the data processing unit, and the image algorithm processing unit is used to process the captured image data (including data from the infrared sensor and the front-view camera). The image data is analyzed by algorithms to identify the user's posture, actions, and possible interaction intentions. The processed results are used to convert into operation instructions in the game, and then the projection unit projects the screen indicated by the operation instructions in the interaction area to achieve intelligent linkage and real-time interaction response with the vehicle state.

[0057] The interaction area projection unit creates an interaction area on the ground in front of the vehicle, where users can interact with the projected content through body movements. It should be noted that in other examples, the physical space range where users can carry out game interactions is defined to ensure the safety and effectiveness of the game. That is

[0058] The projection unit includes, but is not limited to, in-vehicle DLP (Digital Light Processing) projection headlights, and the data processed by the image algorithm is used to control the DLP projection headlights. The DLP headlights project corresponding images or animations on the ground or other surfaces in front of the vehicle according to the processing results.

[0059] The in-vehicle display unit includes, but is not limited to, a central control screen, which is used to display the game interface and the pictures of the interaction process. Among them, the game system has a built-in game animation library, which contains a variety of game animations and visual effects to enhance the game experience. The game system also has a built-in rule library corresponding to the animations, and the game rule library defines the gameplay and rules of the game to ensure the fairness and consistency of the game interaction. Therefore, the interaction between the user and the game interface is captured by the infrared sensor and the front-view camera, and these actions are converted into operation instructions in the game based on the corresponding game animation library and game rule library.

[0060] The above-mentioned vehicle headlight entertainment interaction system is a closed-loop system. Starting from the capture of the user's actions, through data processing and game logic judgment, the game results are finally fed back to the user through the DLP projection headlights, forming a complete interaction experience. This system can provide a novel in-vehicle entertainment method and enhance the interaction between the user and the vehicle.

[0061] Specifically, it includes the following steps:

[0062] Figure 2 It is a schematic flowchart of a vehicle headlight entertainment interaction method provided by an embodiment of this specification, including the following steps:

[0063] Step S100, obtain a real-scene image of the interaction area projected by the projection headlight in front of the vehicle, and the real-scene image indicates the behavior performance of the interaction object participating in the game in the interaction area.

[0064] The projection unit of the vehicle projects an area for user interaction in front of the vehicle. The information acquisition unit collects the user behavior and environmental information of the interaction area to obtain a real-scene image of the interaction area. This real-scene image is used for subsequent analysis of the gestures of the interaction object, so as to enable the user to interact with the vehicle headlight entertainment system through gestures.

[0065] As an example, the information acquisition unit includes but is not limited to an infrared sensor and a front-view camera. Furthermore, the real-scene image includes an infrared image and an RGB image. The fusion of the two features lays a foundation for subsequent more accurate extraction of human-related information.

[0066] In some embodiments, safety is considered when designing the entertainment function. By monitoring the vehicle state, it is ensured that the entertainment function is activated only under appropriate conditions, reducing potential safety risks.

[0067] As an example, before obtaining the real-scene image of the interaction area projected by the projection headlight in front of the vehicle, the method further includes:

[0068] Obtain the state data of the vehicle;

[0069] Combined with preset game trigger conditions, determine whether to activate the entertainment function of the vehicle; the game trigger conditions include that the state data meets the set safety threshold;

[0070] After activating the entertainment function of the vehicle, execute obtaining the real-scene image of the interaction area projected by the projection headlight in front of the vehicle.

[0071] The state data of the vehicle described in this specification includes data such as vehicle speed, gear position, and light state. According to these states, combined with preset trigger conditions, it is intelligently determined whether to activate or deactivate the vehicle headlight entertainment function to ensure safety.

[0072] In this embodiment, the state data is collected through vehicle sensors, and it is intelligently determined whether it is currently suitable to activate the entertainment function to avoid potential safety hazards and ensure an entertainment experience in a safe situation.

[0073] Step S200: Detect the human key points in the real-scene image to determine the interaction behavior of the interaction object.

[0074] The headlight entertainment interaction method involves the information collection unit collecting the behavior performance of the interaction object in the interaction area and the environmental information in the interaction area, and processing the collected information to determine the interaction behavior of the interaction object.

[0075] In some embodiments, the detecting the human key points in the real-scene image to determine the interaction behavior of the interaction object includes:

[0076] Step S210: Detect the human key points in the real-scene image to determine the human key point information;

[0077] Step S220: Based on the human key point information, determine the interaction behavior of the interaction object through the convolutional neural network of deep learning.

[0078] In step S210, with reference to Figure 3 , Figure 3 is a schematic structural diagram of a human detection model provided by an embodiment of this specification. The neural network structure of the human detection model aims to accurately output the human detection frame and the human key point information from the input image.

[0079] The human detection model includes a basic feature extraction module, a human region focusing module, a detection frame generation module, and a key point detection module.

[0080] The basic feature extraction module adopts the architecture of a two-stream network enhanced by a non-local module as the basic feature extractor. Its core is to utilize the powerful feature extraction ability of the deep convolutional neural network. Here, some modules in ResNet50 are used to respectively perform preliminary feature extraction on the input real-scene image, and finally the two features will be fused together by superposition.

[0081] As an example, the basic feature extraction module is used to perform feature extraction on the real-scene image and output the basic feature map, including:

[0082] The basic feature extraction module includes a two-stream network, and the two-stream network includes a residual network. Through the convolutional layer of the residual network, convolutional operation is performed on the RGB image to extract the feature map of the RGB image;

[0083] Through the convolutional layer of the residual network, feature extraction is performed on the infrared image to obtain the feature map of the infrared image;

[0084] Overlay and fuse the feature maps of the RGB image and the infrared image to obtain the basic feature map.

[0085] Send the input RGB image and infrared image into some modules in ResNet50 (ResNet50 is a deep convolutional neural network that includes multiple convolutional layers, pooling layers, and residual connections, etc.) for preliminary feature extraction. In this process, the convolutional layer performs convolution operations by sliding the convolution kernel on the image to extract local features of the image, and the pooling layer reduces the dimensionality of the features to reduce the amount of calculation and the risk of overfitting. After being processed by ResNet50, the RGB image and the infrared image respectively obtain their own feature representations, and finally these two features are fused together by superposition to obtain the basic feature map with rich information and a general feature representation.

[0086] Through the above method, a general feature representation of the image, a fused feature map can be obtained. This feature map contains both information such as color and texture in the RGB image and thermal imaging information in the infrared image, providing rich basic features for subsequent focusing on the human body area.

[0087] On the basis of basic feature extraction, the human body area focusing module designs a dedicated human body area focusing sub-network. It is used to perform convolution and pooling operations on the basic feature map to focus on the features of the human body area where the interaction object is located in the real scene image, and output a human body area feature map.

[0088] Specifically, perform a series of combinations of convolutional layers and pooling layers on the feature map output by the basic feature extraction module. Among them, the parameters of the convolutional layer are carefully designed and trained to enhance the sensitivity to features such as human body contours and postures, enabling the network to automatically learn and focus on the human body area in the image, thereby effectively distinguishing the human body from the background information. That is, the features of the human body area are gradually highlighted through the above processing.

[0089] At the same time, an adaptive receptive field adjustment mechanism is adopted to dynamically adjust the receptive field of the convolution kernel according to the size and position of the human body in the input image to ensure that human body areas of different scales and positions can be accurately captured.

[0090] As an example, in the process where the human body area focusing module is used to perform convolution and pooling operations on the basic feature map to focus on the features of the human body area in the real scene image and output a human body area feature map, the receptive field of the convolution kernel is dynamically adjusted according to the scale size and position information of the interaction object in the human body area feature map; as the human body scale increases, the receptive field is increased.

[0091] The adaptive receptive field adjustment mechanism described in this specification dynamically adjusts the receptive field of the convolution kernel according to the actual situation of the human body in the input image (such as the size and position of the human body). If the human body occupies a large area in the image, the receptive field will increase accordingly to cover the entire human body; if the human body is small or at the edge of the image, the receptive field will decrease to ensure that the human body area can be accurately captured.

[0092] The detection frame generation module is used to perform a regression operation on the human body region feature map to output a detection frame representing the human body boundary, and is also used to analyze the features in the detection frame through convolution and fully connected layers to output the human body features in the detection frame;

[0093] Based on the feature map output by the human body region focusing module, a detection frame generation module is further constructed. This module contains multiple branches, one of which learns the coordinate information of the human body's bounding box through a regression layer. The regression layer is trained using a loss function based on a distance metric. The training can be performed using, but is not limited to, a Smooth L1 loss function that minimizes the error between the predicted detection frame and the real human body detection frame, that is, when the error between the predicted detection frame and the real human body detection frame is large, the position and size of the predicted frame can be quickly adjusted so that the predicted frame gradually approaches the real frame.

[0094] Another branch of the human body area focusing module further refines and analyzes the features within the detection frame. Through a combination of a series of convolutional and fully connected layers, it extracts more detailed feature information of the human body in the detection frame for subsequent key point detection tasks.

[0095] The key point detection module is used to perform convolution and deconvolution operations on the human body features in the detection frame to extract the features of the human body, and convert the feature map obtained after the feature extraction into a heat map to determine the key point information of the human body.

[0096] This module uses the human features in the detection frame output by the detection frame generation module as input, and detects the key points of the human body through a specially designed key point prediction subnetwork. The key point prediction subnetwork consists of multiple convolutional layers and deconvolutional layers. The convolutional layers are used to extract higher-level feature representations, while the deconvolutional layers are used to upsample and restore the resolution of the feature map to accurately locate the positions of the key points of the human body. Specifically, first, higher-level feature representations are extracted through multiple convolutional layers. These convolutional layers can further explore the key features of the human body, such as features around joints. Then, the deconvolutional layer is used to upsample and restore the resolution of the feature map, because in the previous processing process, the feature map undergoes multiple convolution and pooling operations, and the resolution will be reduced. The deconvolutional layer can restore the resolution of the feature map to a level close to that of the original image, so as to more accurately locate the positions of the key points of the human body.

[0097] In the above process, a heatmap generation mechanism is adopted to enable the heatmap to more accurately correspond to the positions of human key points in the original image.

[0098] As an example, the conversion of the feature map obtained after feature extraction into a heatmap to determine the information of human key points includes:

[0099] Convert the feature map obtained after feature extraction into a heatmap, where each pixel value in the heatmap represents the probability that the position is a key point;

[0100] Perform threshold processing and non-maximum suppression operations on the heatmap to determine the human key points;

[0101] Convert the coordinates of the human key points in the heatmap to determine the coordinate information of the human key points in the image where the detection box is located as the human key point information

[0102] Convert the position information of the key points into the form of a heatmap, where the peak position in the heatmap corresponds to the position of the human key points. By performing operations such as threshold processing and non-maximum suppression on the heatmap, the coordinate information of the human key points can be accurately extracted. For example, for joint points such as the shoulders, elbows, and wrists of the human body, corresponding peaks will be formed in the heatmap, and through operations such as threshold processing and non-maximum suppression, the coordinate information of these key points can be accurately extracted, such as the coordinates of the shoulder key point being (x1, y1), the coordinates of the elbow key point being (x2, y2), etc.

[0103] In the network structure of the above entire human detection model, the output features of the human region focusing module will be fused with some features of the basic feature extraction module, combining the global information in the basic features with the local human features extracted by the human region focusing module to supplement more comprehensive human-related information. At the same time, there will also be feature interaction and fusion between the detection box generation module and the key point detection module. The information of the detection box can assist in the detection of key points. For example, the position and size of the detection box can provide a reference for the search range of key points; the information of key points can also be fed back to the detection box for further optimization and adjustment, such as more accurately adjusting the boundary of the detection box according to the position and distribution of key points.

[0104] Through the above embodiments, the network structure of the human body detection model in this specification can effectively output accurate human body detection boxes and human key point information from the input image through the above unique design and innovative mechanism, with high accuracy and practicality, providing an effective technical solution for related fields of human body analysis. Among them, an adaptive learning mechanism is adopted in multiple modules. The adaptive receptive field adjustment mechanism of the human body area focusing module can automatically adjust the receptive field of the convolution kernel according to different input images, improving the adaptability of the network to different human postures and positions. The heat map generation mechanism in the key point detection module also adaptively adjusts the threshold and parameters of the heat map according to the training data to improve the accuracy of key point detection.

[0105] In step S220, the coordinate information of the human key points is converted into a feature representation suitable for the input of the convolutional neural network. After the human key point information is subjected to feature extraction, it is input into the trained convolutional neural network model, and the model will output the prediction result of the interaction behavior. For example, by analyzing the position and movement trajectory of the key points of the player's hand, it is determined whether the player is waving, making a fist or doing other specific actions, so as to determine the player's interaction behavior.

[0106] It can be understood that when using the convolutional neural network of deep learning to classify and identify the behavior of the interaction object, the prediction result is a probability distribution, indicating the possibility of each interaction behavior.

[0107] Referring to Fig. 4(a), by detecting the human key points in the real scene image, the interaction behavior of the interaction object is determined. For example, when the user moves left and right - the tracking box moves left and right; when the user moves to area B - the bottom edge of the tracking box moves to area B; when the user jumps in place - the center of the tracking box moves up by > 5 cm.

[0108] Subsequently, according to the interaction behavior of the interaction object and the interaction rules in the game library, the operation instructions in the game are determined. The operation instructions are used to control the projection screen of the projection unit in the interaction area subsequently, and this content will be described in detail in the subsequent chapters.

[0109] Through this embodiment, through the real-time gesture recognition technology, the user is allowed to directly control the vehicle lamp entertainment system through gestures, providing a more intuitive and natural interaction method than traditional buttons or touch screens.

[0110] Step S300, convert the interaction behavior into operation instructions in the game.

[0111] The game animation library includes various game animations, visual effects, and game rules. Based on the processing results of the image algorithm, DLP can project corresponding images (such as game images / animations) in the ground projection game library, and the projected images have a mapping / corresponding relationship with the recognized actions of the user. Therefore, after the behavior of the interaction object is recognized through the above visual algorithm, according to the preset mapping rules, it is converted into operation instructions that the game can recognize and execute. These instructions will be sent to the game engine to control the game process.

[0112] As an example, the conversion of the interaction behavior into operation instructions in the game includes:

[0113] Match the interaction behavior with the pre-stored operation behavior information of user game interactions in the game animation library to obtain the target operation behavior;

[0114] Retrieve the pre-stored images in the game animation library according to the target operation behavior; the images include game animations and result feedback images;

[0115] Generate corresponding operation instructions based on the target operation behavior, and the operation instructions include the images corresponding to the target operation behavior.

[0116] After obtaining the determined interaction behavior, it will be compared with the pre-stored user game interaction operation behavior information in the game animation library. The game animation library stores a large amount of pre-set standard behavior information corresponding to various game operations. Continuing to refer to Figure 4(b), when the interaction object walks into the specified area and starts the game, the squares in the area start to fall. When the interaction object jumps into the area where the squares are falling, the speaker plays "perfect", the corresponding squares disappear, and the projection refers to the image in Figure 4(c).

[0117] Specifically, the actual interaction behavior of the detected interaction object (such as the player's body posture and movement data obtained through human key point detection) is matched with the jumping action characteristics in the game library. By calculating the similarity between the two, the most matching pre-stored operation behavior, that is, the target operation behavior, is found. Once the target operation behavior is determined, the system will retrieve the corresponding pre-stored images from the game animation library. The game animation library not only contains the animations when the player performs operations, but also includes result feedback images. Continuing to refer to Figure 4, in the above jumping and eliminating squares game, when the interaction object jumps (target operation behavior), the animation of the interaction object jumping will be retrieved to show the action of the character leaping; if the squares are successfully eliminated by jumping, result feedback images such as the disappearance of the squares and the increase in score will also be retrieved. These images are pre-produced and stored in the library, and they are saved in the form of image sequences, animation scripts, etc., and are quickly retrieved and called when needed.

[0118] Finally, based on the determined target operation behavior, a corresponding operation instruction is generated. This operation instruction not only contains the identifier of the target operation behavior but also associates the corresponding screen information.

[0119] In this embodiment, the vehicle-mounted camera captures the user's gestures, which are analyzed and recognized using vision algorithms. Advanced computer vision technologies, deep learning, and image processing algorithms are adopted to real-time identify the user's key points and body movements, and the recognition results are converted into headlight control instructions, thereby enabling the user to interact with the headlight entertainment system through gestures. It provides a more intuitive and natural interaction method than traditional buttons or touchscreens.

[0120] In some embodiments, the method further includes:

[0121] Collect the voice information of the interaction object;

[0122] Combine the voice information and the interaction behavior to generate operation instructions in the game.

[0123] During the interaction process, a speech recognition process is also involved. It can be understood that a variety of sensors and input devices such as integrated microphones and cameras are supported, and various multimodal interaction and feedback methods such as gestures and voices are supported, and feedback is given to the user through the headlight changes.

[0124] In this embodiment, a variety of interaction methods such as key point and gesture recognition, touch screen operation, and voice control are combined, providing users with a more diverse control selection, and increasing the usability and convenience of the system.

[0125] Step S400, control the projection headlight to project the screen indicated by the operation instruction in the interaction area.

[0126] The game system generates corresponding game screens in real time according to the received operation instruction. These screens will change dynamically according to the logic and rules of the game, including result feedback screens. The generated game screen data is transmitted to the control system of the projection headlight, and the projection headlight will project the corresponding screen in the interaction area in front of the vehicle, enabling the interaction object to see the feedback of their operations in the game in real time, thereby realizing the interaction with the game.

[0127] In some embodiments, when the game system receives this operation instruction, it can play the corresponding game animation in the game scene displayed on the in-vehicle display unit according to the information therein, display the presentation and feedback of the real-time picture of the interaction object in the game, and realize the smooth interaction and dynamic update of the game.

[0128] In this embodiment, an infrared sensor and a front-view camera are used to capture user actions and environmental images. After being processed by an image algorithm, the DLP projection headlight is controlled to project images or animations in the interaction area. The user interacts with the game interface within the game activity range. The game animation library and rule library cooperate, and the data processing module ensures the smooth operation of the game, forming a closed-loop interaction experience.

[0129] This application provides a method, device, and computer-readable storage medium for headlight entertainment interaction. Compared with the current situation where immediate interaction between the user and the headlight cannot be achieved, by obtaining the real-scene image of the interaction area projected by the projection headlight in front of the vehicle, detecting the human key points in the behavior performance of the interaction object participating in the game indicated by the real-scene image in the interaction area, analyzing and identifying through a vision algorithm, and determining the interaction behavior of the interaction object. The interaction behavior is converted into an operation instruction in the game, and the projection headlight is controlled to project the picture indicated by the operation instruction in the interaction area, so that the user can directly control the headlight entertainment system through gestures, greatly enriching and enhancing the user's driving experience.

[0130] In the in-vehicle application scenario, in order to enable the deep learning model to run efficiently on the neural processing unit (NPU) supported by the Qualcomm Neural Processing SDK, based on the above first embodiment, this specification proposes a set of methods for model compression and deployment of a complete headlight entertainment interaction program.

[0131] This method involves the following key steps:

[0132] Quantization-Aware Training (QAT): Deep learning models usually use floating-point numbers to store parameters, which has relatively high costs in both calculation and storage. The core of QAT technology is to simulate the conversion of model parameters from floating-point numbers to low-precision integers during the training process. In this step, the quantization-aware training technology is adopted. During the training process, the quantization parameters required for the conversion of deep learning model parameters from floating-point numbers to low-precision integers are accurately captured. This not only maintains the prediction accuracy of the model but also significantly reduces the calculation load and memory requirements of the model.

[0133] Model conversion: Different hardware platforms have different requirements for the file format of the model, and in-vehicle NPUs are no exception. Therefore, after completing the quantization-aware training, a model conversion tool of Qualcomm Neural Network (QNN) is used to convert the trained deep learning model into a specific file format suitable for the in-vehicle environment, ensuring that the model can run efficiently on the Qualcomm NPU.

[0134] Quantization Parameter Embedding: The quantization parameters are obtained during the quantization-aware training process, and these parameters determine the performance of the model in low-precision representation. Therefore, the quantization parameters are cleverly embedded into the json configuration file of the converted neural network model, so that when the model is deployed and run on the NPU, the quantization parameters are automatically applied, and the inference calculation is performed according to the preset quantization rules, further improving the inference efficiency of the model.

[0135] Specifically, by parsing the quantization parameters obtained from the quantization-aware training and using the automatically written network structure alignment tool, the network structure in the quantization-aware training is aligned with the network structure in the file after being converted into a QNN model, and the quantization parameters are overwritten.

[0136] Compilation: The json configuration file embedded with quantization parameters is then compiled into a binary format executable by QNN. Through compilation, the model is converted into an instruction set that can be directly executed by the NPU, providing the necessary format support for the efficient inference of the model on the NPU and optimizing the running efficiency of the model on the NPU.

[0137] Through this series of processes, this embodiment achieves a more than 10-fold improvement in inference performance on the NPU compared to the traditional CPU solution. This method not only improves the inference efficiency of the deep learning model on the NPU but also maintains the prediction accuracy of the model. This significant performance leap makes the method of this embodiment particularly suitable for in-vehicle application scenarios with extremely high real-time requirements.

[0138] Figure 5 An example of the physical structure diagram of a vehicle lamp entertainment interaction device is shown as Figure 5 As shown, the vehicle lamp entertainment interaction device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the vehicle lamp entertainment interaction method.

[0139] In addition, when the logical instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0140] On the other hand, this application also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the headlight entertainment interaction method provided by the above-mentioned various methods.

[0141] On another aspect, this application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the headlight entertainment interaction method provided by the above-mentioned various methods.

[0142] It should be noted that the technical solutions or technical features described in the above embodiments can be combined or supplemented with each other without conflict. The scope of protection of this application is not limited to the precise structures described in the above embodiments and shown in the drawings; all modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the scope of protection of this application.

Claims

1. A method for vehicle lamp entertainment interaction, characterized in that, The method includes: Obtaining a real - scene image of an interaction area projected by a projection headlight in front of a vehicle, where the real - scene image indicates the behavior performance of an interaction object participating in a game in the interaction area; Detecting human key points in the real - scene image to determine the interaction behavior of the interaction object; Converting the interaction behavior into an operation instruction in the game; Controlling the projection headlight to project the picture indicated by the operation instruction in the interaction area.

2. The vehicle - lamp entertainment interaction method according to claim 1, wherein The detecting human key points in the real - scene image to determine the interaction behavior of the interaction object includes: Detecting human key points in the real - scene image to determine human key - point information; Based on the human key - point information, determining the interaction behavior of the interaction object through a convolutional neural network of deep learning.

3. The headlight entertainment interaction method according to claim 2, characterized in that, Using a human - detection model to detect human key points in the real - scene image, where the human - detection model includes a basic feature extraction module, a human - area focusing module, a detection - box generation module, and a key - point detection module; The basic feature extraction module is used to extract features from the real - scene image and output a basic feature map; The human - area focusing module is used to perform convolution and pooling operations on the basic feature map, focus on the features of the human area where the interaction object is located in the real - scene image, and output a human - area feature map; The detection - box generation module is used to perform a regression operation on the human - area feature map to output a detection box representing the human boundary, and is also used to analyze the features within the detection box through convolution and fully - connected layers, and output the human features within the detection box; The key - point detection module is used to perform convolution and de - convolution operations on the human features within the detection box to extract human features, and convert the feature map obtained after extracting the features into a heat map to determine human key - point information.

4. The vehicle - lamp entertainment interaction method according to claim 3, wherein The converting the feature map obtained after extracting features into a heat map to determine the information of human key points includes: Converting the feature map obtained after extracting features into a heat map, where each pixel value in the heat map represents the probability that the position is a key point; Performing threshold processing and non - maximum suppression operations on the heat map to determine the human key points; Converting the coordinates of the human key points in the heat map to determine the coordinate information of the human key points in the image where the detection box is located as the human key - point information.

5. The vehicle - lamp entertainment interaction method according to claim 3, wherein During the process that the human - area focusing module is used to perform convolution and pooling operations on the basic feature map, focus on the features of the human area in the real - scene image, and output a human - area feature map, the receptive field of the convolution kernel is dynamically adjusted according to the scale size and position information of the interaction object in the human - area feature map; As the scale of the human body increases, the receptive field is increased.

6. The vehicle - lamp entertainment interaction method according to claim 3, wherein The real - scene image includes an infrared image and an RGB image.

7. The vehicle - lamp entertainment interaction method according to claim 6, wherein The basic feature extraction module is used to extract features from the real-scene image and output a basic feature map, including: The basic feature extraction module includes a dual-stream network, and the dual-stream network includes a residual network. The convolutional layer of the residual network performs a convolutional operation on the RGB image to extract the feature map of the RGB image; The convolutional layer of the residual network is used to extract features from the infrared image to obtain the feature map of the infrared image; The feature map of the RGB image and the feature map of the infrared image are superimposed and fused to obtain the basic feature map.

8. The vehicle lamp entertainment interaction method according to any one of claims 1, wherein The step of converting the interaction behavior into an operation instruction in the game includes: Matching the interaction behavior with the operation behavior information of the user game interaction pre-stored in the game animation library to obtain a target operation behavior; Retrieving the pre-stored pictures in the game animation library according to the target operation behavior; the pictures include game animations and result feedback pictures; Generating a corresponding operation instruction based on the target operation behavior, and the operation instruction includes the picture corresponding to the target operation behavior.

9. The vehicle lamp entertainment interaction method according to claim 1, wherein The method further includes: Collecting the voice information of the interaction object; Combining the voice information and the interaction behavior to generate an operation instruction in the game.

10. The vehicle lamp entertainment interaction method according to claim 1, wherein Before obtaining the real-scene image of the interaction area projected by the projection headlamp in front of the vehicle, the method further includes: Obtaining the state data of the vehicle; Combining preset game trigger conditions to determine whether to activate the entertainment function of the vehicle; the game trigger conditions include that the state data meets the set safety threshold; After activating the entertainment function of the vehicle, execute obtaining the real-scene image of the interaction area projected by the projection headlamp in front of the vehicle.

11. A vehicle lamp entertainment interaction device, characterized in that, It includes a memory, a processor, and a vehicle lamp entertainment interaction program stored on the memory and executable on the processor. When the processor executes the vehicle lamp entertainment interaction program, it implements the steps of the vehicle lamp entertainment interaction method according to any one of claims 1-10.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a vehicle lamp entertainment interaction program, and when the vehicle lamp entertainment interaction program is executed, it implements the steps of the vehicle lamp entertainment interaction method according to any one of claims 1-10.

Citation Information

Cited By

  • Human-computer interaction control method and device, electronic equipment and storage medium

    CN120469602A

  • Game interaction control method and device, computer equipment and storage medium

    CN120550406A

  • In-vehicle light control method and device, vehicle end control equipment and readable storage medium

    CN121510424A