Interaction early warning method and system of intelligent cockpit and vehicle
Patent Information
- Application Number
- CN202610711959.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]然而,现有技术的座舱图像采集装置大多仅限于基础的画面传输,未能对乘员的生理特征与行为状态进行深度的特征挖掘与分析
[0071] This application provides an interactive early warning method, system, and vehicle for a smart cockpit. It acquires image data of the target occupant in real time using an image acquisition device and simultaneously extracts the occupant's physiological characteristics and behavioral status based on the image data. This allows for the full utilization of cockpit image data while simultaneously obtaining the individual characteristics and riding behavior of the target occupant. By generating or updating a virtual assistant avatar corresponding to the target occupant based on physiological characteristic information, the virtual assistant avatar can be matched to the target occupant's age, physical appearance, or growth changes, thereby improving the personalization and emotional companionship effect of the smart cockpit interaction. By determining the target occupant's behavioral risk level based on behavioral status information, different riding behaviors of the target occupant can be graded, thereby improving the precision of safety monitoring and the accuracy of risk identification. By outputting corresponding interactive information and/or early warning information based on the virtual assistant avatar and behavioral risk level, the interactive content, reminder methods, and warning intensity can be adapted to the actual state of the target occupant, thereby enhancing the smart cockpit's personalized interaction capabilities, proactive safety early warning capabilities, and riding safety assurance effect.
Smart Images

Figure CN122598397A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle technology, and in particular to an interactive warning method, interactive warning system and vehicle for an intelligent cockpit. Background Technology
[0002] With the rapid development of intelligent connected vehicles, occupant safety monitoring and human-machine interaction technologies in smart cockpits have been widely applied, playing an irreplaceable role, especially in scenarios involving the supervision of rear-seat child occupants. Current in-vehicle rear-seat monitoring systems are typically equipped with image acquisition devices such as cameras, allowing drivers to view the rear-seat occupants' seating positions in real time, thereby enhancing driving safety.
[0003] However, most existing cockpit image acquisition devices are limited to basic image transmission and fail to perform in-depth feature mining and analysis of occupants' physiological characteristics and behavioral states. Meanwhile, existing cockpit interaction systems lack specificity, failing to generate and dynamically update personalized virtual interactive avatars based on occupants' physiological characteristics, thus making it difficult to provide an interactive experience with emotional companionship attributes. Furthermore, traditional safety monitoring systems often employ a single alarm mechanism, failing to conduct refined risk level assessments based on occupants' specific behavioral states, resulting in simplistic and untargeted warning and interaction methods.
[0004] In summary, existing technologies cannot fully utilize cockpit image data to achieve in-depth physiological perception and behavioral risk classification, resulting in a lack of personalization in the cockpit interaction experience and insufficient accuracy in safety warnings. Summary of the Invention
[0005] In view of this, the purpose of this application is to provide an interactive early warning method, interactive early warning system and vehicle for a smart cockpit. By collecting image data of the target occupant in real time, and extracting the physiological characteristic information and behavioral state information of the target occupant based on the image data, the system can generate or update a virtual assistant image corresponding to the target occupant using the physiological characteristic information, and determine the behavioral risk level of the target occupant using the behavioral state information. This enables the smart cockpit to output corresponding interactive information and / or early warning information according to the individual characteristics and behavioral risk state of the target occupant, thereby improving the personalization of cockpit interaction and the accuracy of safety early warning.
[0006] In a first aspect, the present invention provides an interactive early warning method for a smart cockpit, comprising: Image data of the target occupants is acquired in real time using an image acquisition device.
[0007] Physiological characteristics and behavioral status information of the target occupants are extracted based on image data.
[0008] Based on physiological characteristics, generate or update the virtual assistant image corresponding to the target occupant.
[0009] Based on behavioral status information, the behavioral risk level of the target occupant is determined.
[0010] Based on the virtual assistant's appearance and behavioral risk level, it outputs interactive information and / or warning information corresponding to the target occupant.
[0011] In an optional implementation, the step of extracting physiological characteristic information and behavioral state information of the target occupant based on image data includes: The occupant image region in the image data is identified to obtain the facial features, body proportion features, and posture features of the target occupant.
[0012] Physiological characteristics are determined based on facial features and body proportions.
[0013] Behavioral state information is determined based on posture features and the positional relationship between the target occupant and the target components inside the vehicle.
[0014] In an optional implementation, the in-vehicle target components include at least one of seat belts, windows, doors, and child safety seats.
[0015] When the target occupant is a child occupant, the steps for determining behavioral state information based on posture characteristics and the positional relationship between the target occupant and target components inside the vehicle include: The system identifies key human body points, limb orientation, and positional changes of key human body points relative to target components within the vehicle for child occupants; key human body points include at least one of the hands, head, and torso.
[0016] Based on information about changes in limb orientation and position, determine whether the child occupant exhibits at least one of the following behavioral states: signs of seatbelt release, signs of leaning out of the window, signs of leaning forward, signs of tilting, and signs of leaving the child seat.
[0017] In an optional implementation, the step of generating a virtual assistant image corresponding to the target occupant based on physiological characteristic information includes: Obtain the basic configuration information corresponding to the target occupant; wherein, the basic configuration information includes at least one of age information, gender information, and image preference information.
[0018] The physical appearance parameters of the target occupant are determined based on physiological characteristic information; wherein, the physical appearance parameters include at least one of facial structure parameters, head-to-body ratio parameters, height ratio parameters, skin color parameters, hairstyle parameters, and clothing adaptation parameters.
[0019] Based on basic configuration information and physical parameters, an initial virtual assistant image corresponding to the target occupant is generated.
[0020] The initial virtual assistant avatar is associated with and stored with the identity identifier of the target passenger.
[0021] In an optional implementation, the step of updating the virtual assistant image corresponding to the target occupant based on physiological characteristic information includes: The physiological characteristic information of the target occupant is reacquired according to the preset update cycle or when changes in physiological characteristics are detected to exceed a preset threshold.
[0022] By comparing the newly acquired physiological characteristic information with historical physiological characteristic information, growth and change information is obtained.
[0023] If the growth and change information meets the preset growth and change conditions, update the physical parameters of the virtual assistant's image based on the growth and change information.
[0024] If the growth and change information does not meet the preset growth and change conditions, the physical parameters of the virtual assistant's image will remain unchanged.
[0025] In an optional implementation, the step of determining the behavioral risk level of a target occupant based on behavioral state information includes: Acquire consecutive multi-frame image data to determine behavioral state information.
[0026] Input multiple consecutive frames of image data into a pre-trained behavior recognition model.
[0027] Spatial features are extracted from each frame of image data using a behavior recognition model to obtain a spatial feature vector corresponding to each frame of image data. The spatial feature vector is used to characterize the posture features of the target occupant and its positional relationship with the target components inside the vehicle.
[0028] Multiple spatial feature vectors are combined according to the acquisition time sequence to obtain a temporal feature sequence.
[0029] Temporal change analysis is performed on the temporal feature sequence to obtain the behavior recognition results corresponding to the target occupant; the behavior recognition results include behavior category, behavior confidence level and behavior development trend.
[0030] The behavioral risk level is determined based on the behavioral category, behavioral confidence level, and behavioral development trend.
[0031] In an optional implementation, the step of determining the behavioral risk level based on behavioral category, behavioral confidence level, and behavioral development trend includes: Obtain the risk assessment criteria corresponding to the behavior category.
[0032] Determine whether the confidence level of the behavior meets the confidence level requirements corresponding to the risk assessment conditions.
[0033] If the confidence level of the behavior meets the confidence level requirements, determine whether the trend of the behavior is pointing to a pre-set dangerous component or a pre-set dangerous area.
[0034] When the trend of behavior points to a pre-set dangerous component or pre-set dangerous area, the behavior category is determined to be a valid risk behavior.
[0035] The risk level of a behavior is determined based on the degree of risk corresponding to an effective risky behavior.
[0036] In an optional implementation, the behavior category is identified as a valid risky behavior, including: Within a preset time window, consistency is determined for the behavior categories corresponding to multiple frames of image data.
[0037] If multiple consecutive frames of image data within a preset time window correspond to the same behavior category, and the confidence levels of the behaviors corresponding to the same behavior category all meet the confidence level requirements, then the same behavior category will be identified as a candidate risk behavior.
[0038] Short-term interference filtering is applied to candidate risky behaviors.
[0039] If a candidate risky behavior is not identified as a short-term disruptive behavior, the candidate risky behavior will be identified as a valid risky behavior.
[0040] In an optional implementation, when the target occupant is a child occupant, after the step of obtaining the risk assessment conditions corresponding to the behavior category, the method further includes: Obtain historical behavioral data and historical early warning data for child occupants.
[0041] Based on historical behavioral data and historical early warning data, individualized risk assessment parameters are determined for child occupants. These individualized risk assessment parameters are used to characterize the risk assessment standards for child occupants at different stages of travel, at different ages, and with different historical behavioral habits.
[0042] Based on individualized risk assessment parameters, the risk assessment conditions, confidence requirements, preset hazardous components and / or preset hazardous areas corresponding to the behavior category are adjusted.
[0043] Based on the adjusted risk assessment criteria, behavior category, behavior confidence level, and behavior development trend, the behavioral risk level of child occupants is determined.
[0044] In an optional implementation, the step of determining individualized risk assessment parameters for child occupants based on historical behavioral data and historical warning data includes: The study analyzed the routine and risky movement patterns of child occupants during their historical rides.
[0045] The range of typical behaviors for child occupants is determined based on their regular movement trajectories.
[0046] The range of risky behaviors corresponding to child occupants is determined based on historical risky action trajectories.
[0047] Based on the difference between the scope of routine behavior and the scope of risky behavior, individualized risk assessment parameters are generated.
[0048] Upon receiving feedback from the guardian regarding the warning information, the individualized risk assessment parameters are updated based on the feedback.
[0049] In an optional implementation, the step of outputting interaction information corresponding to the target occupant based on the virtual assistant's image and behavioral risk level includes: Based on behavioral state information, the target interaction scenario corresponding to the target occupant is determined.
[0050] Based on the target interaction scenario and physiological characteristics, determine the interaction content that matches the target occupant.
[0051] Control the virtual assistant avatar to output corresponding voice interaction information and / or visual interaction information according to the interaction content.
[0052] While outputting voice interaction information and / or visual interaction information, continue to acquire the target occupant's facial expression status information.
[0053] Adjust interactive content based on facial expression status information.
[0054] In an optional implementation, the step of adjusting the interactive content based on facial expression state information includes: If the facial expression indicates that the target passenger is crying, the interaction content will be adjusted to a calming or soothing manner.
[0055] When facial expression information indicates that the target occupant is in a pleasant state, maintain the current interaction content or output continued interaction content related to the current interaction content.
[0056] When the behavioral status information indicates that the target occupant has a demand expression action, the demand prompt information is determined based on the demand expression action and sent to the vehicle display device.
[0057] In an optional implementation, the step of outputting warning information corresponding to the target occupant based on the virtual assistant's image and behavioral risk level includes: When the behavioral risk level is low, control the virtual assistant avatar to output the first warning information.
[0058] When the behavioral risk level is medium risk, control the virtual assistant to output a second warning message and control the in-vehicle prompting device to output auxiliary prompt messages.
[0059] When the behavioral risk level is high, the system extracts the specific dangerous action characteristics corresponding to the behavioral status information, controls the virtual assistant image to present a warning posture corresponding to the specific dangerous action characteristics, outputs a third warning message, and controls the in-vehicle alarm device to output a danger warning message to the driver.
[0060] In an optional implementation, after the step of controlling the in-vehicle alarm device to output a hazard warning message to the driver, the method further includes: Obtain the behavioral status information of the target occupant and the status change results after the output of the danger warning information.
[0061] When the state change result indicates a decrease in the behavioral risk level of the target occupant, the output intensity of the warning information should be reduced or the warning information should be stopped.
[0062] If the risk level of the target occupant's behavior as a result of the state change has not decreased, the output intensity of the warning information should be increased, and the vehicle should be controlled to perform auxiliary safety control operations.
[0063] Secondly, the present invention provides an interactive early warning system for an intelligent cockpit, comprising: The image data acquisition module is used to acquire image data of the target occupants in real time through an image acquisition device.
[0064] The data processing module is used to extract physiological characteristics and behavioral status information of the target occupants based on image data.
[0065] The virtual assistant generation module is used to generate or update the virtual assistant image corresponding to the target occupant based on physiological characteristic information.
[0066] The risk assessment module is used to determine the behavioral risk level of a target occupant based on behavioral status information.
[0067] The interactive warning module is used to output interactive information and / or warning information corresponding to the target occupant based on the virtual assistant's image and behavioral risk level.
[0068] Thirdly, the present invention provides a vehicle including an image acquisition device, an in-vehicle central control system, and an interactive early warning system for an intelligent cockpit as described in the foregoing embodiments.
[0069] The intelligent cockpit's interactive warning system is integrated into the vehicle's central control system.
[0070] The image acquisition device is connected to the vehicle's central control system and includes a rear-seat monitoring camera located on the top of the rear seat or on the B-pillar of the vehicle.
[0071] This application provides an interactive early warning method, system, and vehicle for a smart cockpit. It acquires image data of the target occupant in real time using an image acquisition device and simultaneously extracts the occupant's physiological characteristics and behavioral status based on the image data. This allows for the full utilization of cockpit image data while simultaneously obtaining the individual characteristics and riding behavior of the target occupant. By generating or updating a virtual assistant avatar corresponding to the target occupant based on physiological characteristic information, the virtual assistant avatar can be matched to the target occupant's age, physical appearance, or growth changes, thereby improving the personalization and emotional companionship effect of the smart cockpit interaction. By determining the target occupant's behavioral risk level based on behavioral status information, different riding behaviors of the target occupant can be graded, thereby improving the precision of safety monitoring and the accuracy of risk identification. By outputting corresponding interactive information and / or early warning information based on the virtual assistant avatar and behavioral risk level, the interactive content, reminder methods, and warning intensity can be adapted to the actual state of the target occupant, thereby enhancing the smart cockpit's personalized interaction capabilities, proactive safety early warning capabilities, and riding safety assurance effect.
[0072] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. The objectives and other advantages of this application are realized and obtained through the structures particularly pointed out in the description, claims and drawings.
[0073] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0074] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0075] Figure 1 This is a schematic diagram of a vehicle provided in an embodiment of this application; Figure 2 A flowchart of the interactive early warning method for a smart cockpit provided in this application embodiment; Figure 3 Flowchart of the method for extracting physiological feature information and behavioral state information provided in the embodiments of this application; Figure 4 Flowchart of the virtual assistant avatar generation method provided in the embodiments of this application; Figure 5 Flowchart of the virtual assistant avatar update method provided in the embodiments of this application; Figure 6 A schematic diagram of an interactive early warning system for an intelligent cockpit provided in an embodiment of this application.
[0076] Icons: 11-Image acquisition device; 12-Vehicle central control system; 13-Interactive early warning system for smart cockpit; 21-Image data acquisition module; 22-Data processing module; 23-Virtual assistant generation module; 24-Risk assessment module; 25-Interactive early warning module. Detailed Implementation
[0077] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0078] To facilitate a better understanding of this application by those skilled in the art, a brief introduction to the application scenarios and design concepts of this application is provided.
[0079] Vehicles typically use rear-seat cameras to capture images of rear-seat occupants and display the images on an in-vehicle screen so that the driver or caregiver can monitor their seating situation. However, such solutions usually only use the image acquisition device as a video transmission device, failing to extract physiological and behavioral information of the target occupants from the image data. This results in low utilization of in-vehicle image data and makes it difficult to comprehensively analyze the target occupants' growth changes, posture, body movements, and positional relationships with target components within the vehicle.
[0080] Meanwhile, existing virtual assistants or voice interaction functions in smart cockpits typically use fixed images and preset interaction content. This makes it difficult to generate a virtual assistant image that corresponds to the target occupant's age, facial features, body proportions, and other physiological characteristics. Furthermore, it's difficult to dynamically update the virtual assistant image as the target occupant grows and changes. Therefore, the existing interaction methods have a low degree of matching with the target occupant's own characteristics, failing to create a companionable and personalized interactive experience.
[0081] Furthermore, existing safety monitoring methods largely rely on single alarm rules, issuing alerts directly upon detecting abnormal states. They typically fail to categorize risk levels based on the target occupant's behavioral status information, nor do they output interactive information and / or warnings of varying intensities and forms according to different behavioral risk levels. Therefore, existing technologies suffer from limitations such as simplistic warning methods, lack of targeted alerts, and insufficiently precise risk identification, making it difficult to intervene in dangerous behaviors of target occupants in a timely and accurate manner.
[0082] Based on this, this application provides an interactive early warning method, interactive early warning system, and vehicle for a smart cockpit. It acquires image data of the target occupant in real time using an image acquisition device, and extracts physiological characteristic information and behavioral state information of the target occupant based on the image data. This allows the cockpit image data to be used not only for screen display but also for individual characteristic recognition and riding behavior analysis of the target occupant. Furthermore, this application generates or updates a virtual assistant image corresponding to the target occupant based on the physiological characteristic information, ensuring that the virtual assistant image matches the target occupant's age, physical appearance, or growth changes, thereby enhancing the personalization of the interactive experience. Further, this application determines the behavioral risk level of the target occupant based on behavioral state information and outputs corresponding interactive information and / or early warning information based on the virtual assistant image and behavioral risk level. This ensures that the interactive content, early warning method, and reminder intensity are adapted to the actual state of the target occupant, thereby improving the personalized interactive capabilities and proactive safety assurance capabilities of the smart cockpit. To facilitate understanding of this embodiment, the embodiments of this application will be described in detail below.
[0083] This application provides a vehicle, referring to... Figure 1 The vehicle provided in this application embodiment includes an image acquisition device 11, an in-vehicle central control system 12, and an interactive early warning system 13 for the smart cockpit.
[0084] The intelligent cockpit's interactive warning system 13 is integrated into the vehicle's central control system 12.
[0085] The image acquisition device 11 is communicatively connected to the vehicle central control system 12. The image acquisition device 11 includes a rear-seat monitoring camera installed on the top of the rear seat or the B-pillar of the vehicle.
[0086] Here, the intelligent cockpit's interactive early warning system 13 is used to perform physiological feature recognition, behavioral status recognition, virtual assistant image generation or update, behavioral risk level judgment, and generation of interactive information and / or early warning information based on the image data collected by the image acquisition device 11.
[0087] The rear-seat monitoring camera can be an RMS (Rear Monitoring System) camera. The camera's acquisition direction is towards the rear passenger area of the vehicle, ensuring coverage of child safety seats, rear seats, rear windows, rear doors, and the area where the target occupant is located. The rear-seat monitoring camera is used to acquire real-time images of the target occupant's face, body movements, posture, and the surrounding in-vehicle environment. Facial images can be used to identify the target occupant's facial features, expressions, and identity. Body movement images can be used to identify the target occupant's hand movements, head movements, torso movements, and body orientation. Posture images can be used to identify whether the target occupant is properly leaning against the seat, whether their body is leaning forward, whether their body is tilted, and whether they have moved out of the child safety seat. In-vehicle environment images can be used to identify the position and status of target components inside the vehicle, such as seat belts, windows, doors, and child safety seats.
[0088] In one optional embodiment, the rear-seat monitoring camera has high-definition imaging capabilities, with an image resolution of at least 1080P. The image acquisition frame rate of the rear-seat monitoring camera can be at least 30 frames per second, enabling the in-vehicle central control system 12 to analyze the target occupant's movement changes based on continuous multi-frame image data. The rear-seat monitoring camera can have a wide-angle acquisition range of 120° to 150° to expand the image coverage of the rear area of the vehicle and reduce blind spots caused by the target occupant's limb movements or obstructed interior components. The rear-seat monitoring camera can also have infrared night vision capabilities to continue acquiring image data of the target occupant at night, in tunnels, underground parking garages, or in situations with insufficient interior lighting, thereby ensuring that the intelligent cockpit's interactive warning system 13 can perform occupant monitoring and warning processing under different lighting conditions.
[0089] The in-vehicle central control system 12 is used to receive image data sent by the image acquisition device 11 and to run the interactive warning system 13 of the smart cockpit. The in-vehicle central control system 12 may include a processor, a memory, and a communication interface. The processor is used to execute the processing logic corresponding to the interactive warning system 13 of the smart cockpit, the memory is used to store the target occupant's historical physiological characteristics, historical behavioral data, virtual assistant image data, risk assessment parameters, and interactive content data, and the communication interface is used to communicate with the image acquisition device 11, the virtual assistant interaction device, and the warning device.
[0090] The virtual assistant interaction device is communicatively connected to the in-vehicle central control system 12. Under the control of the in-vehicle central control system 12, the virtual assistant interaction device displays a virtual assistant avatar and outputs interactive information corresponding to the target occupant. The virtual assistant avatar can be a VPA (Virtual Personal Assistant) avatar. The virtual assistant interaction device may include a voice output device and a display device. The voice output device may include a rear-seat dedicated speaker or in-vehicle audio system, and the display device may include a rear-seat display screen, a central control display screen, or a split-screen display area within the central control display screen. The in-vehicle central control system 12 can control the virtual assistant interaction device to output nursery rhymes, stories, soothing voices, request confirmation voices, or visual interactive content based on the target occupant's physiological characteristics and the interaction scenario.
[0091] The warning device is communicatively connected to the vehicle central control system 12. Under the control of the vehicle central control system 12, the warning device outputs warning information corresponding to the behavioral risk level. The warning device may include at least one of the following: vehicle audio system, interior lighting system, instrument panel display device, seat vibration device, and vehicle display device. The vehicle central control system 12 can control the virtual assistant interaction device to output reminder information when the behavioral risk level is low, control the interior lighting system or vehicle display device to output auxiliary prompt information when the behavioral risk level is medium, and control the vehicle audio system, seat vibration device, and vehicle display device to output danger warning information when the behavioral risk level is high.
[0092] In one specific embodiment, when the target occupant is a child, the rear-seat monitoring camera captures facial images, posture images, and body movement images of the child occupant and sends the captured image data to the in-vehicle central control system 12. The in-vehicle central control system 12 operates the intelligent cockpit interactive warning system 13, generating a virtual assistant image corresponding to the child occupant based on facial images and body proportion features, and controlling the rear-seat display screen to display the virtual assistant image. The in-vehicle central control system 12 can also control the rear-seat dedicated speakers to play appropriate nursery rhymes or stories based on the child occupant's age stage. When the rear-seat monitoring camera captures the child occupant's hand gradually approaching the seat belt buckle, the in-vehicle central control system 12 operates the intelligent cockpit interactive warning system 13, identifying the signs of seat belt release based on continuous multi-frame image data and determining the child occupant's behavioral risk level. The in-vehicle central control system 12 can control the rear-seat display screen to display the warning posture corresponding to the virtual assistant image, control the rear-seat dedicated speakers to output voice reminders, and simultaneously control the in-vehicle lighting devices to provide a prompt. If a child occupant continues to perform a dangerous action, the in-vehicle central control system 12 can control the seat vibration device to output a vibration alert, control the in-vehicle display device to display a danger warning text, and control the in-vehicle audio system to output a danger warning voice to the driver. The in-vehicle central control system 12 can also reduce the volume of the in-vehicle multimedia system to increase the probability that the driver receives the warning information.
[0093] Based on this, this application provides an interaction method for an intelligent cockpit, referring to... Figure 2 The interactive early warning method for a smart cockpit provided in this application includes: Step S101: Real-time image data of the target occupant is acquired using an image acquisition device.
[0094] Here, the image acquisition device can be installed in the rear roof, B-pillar, seat back, center console area, or any location within the vehicle that covers the area where the target occupant is seated. The image acquisition device can include an in-vehicle camera, a rear-seat monitoring camera, an infrared camera, a depth camera, or a multi-view camera. The rear-seat monitoring camera can be an RMS camera. The target occupant can be a child passenger, an elderly passenger, or an occupant requiring special care.
[0095] The image data includes facial images, full-body images, upper-body images, body movement images, and sitting posture images of the target occupant, as well as images of the vehicle's interior environment surrounding the target occupant. The interior environment images include images of seat belts, windows, doors, child safety seats, rear seats, in-vehicle displays, cup holders, or other interior components. By simultaneously acquiring images of the target occupant and the interior environment, it is possible to determine the target occupant's own condition and the relative relationship between the target occupant and interior components.
[0096] Real-time data acquisition can be continuous after the vehicle starts, or triggered by detections such as occupied rear seats, child safety seats being in use, closed rear doors, the vehicle entering motion, or a guardian activating the monitoring function. The image acquisition device can acquire multiple consecutive frames of image data at a preset frame rate to form a sequence of image frames arranged in chronological order. The image acquisition device can also automatically adjust exposure parameters, image gain, acquisition frame rate, or acquisition mode based on in-vehicle lighting conditions, and switch to infrared acquisition mode at night, in tunnels, underground parking garages, or in situations with insufficient in-vehicle lighting to ensure that the image data meets the requirements for feature extraction and behavior recognition.
[0097] Step S102: Extract physiological feature information and behavioral status information of the target occupant based on image data.
[0098] Here, after obtaining the image data, image preprocessing is performed. Image preprocessing includes at least one of image denoising, brightness compensation, distortion correction, image enhancement, image cropping, image frame selection, image alignment, and target region extraction. Through image preprocessing, the impact of changes in in-vehicle lighting, vehicle vibration, partial occupant occupancy, or image noise on the feature extraction results can be reduced.
[0099] After image preprocessing, occupant image regions are identified from the image data. These regions are obtained through face detection, human detection, human keypoint detection, semantic segmentation, object detection, or background subtraction. Once the occupant image regions are identified, facial features, body proportions, posture features, and expression features of the target occupant are extracted. Facial features include face shape, facial proportions, eye shape, hairstyle, skin color, and facial contours. Body proportions include height ratio, head-to-body ratio, torso ratio, limb proportions, and body contours. Posture features include sitting angle, body tilt, head orientation, torso direction, hand position, arm extension direction, and leg posture.
[0100] Physiological characteristic information is used to characterize the appearance, growth status, or individual condition of a target occupant. Physiological characteristic information includes at least one of the following: facial features, body proportions, height estimation, body shape, age group, skin color, hairstyle, and facial expression. For example, the height variation of a target occupant can be estimated based on a seat reference, child safety seat markings, or dimensions or depth images of fixed components inside the vehicle. Alternatively, the target occupant's facial features and body proportions can be used to determine whether the target occupant is an infant, a child, or another age group.
[0101] Behavioral state information is used to characterize the target occupant's actions, postures, behavioral trends, and positional relationships between the target occupant and target components within the vehicle during the riding process. Behavioral state information includes at least one of the following: sitting posture, limb movement state, key body point positions, limb orientation, head orientation, torso tilt direction, hand movement trajectory, positional relationships between the target occupant and the seatbelt, positional relationships between the target occupant and the vehicle window, positional relationships between the target occupant and the vehicle door, and positional relationships between the target occupant and the child safety seat. For example, when the target occupant's hand key points continuously approach the seatbelt buckle, behavioral state information corresponding to the premonition of seatbelt release is extracted. When the target occupant's head and torso key points continuously move towards the vehicle window, behavioral state information corresponding to the premonition of leaning out of the window is extracted. When the target occupant's torso angle exceeds the normal sitting posture range, behavioral state information corresponding to leaning forward or tilting the body is extracted.
[0102] Step S103: Based on physiological characteristic information, generate or update the virtual assistant image corresponding to the target occupant.
[0103] Here, the virtual assistant avatar can be a VPA (Virtual Assistant Image). The virtual assistant avatar can be a 2D cartoon character, a 3D virtual model, an anthropomorphic avatar, an animal-like companion image, or another visually interactive avatar. The virtual assistant avatar is used to provide interactive content to the target occupant, and can also be used to provide warning prompts to the driver or guardian.
[0104] When initially generating the virtual assistant avatar, basic configuration information corresponding to the target passenger is obtained. This basic configuration information includes at least one of the following: age, gender, nickname, appearance preference, interaction preference, and guardian settings. Based on the basic configuration information and physiological characteristics, the target passenger's physical appearance parameters are determined. These parameters include at least one of the following: facial structure parameters, head-to-body ratio parameters, height ratio parameters, body shape ratio parameters, skin color parameters, hairstyle parameters, clothing compatibility parameters, and facial expression style parameters. Subsequently, an initial virtual assistant avatar matching the target passenger is generated based on these physical appearance parameters, and this initial virtual assistant avatar is associated with and stored as the target passenger's identity identifier. The target passenger's identity identifier is determined through facial features, seat position, guardian binding information, or user account information.
[0105] When updating the virtual assistant's avatar, the system can reacquire the target passenger's physiological characteristics according to a preset update cycle, or when changes in the target passenger's physiological characteristics exceed a preset threshold. The preset update cycle can be one week, one month, three months, or a time period set according to actual needs. The reacquired physiological characteristics are compared with historical physiological characteristics to obtain growth change information. Growth change information includes at least one of the following: height changes, body shape changes, facial shape changes, hairstyle changes, and age changes. When the growth change information meets preset growth change conditions, the virtual assistant's height proportions, body shape proportions, facial structure, hairstyle, clothing style, or voice style are updated. Thus, the virtual assistant's avatar can dynamically update as the target passenger grows, maintaining a high degree of matching between the virtual assistant and the target passenger.
[0106] Step S104: Determine the behavioral risk level of the target occupant based on the behavioral status information.
[0107] Here, behavioral risk level is used to characterize the degree of safety risk corresponding to the current behavior or behavioral development trend of the target occupant. Behavioral risk level can include low risk level, medium risk level, and high risk level, and can also be set to risk-free level, attention level, warning level, and emergency level according to actual early warning needs.
[0108] In one implementation, the behavioral risk level is determined based on preset rules. These preset rules may include behavioral category rules, movement amplitude rules, duration rules, distance threshold rules, direction rules, confidence level rules, and vehicle state rules. For example, when the target occupant only slightly twists their body or briefly shifts their sitting posture, the behavioral risk level is determined to be low. When the target occupant exhibits significant body tilting, obvious forward leaning, or persistent improper sitting posture, the behavioral risk level is determined to be medium. When the target occupant shows signs of unfastening their seatbelt, leaning out of the window, leaving the child safety seat, or approaching dangerous components, the behavioral risk level is determined to be high.
[0109] In another implementation, a behavior risk level is determined based on a behavior recognition model. The behavior recognition model extracts spatial features and analyzes temporal changes from multiple consecutive frames of image data. The behavior recognition model can include CNN (Convolutional Neural Network), LSTM (Long Short-Term Memory), temporal convolutional networks, pose estimation networks, object detection networks, or multi-model fusion networks. The CNN extracts spatial features from each frame of image data to obtain spatial features such as human keypoints, the location of target components inside the vehicle, hand posture, limb orientation, and occlusion relationships. The LSTM analyzes the temporal feature sequence formed by multiple spatial feature vectors to identify temporal features such as hand movement trajectories, changes in torso tilt, changes in seatbelt status, and trends in window orientation. The behavior recognition model outputs behavior category, behavior confidence, and behavior development trend. Behavior categories include warning signs of seatbelt release, warning signs of leaning out of the window, warning signs of leaning forward, warning signs of getting out of a child seat, expressions of need, or normal riding behavior. Behavior confidence indicates the reliability of the behavior recognition results. Behavioral development trends indicate the direction or risk level of a behavior's continued development in the near future.
[0110] When determining the risk level of a behavior, a comprehensive assessment is taken into account, including the behavior category, behavior confidence level, behavior development trend, number of consecutive trigger frames, action duration, distance between the target occupant and the hazardous area, distance between the target occupant and the hazardous component, vehicle driving status, and historical behavior data. To reduce false alarms, consistency judgment is performed on the behavior categories corresponding to multiple consecutive frames of image data, and short-term interfering actions are filtered out. When the same behavior category appears consecutively within a preset time window, and the behavior confidence level meets the preset confidence requirements, the same behavior category is identified as a valid risk behavior. Subsequently, the behavior risk level is determined based on the degree of risk corresponding to the valid risk behavior.
[0111] Step S105: Based on the virtual assistant's image and behavioral risk level, output the interaction information and / or warning information corresponding to the target occupant.
[0112] Here, interactive information includes at least one of the following: voice interactive information, visual interactive information, animated interactive information, facial expression interactive information, content recommendation information, and demand response information. Warning information includes at least one of the following: voice warning information, visual warning information, light warning information, vibration warning information, text prompts, and vehicle auxiliary control prompts. Interactive and warning information are output through a virtual assistant avatar, in-vehicle display device, rear-seat display screen, in-vehicle audio system, rear-seat speakers, interior lighting devices, seat vibration devices, instrument panel display devices, or a guardian terminal.
[0113] When outputting interactive information, the target interaction scenario is determined based on the physiological characteristics and behavioral state information of the target occupant. Target interaction scenarios include daily companionship scenarios, soothing scenarios, entertainment interaction scenarios, need recognition scenarios, and safety reminder scenarios. For child occupants, the interaction content is determined according to their age group. For example, when the child occupant is younger, soothing nursery rhymes, calming music, or simple voice interaction are output. When the child occupant is older, story content, question-and-answer content, or gamified interactive content are output. During the interaction, the target occupant's facial expression status information continues to be acquired. When the facial expression status information indicates that the target occupant is crying, the interaction content is adjusted to soothing content. When the facial expression status information indicates that the target occupant is in a happy state, the current interaction content is maintained or a continuation of the current interaction content is output. When the behavioral state information indicates that the target occupant is pointing to a water cup, pointing to a toy, or performing other actions to express a need, a need prompt message is generated and sent to the in-vehicle display device or the guardian's terminal.
[0114] When issuing warning information, the warning intensity and method are determined based on the behavioral risk level. When the behavioral risk level is low, the virtual assistant provides a gentle reminder, such as prompting the target occupant to maintain a seated posture. When the behavioral risk level is medium, the virtual assistant provides a stronger reminder, and the in-vehicle notification devices provide auxiliary prompts, such as flashing lights or a display screen prompting the driver to pay attention to the rear seats. When the behavioral risk level is high, the specific dangerous action characteristics corresponding to the behavioral state information are extracted, and the virtual assistant displays a warning posture corresponding to the specific dangerous action characteristics, and a third warning message is issued. For example, when the specific dangerous action characteristic is a hand approaching or reaching towards the window, the virtual assistant displays a worried expression and waves a hand to stop the movement. When the specific dangerous action characteristic is a hand approaching the seatbelt buckle or pulling the seatbelt, the virtual assistant displays a serious reminder expression and points towards the seatbelt. When the specific dangerous action characteristic is leaning forward or bending down to pick up an object, the virtual assistant displays an anxious expression and gestures to sit upright.
[0115] After outputting a high-risk warning, the system continues to acquire the behavioral status information of the target occupant to determine the outcome of the status change after the warning output. If the status change indicates a decrease in the target occupant's behavioral risk level, the system reduces the intensity of the warning output or stops outputting the warning. If the status change indicates no decrease in the target occupant's behavioral risk level, the system increases the intensity of the warning output and controls the vehicle to perform auxiliary safety control operations. These auxiliary safety control operations include at least one of the following: reducing the volume of the in-vehicle multimedia system, controlling the seat vibration device to output vibration alerts, controlling the in-vehicle display device to output text alerts, controlling the interior lights to flash, sending alert information to the guardian's terminal, and suspending cruise control if it is active. Thus, the smart cockpit outputs matching interactive information and / or warning information based on the target occupant's actual behavioral risk status, thereby improving the relevance of the interactive content and the accuracy of the safety warnings.
[0116] In an optional implementation, refer to Figure 3 Step S102 includes the following steps S201-S203.
[0117] Step S201: Identify the occupant image region in the image data to obtain the facial features, body proportion features, and posture features of the target occupant.
[0118] Here, after receiving the image data, the image data is first preprocessed. The preprocessing process may include at least one of image denoising, brightness compensation, infrared image enhancement, image distortion correction, image frame selection, image cropping, and target area magnification.
[0119] After preprocessing, occupant image region recognition is performed on the image data. The occupant image region is the image area containing the target occupant's face, torso, limbs, and seated posture outline. The target occupant's image region is determined from the image data through face detection, human detection, human keypoint detection, semantic segmentation, object detection, or background subtraction. When the target occupant is a child, the location of the occupant image region is determined by considering the position of the child safety seat, the position of the rear seats, and the installation position of the rear-seat monitoring camera, to reduce interference from background objects inside the vehicle.
[0120] Facial features are extracted from the occupant image region. These features include at least one of the following: face shape, facial proportions, eye shape, nose contour, mouth contour, facial borders, hairstyle, skin color, and facial orientation. Facial features can be used to identify the target occupant and can also be used to subsequently generate or update the corresponding virtual assistant avatar. When the target occupant is a child, facial features can also be used to determine the child's growth and development over different time periods.
[0121] Body proportion features are extracted from the occupant image region. These features include at least one of head-to-body ratio, height ratio, torso ratio, shoulder width ratio, arm length ratio, leg length ratio, and body contour. In one embodiment, based on images of child safety seats, rear seat edges, seat markings, and dimensions or depth of in-vehicle fixed components, dimensional conversion relationships in the image data are determined, and the target occupant's height and body proportions are estimated based on the target occupant's head, shoulder, torso, and leg positions. In another embodiment, body proportion features in the current image data are compared with body proportion features in historical image data to identify the target occupant's growth trend.
[0122] Postural features are extracted from the occupant image region. These features include at least one of the following: head orientation, torso tilt angle, shoulder posture, hand position, arm extension direction, leg posture, sitting profile, body center of gravity position, and limb orientation. Human keypoint detection methods are used to identify head, shoulder, elbow, wrist, torso, and leg keypoints of the target occupant, and postural features are determined based on the positional relationships between these keypoints. Changes in postural features are identified through continuous multi-frame image data, such as the movement of hands from near the body to near the seatbelt buckle, the gradual movement of the head and torso towards the window, and the gradual shift from a normal sitting posture to a forward-leaning posture.
[0123] Step S202: Determine physiological feature information based on facial features and body proportion features.
[0124] Here, physiological characteristic information includes at least one of the following: identity recognition features, facial structure information, skin color information, hairstyle information, height estimation information, body proportion information, body shape information, age group information, and growth and change information. Physiological characteristic information is mainly used to characterize the individual physical appearance and growth status of the target occupant.
[0125] In one implementation, facial structure information of the target occupant is determined based on their face shape, facial feature proportions, hairstyle, and skin tone. Body proportion information of the target occupant is determined based on their head-to-body ratio, height ratio, and body contour. Then, the facial structure information and body proportion information are correlated to generate the target occupant's physiological characteristic information. This physiological characteristic information serves as the foundational data for generating or updating the virtual assistant avatar, ensuring a high degree of visual matching between the virtual assistant avatar and the target occupant.
[0126] In another implementation, based on currently acquired facial and body proportion features, combined with historically stored physiological information, it is determined whether the target occupant has undergone growth changes. Growth changes include at least one of the following: height increase, changes in head-to-body ratio, changes in body shape, changes in facial shape, changes in hairstyle, and changes in age. For example, facial and body images of the target occupant are acquired within a preset time period, and the current height proportion is compared with historical height proportions. When the comparison results indicate a significant change in the target occupant's height proportion or body shape, the change result is written into the target occupant's corresponding growth database, and this result serves as the basis for subsequent updates to the virtual assistant's image.
[0127] In the process of determining physiological characteristic information, the quality of data acquisition can also be assessed. For example, if the target occupant's face is obscured, the target occupant is not within the image acquisition range, the image brightness is insufficient, or the image is highly blurry, the update of physiological characteristic information can be temporarily suspended, or the image acquisition device can be restarted to acquire image data. By assessing the acquisition quality, errors in updating physiological characteristic information due to poor image quality in a single instance can be avoided.
[0128] Step S203: Determine behavioral state information based on posture features and the positional relationship between the target occupant and the target components inside the vehicle.
[0129] Here, the target components inside the vehicle include at least one of the following: seat belts, windows, doors, child safety seats, rear seats, cup holders, in-vehicle display devices, and armrests. Positional relationships include at least one of the following: distance relationship, directional relationship, overlapping relationship, proximity relationship, distance relationship, and contact relationship.
[0130] In one implementation, the spatial relationship between key points on the target occupant's body and target components within the vehicle is established. For example, the distance between the key points of the target occupant's hands and the seatbelt buckle is determined; the distance between the key points of the target occupant's head and the window boundary is determined; the angle between the key points of the target occupant's torso and the back of the child safety seat is determined; and the degree of overlap between the target occupant's body contour and the child safety seat area is determined. Based on these spatial relationships, it is determined whether the target occupant's actions pose a safety risk.
[0131] In another implementation, temporal analysis is performed on the posture features and positional relationships in consecutive multi-frame image data. This temporal analysis determines the movement development trend of the target occupant. For example, when the target occupant's hand key points continuously approach the seatbelt buckle in consecutive multi-frame image data, it indicates a premonition of seatbelt release; when the target occupant's head and torso key points continuously move towards the window, it indicates a premonition of leaning out of the window; when the target occupant's torso tilt angle continuously increases, it indicates a premonition of forward leaning or tilting; and when the target occupant's body contour gradually deviates from the child safety seat area, it indicates a premonition of leaving the child seat.
[0132] Behavioral status information includes the target occupant's current behavior category, premonitory signs, posture, limb movement status, duration, direction, amplitude, and trend of the movement, as well as information on the relative positional changes between the target occupant and target components within the vehicle. This behavioral status information is used to subsequently determine the target occupant's behavioral risk level. It is not limited to completed dangerous actions but can also include the trend of actions preceding a dangerous action. By identifying premonitory signs, interactive or warning information can be output before the target occupant actually completes the dangerous action.
[0133] In scenarios where the target occupant is a child, behavioral status information is used to characterize whether the child occupant is unfastening their seatbelt, reaching for the seatbelt buckle, approaching the window with their head, leaning forward significantly, tilting their body noticeably, bending down to pick up an object, or moving away from their child safety seat. Appropriate follow-up actions are selected based on this behavioral status information. For example, if the child occupant slightly twists their body, the behavioral status information is recorded as a low-risk behavior; if the child occupant tilts their body significantly, the behavioral status information is recorded as a medium-risk behavior; if the child occupant's hand approaches the seatbelt buckle and continuously pulls on the seatbelt, the behavioral status information is recorded as a precursor to a high-risk behavior.
[0134] In some implementations, behavioral state information may also include actions expressing needs. These actions include the target occupant pointing to a water cup, pointing to a toy, waving, shaking their head, nodding, or expressing a need through facial expressions. A need prompt is generated based on these actions and sent to the in-vehicle display device or a guardian's terminal. By incorporating these actions into the behavioral state information, the smart cockpit can not only identify safety risks but also recognize the target occupant's interaction needs, thereby improving the targeting of in-vehicle interactions.
[0135] In an optional implementation, the in-vehicle target components include at least one of seat belts, windows, doors, and child safety seats.
[0136] Here, the target components inside the vehicle include at least one of the following: seat belts, windows, doors, and child safety seats. Seat belts may include seat belt webbing, seat belt buckles, seat belt latches, and seat belt anchor points. Windows include rear window glass, window borders, window opening areas, and window control areas. Doors include rear doors, door handles, door locking areas, and door trim panels. Child safety seats include child safety seat backrests, child safety seat cushions, child safety seat side wings, child safety seat restraint straps, and child safety seat mounting boundaries.
[0137] When the target occupant is a child occupant, step S203 includes the following steps S301-S302.
[0138] Step S301: Identify the key points of the child occupant's body, limb orientation, and positional changes of the key points of the body relative to the target component inside the vehicle; wherein, the key points of the body include at least one of the hands, head, and torso.
[0139] Here, after determining the occupant image region in the image data, human keypoint recognition is performed on the occupant image region. Human keypoints include hand keypoints, head keypoints, and torso keypoints. Hand keypoints include at least one of the following: palm center, wrist, elbow, and finger direction. Head keypoints include at least one of the following: head center, facial orientation point, and head boundary point. Torso keypoints include at least one of the following: shoulder center, chest center, waist center, and body center of gravity. The positions of the child occupant's human keypoints in consecutive frames of image data are determined using a human pose recognition model, a human skeleton recognition model, an object detection model, or an image segmentation model.
[0140] Limb orientation is used to characterize the direction of movement of a child occupant's hands, head, or torso. Limb orientation is determined based on changes in the position of key human body points in adjacent image frames. For example, when a hand key point moves from the inside of the child occupant's body towards the seatbelt buckle, the hand is determined to be facing the seatbelt buckle. When a hand or head key point moves from the inside of the seat towards the window, the hand or head is determined to be facing the window. When a torso key point moves from the back of the child safety seat towards the front of the vehicle, the torso is determined to be facing the front of the vehicle.
[0141] Position change information includes at least one of the following: changes in distance, direction, angle, overlap, and contact trend between key human body points and target components inside the vehicle. The positions of key human body points and target components inside the vehicle are determined in consecutive multi-frame image data, and position change information is calculated based on positional differences between different image frames. For example, it calculates whether the distance between the hand key point and the seatbelt buckle is continuously decreasing, whether the distance between the head key point and the window boundary is continuously decreasing, whether the angle between the torso key point and the child safety seat back is continuously increasing, and whether the overlap area between the child occupant's body contour and the child safety seat area is continuously decreasing.
[0142] When identifying position change information, the judgment is made by combining the fixed position of the target component inside the vehicle or by identifying its position in real time. For components with relatively fixed positions, such as seat belt buckles, window edges, door handles, and child safety seat edges, reference areas for the target components inside the vehicle are pre-established. For objects whose positions change over time, such as seat belt webbing, child occupant's hands, or child occupant's head, the corresponding positions are re-identified in each frame of image data. By combining fixed reference areas with real-time identification results, the accuracy of position change information can be improved.
[0143] Step S302: Based on limb orientation and position change information, determine whether the child occupant has at least one of the following behavioral state information: signs of seatbelt release, signs of leaning out of the window, signs of leaning forward, signs of tilting, and signs of leaving the child seat.
[0144] Here, a warning sign of a dangerous action refers to a situation where the dangerous action has not yet fully occurred, but the child occupant's direction, trajectory, or body posture already shows a tendency to develop into a dangerous situation. By recognizing these warning signs, interactive or early warning information can be provided to the child occupant before they complete the dangerous action.
[0145] In one implementation, when a child occupant's hand key points continuously move towards the seatbelt buckle, and the distance between the hand key points and the seatbelt buckle continuously decreases within a preset time window, it is determined that the child occupant is showing signs of impending seatbelt release. If the image data further shows that the hand key points overlap with the seatbelt buckle area, or that the seatbelt webbing position abnormally shifts, the risk level corresponding to the signs of impending seatbelt release is increased.
[0146] In one implementation, when a child occupant's head, hand, or torso key points continuously move towards the window, and the distance between the corresponding key points and the window boundary or window opening area continuously decreases, it is determined that the child occupant is showing signs of wanting to stick their head out of the window. If the child occupant's head or hand boundary approaches the window opening area, the behavioral status information is recorded as a high-risk action precursor.
[0147] In one implementation, a child occupant is identified as exhibiting signs of forward leaning when the angle between the child occupant's torso and the back of the child safety seat continuously increases, or when the child's center of gravity continuously shifts forward towards the vehicle. These signs are used to identify situations such as the child occupant bending down to pick up an object, leaning forward, or deviating from a normal reclining posture.
[0148] In one implementation, when a child occupant's shoulder center, waist center, or center of gravity continuously shifts to the left or right side of the vehicle, and the torso tilt angle exceeds a preset range, it is determined that the child occupant exhibits signs of impending body tilt. These signs are used to identify situations such as significant twisting, tilting towards the vehicle door, or tilting towards the vehicle window.
[0149] In one implementation, when the overlap between the child occupant's body contour and the child safety seat area continuously decreases, or when key points of the torso, waist, and legs gradually deviate from the child safety seat cushion area, it is determined that the child occupant is showing signs of leaving the child seat. If the child occupant also exhibits abnormal seatbelt restraint, upward shift of body center of gravity, or forward shift of body center of gravity, the risk level corresponding to the signs of leaving the child seat is increased.
[0150] When determining behavioral status information, judgments are made using multiple consecutive frames of image data to avoid misidentification due to errors in a single frame. For example, it can be determined whether the same warning signs of dangerous actions occur consecutively within a preset time window, or whether the recognition confidence level corresponding to the warning signs meets preset requirements. For interfering actions such as short pauses, occlusion, raising and waving hands, or normal adjustments to sitting posture, filtering is performed based on duration, direction of movement, target area, and amplitude of the action. The filtered action results are used as the behavioral status information of the child occupant.
[0151] Behavioral state information includes at least one of the following: type of dangerous action precursor, trajectory of changes in key points of the human body, limb orientation, magnitude of positional change, duration of action, and confidence level of action.
[0152] Therefore, the interactive early warning method for the smart cockpit provided in this application can not only identify dangerous behaviors that child occupants have already engaged in, but also identify potential dangerous behaviors in advance based on the changing trends of child occupants' movements, thereby improving the initiative and accuracy of safety monitoring for rear-seat child occupants.
[0153] In an optional implementation, refer to Figure 4 In step S103, a virtual assistant image corresponding to the target occupant is generated based on physiological characteristic information, including the following steps S401-S404.
[0154] Step S401: Obtain the basic configuration information corresponding to the target occupant; wherein, the basic configuration information includes at least one of age information, gender information, and image preference information.
[0155] Here, basic configuration information can be input by guardians, drivers, or vehicle users through in-vehicle display devices, mobile terminals, or in-vehicle voice interaction interfaces, or it can be automatically read by the in-vehicle central control system based on the target occupant's historical riding records, user account information, or identity recognition results. When the target occupant is a child, the basic configuration information includes at least one of the following: the child's age, gender, nickname, image preferences, interaction preferences, commonly used languages, and content preferences.
[0156] Age information is used to determine the age-appropriate appearance and interaction style of the virtual assistant avatar. For example, the virtual assistant avatar for younger children uses softer contours, simpler facial expressions and movements, and a slower speech rhythm. The virtual assistant avatar for older children uses richer facial expressions and movements, more complete interactive sentences, and clothing styles more appropriate for their age group. Gender information serves as reference information in the generation process of the virtual assistant avatar's appearance, but it does not require the virtual assistant avatar to adopt a fixed gender appearance. Image preference information includes at least one of the following: cartoon character preference, anthropomorphic character preference, animal character preference, color preference, clothing preference, and voice style preference.
[0157] In some implementations, basic configuration information is associated with the target occupant's identity. This identity is determined based on the target occupant's facial features, seat location, guardian account, vehicle user account, or child safety seat binding information. Basic configuration information can be established when the target occupant first rides in the vehicle, or it can be retrieved from previously stored information when the target occupant rides again. By establishing basic configuration information, the virtual assistant avatar can be generated while taking into account the target occupant's personal preferences and usage habits, avoiding the need to reconfigure the interactive avatar for each ride.
[0158] In one specific implementation, when a guardian uses the rear-seat child care function for the first time, they can input the child's age, gender, and appearance preferences on the in-vehicle display device. For example, the guardian can select preferences such as a cartoon character, a gentle voice, and nursery rhymes. The input information is used as basic configuration information and combined with physiological characteristics to generate a virtual assistant avatar corresponding to the child.
[0159] Step S402: Determine the physical appearance parameters of the target occupant based on physiological characteristic information; wherein, the physical appearance parameters include at least one of facial structure parameters, head-to-body ratio parameters, height ratio parameters, skin color parameters, hairstyle parameters, and clothing adaptation parameters.
[0160] Here, physiological feature information is obtained by analyzing image data acquired by an image acquisition device. This physiological feature information includes at least one of the following: facial features, body proportion features, height estimation information, body contour information, skin color features, hairstyle features, and age group information. This physiological feature information is then converted into body shape parameters suitable for generating a virtual avatar, so that the virtual assistant avatar can correspond in appearance to the target occupant.
[0161] Facial structure parameters are determined based on features such as the target occupant's face shape, facial contours, facial proportions, eye shape, nose contour, mouth contour, and facial orientation. These parameters control the virtual assistant's face shape, eye shape, mouth shape, and overall facial style. To avoid overly realistic portrayals, the facial structure parameters are cartoonized, allowing the virtual assistant to retain the target occupant's highly recognizable features while presenting a more approachable appearance suitable for in-vehicle interaction scenarios.
[0162] The head-to-body ratio and height ratio parameters are determined based on the proportional relationships between the target occupant's head, torso, limbs, and seat reference points. Dimensional reference relationships are established using images of the child safety seat boundary, rear seat edge, seat markings, and the dimensions or depth of fixed components within the vehicle. These dimensional reference relationships are then used to estimate the target occupant's head-to-body ratio and height ratio. When the target occupant is a child, the head-to-body ratio and height ratio parameters can be used to reflect the child's growth stage, allowing the virtual assistant avatar to change in height, body proportions, and clothing size as the child grows.
[0163] Skin color parameters are determined based on the color features of the target occupant's facial or hand areas. Before determining skin color parameters, the image data undergoes brightness compensation and white balance processing to reduce the impact of in-vehicle lighting, night vision mode, or display glare on skin color judgment. Hairstyle parameters are determined based on the boundary shape, hair color features, and hair outline of the target occupant's head area. Hairstyle parameters are used to generate the hairstyle style of the virtual assistant avatar, or to determine whether the virtual assistant avatar's hairstyle appearance needs adjustment during subsequent updates.
[0164] Clothing adaptation parameters are determined based on the target passenger's age, gender, image preferences, developmental stage, and body proportions. These parameters determine the virtual assistant's clothing type, size, color scheme, and seasonal style. When the target passenger is a child, clothing styles more suitable for the child's cognitive characteristics are selected based on their age group; for example, soft colors and simple patterns are appropriate for younger children, while more elaborate clothing elements are suitable for older children.
[0165] In some implementations, the facial features parameters are stabilized. Stabilization includes at least one of outlier removal, fusion of multiple acquisition results, smoothing of historical parameters, and limitation of variation range. For example, if the target occupant briefly lowers their head, turns their head, obscures their face, or experiences abnormal lighting, the facial features parameters can be determined by combining historical facial features parameters and multi-frame image recognition results instead of directly updating them with a single recognition result. Stabilization prevents unreasonable changes to the virtual assistant's image due to a single image recognition error.
[0166] Step S403: Based on the basic configuration information and physical parameters, generate an initial virtual assistant image corresponding to the target occupant.
[0167] Here, the initial virtual assistant image can be a two-dimensional cartoon image, a three-dimensional virtual image, an anthropomorphic image, an animal-like companion image, or another visual interactive image. First, the overall style of the virtual assistant image is determined based on the basic configuration information. Then, the facial structure, head-to-body ratio, height ratio, skin color, hairstyle, and clothing of the virtual assistant image are adjusted according to the physical parameters to obtain an initial virtual assistant image that corresponds to the target passenger.
[0168] In one implementation, a basic avatar template matching the basic configuration information is selected from a preset virtual avatar template library, and then the basic avatar template is parametrically adjusted according to physical appearance parameters. The preset virtual avatar template library includes multiple virtual avatar templates for different age groups, different appearance styles, and different interaction styles. A basic avatar template is selected based on the target occupant's age and appearance preferences, and then adjusted according to facial structure parameters, head-to-body ratio parameters, skin color parameters, and hairstyle parameters.
[0169] In another implementation, a virtual assistant avatar is generated directly based on physical appearance parameters. Facial structure parameters are mapped to the virtual assistant avatar's face shape and facial feature proportions; head-to-body ratio parameters are mapped to the virtual assistant avatar's overall body proportions; height ratio parameters are mapped to the virtual assistant avatar's growth stage; skin color and hairstyle parameters are mapped to the virtual assistant avatar's skin color and hairstyle; and clothing adaptation parameters are mapped to the virtual assistant avatar's clothing appearance. Thus, the initial virtual assistant avatar can reflect the target passenger's main physical characteristics while also satisfying the image preferences of the target passenger or guardian.
[0170] After generating the initial virtual assistant avatar, configure its interactive attributes. These attributes include at least one of the following: voice style, facial expressions and gestures, tone of voice, type of reassuring content, commonly used interactive phrases, and warning gestures. For example, if the target passenger is a young child, configure the virtual assistant avatar with a gentle voice, a smiling expression, soothing music, and brief prompts. If the target passenger is an older child, configure story interaction, question-and-answer interaction, and clearer safety prompts.
[0171] Step S404: Associate and store the initial virtual assistant image with the identity identifier of the target passenger.
[0172] Here, the identity identifier includes the target occupant's facial feature identifier, user account identifier, seat identifier, child safety seat binding identifier, guardian account identifier, or occupant profile identifier generated by the vehicle's central control system. After being associated and stored, when the target occupant re-enters the vehicle or is recognized again by the image acquisition device, the already generated virtual assistant avatar is invoked, without needing to complete the entire avatar generation process again.
[0173] The associated storage data includes at least one of the following: basic configuration information, physical parameters, initial virtual assistant image data, interaction attribute data, historical physiological characteristic information, and historical update records. This associated storage data can be stored locally or synchronized to a cloud account or guardian's terminal after user authorization. Data involving the target occupant's identity and image characteristics is encrypted, subject to access control, or anonymized to reduce the risk of privacy breaches.
[0174] By associating and storing identity tags, different virtual assistant avatars can be created for different target occupants, thereby improving the personalization and continuity of in-vehicle interaction.
[0175] In an optional implementation, refer to Figure 5 In step S103, the virtual assistant image corresponding to the target occupant is updated based on physiological characteristic information, including the following steps S501-S504.
[0176] Step S501: Reacquire the physiological characteristic information of the target occupant according to the preset update cycle or when the detected physiological characteristic change exceeds the preset threshold.
[0177] Here, the preset update cycle is determined based on the target occupant's age group, growth rate, guardian settings, or the vehicle's default policy. For example, when the target occupant is a child, the preset update cycle is set to one month, three months, or six months. When the child occupant is in an age group where height and physical appearance change rapidly, the preset update cycle can be relatively shorter. When the child occupant is in an age group where physical appearance changes slowly, the preset update cycle can be extended according to the actual situation.
[0178] In addition to updating according to a preset update cycle, the system can also reacquire the target occupant's physiological characteristic information when changes in physiological characteristics exceed a preset threshold. Changes in physiological characteristics can include at least one of the following: changes in height ratio, head-to-body ratio, body shape, facial contour, hairstyle, skin color, and age group. The preset thresholds can be set separately for different types of physiological characteristics. For example, changes in height ratio correspond to a first preset threshold, changes in body shape correspond to a second preset threshold, and changes in facial contour correspond to a third preset threshold.
[0179] When reacquiring physiological feature information, the image acquisition device is invoked to collect the current image data of the target occupant, and facial features, body proportion features, and posture auxiliary features are extracted from the current image data. Facial features are used to update information such as face shape, facial feature proportions, hairstyle, and skin color. Body proportion features are used to update information such as height ratio, head-to-body ratio, and body contour. Posture auxiliary features are used to eliminate the influence of non-standard postures such as looking down, turning to the side, or occlusion on the physiological feature acquisition results.
[0180] Before reacquiring physiological feature information, the current image data is assessed for acquisition quality. This assessment includes whether the face is clear, whether the target occupant is within the image acquisition range, whether the target occupant's body outline is complete, whether the image brightness meets requirements, whether the image blur level is below a preset blur threshold, and whether the target occupant is in a relatively stable sitting posture. If the current image data does not meet the acquisition quality requirements, the update is paused, or image data is acquired again at the next preset time. This acquisition quality assessment reduces erroneous updates caused by accidental occlusion, changes in lighting, and vehicle vibration.
[0181] Step S502: Compare the newly acquired physiological characteristic information with the historical physiological characteristic information to obtain growth change information.
[0182] Here, historical physiological characteristic information can be the physiological characteristic information stored when the target occupant first generates the virtual assistant image, the physiological characteristic information stored when the virtual assistant image is last updated, or a set of physiological characteristic information corresponding to multiple historical time points.
[0183] The comparison process is performed separately for different feature types. For height proportion, the difference between the current height proportion and the historical height proportion is compared to obtain height change information. For head-to-body proportion, the difference between the current head-to-body proportion and the historical head-to-body proportion is compared to obtain body proportion change information. For facial structure, the differences between the current facial contour, facial feature proportions, and facial boundaries and the historical facial structure are compared to obtain facial change information. For hairstyle and skin tone, the differences between the current hairstyle contour, hair color display area, and skin tone display features and the historical records are compared to obtain appearance change information.
[0184] Growth change information includes at least one of the following: changes in height, percentage changes in height, changes in head-to-body ratio, changes in body contour, changes in facial structure, changes in hairstyle, changes in clothing fit, and changes in age group. Growth change information may also include the time of change, trend of change, and confidence level of change. The time of change indicates the data collection time corresponding to this comparison. The trend of change indicates the continuous direction of change of the target occupant across multiple historical time points, such as continuous height increase, gradual changes in body proportion, or changes in hairstyle. The confidence level of change indicates the reliability of the growth change information.
[0185] In one implementation, growth change information is generated based on multiple data acquisitions. For example, multiple sets of image data are continuously acquired within an update cycle, and multiple sets of physiological feature information are extracted from each set. Then, the multiple sets of physiological feature information are averaged, filtered, or weighted and fused to obtain the current physiological feature information. Finally, the current physiological feature information is compared with historical physiological feature information. By fusing multiple acquisition results, the error of a single acquisition can be reduced, and the stability of growth change information can be improved.
[0186] Step S503: If the growth change information meets the preset growth change conditions, update the physical parameters of the virtual assistant image based on the growth change information.
[0187] Here, the preset growth change conditions include at least one of the following: height change exceeding the height change threshold, head-to-body ratio change exceeding the ratio change threshold, body shape change exceeding the body shape change threshold, facial structure change exceeding the facial change threshold, hairstyle change meeting the hairstyle update conditions, age change, or guardian confirmation that an update is needed.
[0188] When updating physical appearance parameters, the height ratio, head-to-body ratio, facial structure, hairstyle, skin tone, and clothing adaptation parameters of the virtual assistant avatar are adjusted according to growth and change information. For example, when the growth and change information indicates an increase in the target passenger's height ratio, the height ratio parameter of the virtual assistant avatar is increased, and the head-to-body ratio parameter is adjusted appropriately to make the virtual assistant avatar appear as a grown-up. When the growth and change information indicates a change in the target passenger's hairstyle, the hairstyle parameter of the virtual assistant avatar is updated. When the growth and change information indicates a change in the target passenger's age group, the clothing adaptation parameters of the virtual assistant avatar are updated to match the virtual assistant avatar's clothing, accessories, or overall style with the new age group.
[0189] During the update process, the basic identifying features of the virtual assistant's image are retained. These basic features include main facial features, key facial proportions, common hairstyles, or the image preferences set by the guardian. By retaining these basic features, the virtual assistant's image can maintain a continuous correspondence with the target passenger after updates, avoiding significant differences between the updated and pre-update images.
[0190] In some implementations, the virtual assistant's appearance can be updated gradually. A gradual update approach breaks down a single change into multiple smaller appearance adjustments, which are then completed progressively over multiple interaction cycles.
[0191] In some implementations, after updating the physical appearance parameters, the interactive attributes associated with those parameters can also be updated simultaneously. Interactive attributes include facial expressions and gestures, voice rhythm, voice content difficulty, interactive dialogue length, nursery rhyme or story content type, and safety reminder tone. For example, as the child passenger's age increases, the virtual assistant avatar gradually switches from simple nursery rhyme interactions to short story interactions or question-and-answer interactions. By synchronously updating interactive attributes, the virtual assistant avatar can not only grow in appearance with the target passenger but also match the interactive content to the target passenger's developmental stage.
[0192] After the update is complete, the updated physical appearance parameters, the updated virtual assistant appearance, growth and change information, and the update time will be stored in the crew file corresponding to the target crew member. The stored data can be used as historical physiological characteristic information or historical physical appearance parameters for the next update.
[0193] Step S504: If the growth change information does not meet the preset growth change conditions, keep the appearance parameters of the virtual assistant image unchanged.
[0194] Here, the fact that the growth change information does not meet the preset growth change conditions indicates that the physiological characteristics of the target occupant have changed only slightly, or that the current recognition result may be due to non-realistic growth factors such as posture changes, occlusion, lighting changes, or image blurring. In this case, keeping the physical parameters unchanged can prevent meaningless changes in the virtual assistant's appearance.
[0195] While keeping the physical appearance parameters unchanged, the facial structure parameters, head-to-body ratio parameters, height ratio parameters, skin color parameters, hairstyle parameters, and clothing adaptation parameters of the current virtual assistant image will be retained. Simultaneously, the physiological characteristic information re-acquired this time will be saved as a temporary record or candidate record. If physiological characteristic changes in the same direction occur in multiple subsequent acquisition cycles, and the cumulative changes meet the preset growth change conditions, then the physical appearance parameters can be updated again in subsequent update processes.
[0196] In some implementations, if the growth change information does not meet the preset growth change conditions, but the guardian actively confirms through the in-vehicle display device or mobile terminal that the target occupant's appearance has changed, some physical parameters of the virtual assistant image can also be updated based on the guardian's confirmation. For example, if the target occupant's hairstyle has changed significantly, but the image recognition result has a low confidence level due to occlusion, the guardian can manually confirm the hairstyle change and then update the hairstyle parameters of the virtual assistant image.
[0197] The virtual assistant avatar provided in this application can be updated when the target passenger undergoes significant changes in growth, and remains stable when the target passenger's changes are not significant or the recognition results are unstable. Therefore, the virtual assistant avatar can be matched to the target passenger's growth status over a long period, while avoiding frequent changes that could affect the continuity of interaction.
[0198] In an optional implementation, step S104 includes the following steps S601-S606.
[0199] Step S601: Acquire continuous multi-frame image data for determining behavioral state information.
[0200] Here, continuous multi-frame image data is a sequence of image frames arranged in chronological order of acquisition time. Continuous multi-frame image data is used to characterize the movement changes of the target occupant over a period of time.
[0201] Continuous multi-frame image data can be obtained by capturing images from a real-time video stream or by reading images cached within a preset time window. The preset time window is set according to the duration of the behavior to be recognized. For example, for behaviors such as a hand approaching a seatbelt buckle, a body gradually leaning forward, or a head gradually approaching a car window, an image frame sequence covering the action process from 0.5 seconds to 3 seconds is selected. The number of frames in the continuous multi-frame image data is determined based on the image acquisition frame rate, onboard computing power, and recognition accuracy requirements; for example, 16, 32, or 64 frames of image data may be selected.
[0202] When acquiring multiple consecutive frames of image data, priority is given to selecting image frames that meet the requirements for image clarity, have a complete target occupant area, are identifiable by key human body points, and have identifiable locations of target components inside the vehicle. If some image frames are blurry, occluded, or have insufficient lighting, the image frames can be enhanced, or image frames that do not meet the quality requirements can be removed, and new image frames can be added from adjacent time periods.
[0203] Step S602: Input multiple consecutive frames of image data into the pre-trained behavior recognition model.
[0204] Here, the behavior recognition model is used to identify the behavior category, behavior confidence, and behavior development trend of a target occupant based on continuous multi-frame image data. The behavior recognition model can adopt a hybrid model combining CNN and LSTM, or a structure that integrates pose recognition model, temporal convolutional model, 3D convolutional network model, object detection model, and classification model.
[0205] The behavior recognition model is pre-trained based on sample image data. This data includes consecutive image frames corresponding to various behaviors of the target occupant, such as normal seating, slight body twisting, leaning forward, tilting, hand near the seatbelt buckle, pulling the seatbelt, head near the window, hand near the window, and leaving a child seat. During training, the sample image data is labeled with the corresponding behavior category, risk level, action start time, action end time, and key component regions. Based on this, the behavior recognition model can learn the correlation between the target occupant's body movements and target components inside the vehicle.
[0206] Before inputting the data into the model, multiple consecutive frames of image data undergo size unification, brightness normalization, target occupant region cropping, and annotation of target component regions or normalization of key point coordinates within the vehicle. This unified input format improves the recognition stability of the behavior recognition model. If the vehicle is in a nighttime or low-light environment, infrared-enhanced image frames are input into the behavior recognition model to ensure its operation under various lighting conditions.
[0207] Step S603: Spatial features are extracted from each frame of image data using a behavior recognition model to obtain a spatial feature vector corresponding to each frame of image data; wherein, the spatial feature vector is used to characterize the posture features of the target occupant and its positional relationship with the target components inside the vehicle.
[0208] Here, after receiving multiple consecutive frames of image data, the behavior recognition model first performs spatial feature extraction on each frame. Spatial feature extraction is used to obtain the target occupant's body posture, limb positions, the positions of target components inside the vehicle, and the spatial relationship between the target occupant and the target components inside the vehicle from a single frame image. The behavior recognition model uses a CNN to extract features from each frame of image data, converting each frame of image data into a corresponding spatial feature vector.
[0209] Spatial feature vectors include human body keypoint features, limb orientation features, seated posture contour features, in-vehicle target component position features, and relative distance features. Human body keypoint features represent the positions of key body parts such as hands, head, shoulders, torso, and waist. Limb orientation features represent the direction of arm extension, head orientation, and torso tilt direction. Seated posture contour features represent whether the target occupant is leaning against the seat, whether the body is leaning forward, or whether the body is leaning to the side. In-vehicle target component position features represent the positions of target components such as seat belts, windows, doors, and child safety seats in the image. Relative distance features represent the distance between the hand and the seat belt buckle, the distance between the head and the window boundary, the angle between the torso and the child safety seat back, and the degree of overlap between the body contour and the child safety seat area.
[0210] In one implementation, the CNN first extracts image features from the human body region, hand region, seat belt region, window region, and child safety seat region, and then fuses these image features to obtain the spatial feature vector corresponding to a single frame image. In another implementation, a residual network ResNet50, a mobile network MobileNetV2, or an efficient network EfficientNet is used as the spatial feature extraction network to balance recognition accuracy and in-vehicle operating speed. The dimension of the spatial feature vector can be determined according to the model structure, for example, it can be 512-dimensional or 1024-dimensional.
[0211] By extracting spatial features, pixel information in a single frame image is converted into structured feature information for behavior judgment, reducing irrelevant background interference in the original image and preserving key information such as human body, target parts and spatial relationships.
[0212] Step S604: Combine multiple spatial feature vectors according to the acquisition time sequence to obtain a temporal feature sequence.
[0213] Here, after obtaining the spatial feature vector corresponding to each frame of image data, multiple spatial feature vectors are arranged and combined according to the acquisition time sequence of the image data to obtain a temporal feature sequence. The temporal feature sequence is represented as a feature set formed by multiple spatial feature vectors in chronological order. The temporal feature sequence can reflect the continuous change process of the target occupant's actions from the initial state to the development state.
[0214] For example, the process of a child occupant's hand moving from near their body to near the seatbelt buckle is represented in the time-series feature sequence as a gradual decrease in the distance between the key points of the hand and the seatbelt buckle. The process of a child occupant's body gradually leaning forward is represented in the time-series feature sequence as a gradual increase in the angle between the key points of the torso and the child safety seat back. The process of a child occupant's head gradually approaching the window is represented in the time-series feature sequence as a gradual decrease in the distance between the key points of the head and the window edge.
[0215] When combining spatial feature vectors, temporal alignment is performed on the spatial feature vectors corresponding to multiple consecutive frames of image data to ensure the accurate temporal order between feature vectors. If some image frames are missing or of low quality, adjacent frame interpolation, feature completion, or frame skipping are used to ensure the continuity of the temporal feature sequence. Inter-frame variations, such as keypoint displacement, limb orientation changes, distance changes, and angle changes, can also be incorporated into the temporal feature sequence for more accurate subsequent temporal variation analysis.
[0216] Step S605: Perform time series change analysis on the time series feature sequence to obtain the behavior recognition results corresponding to the target occupant; the behavior recognition results include behavior category, behavior confidence and behavior development trend.
[0217] Here, temporal variation analysis is used to identify the patterns of change in the target occupant's actions over time and to determine the type of behavior of the current action. Temporal variation analysis can be performed using LSTM, gated recurrent units, temporal convolutional networks, Transformer networks, or other temporal analysis networks.
[0218] The behavior recognition model identifies the movement trajectory and trend of a target occupant based on temporal feature sequences. For example, it identifies where the hand begins to move, where it moves to, and whether the hand eventually approaches the seatbelt buckle or the window area; it identifies whether the torso gradually changes from a normal leaning position to a forward leaning position; it identifies whether the head and torso continuously move towards the window; and it identifies whether the child occupant's body outline gradually deviates from the child safety seat area. By analyzing the temporal changes, the behavior recognition model can distinguish between brief hand waving, normal posture adjustments, and warning signs of dangerous actions, reducing misjudgments caused by single-frame analysis.
[0219] The behavior recognition results include behavior category, behavior confidence score, and behavior development trend. Behavior categories include at least one of the following: normal sitting behavior, slight twisting behavior, signs of leaning forward, signs of tilting forward, signs of seatbelt release, signs of leaning out of the window, signs of leaving the child seat, and actions expressing needs. Behavior confidence score indicates the reliability of the behavior recognition model's judgment of the behavior category. Behavior development trend indicates the direction and possible outcome of the target occupant's continued actions, such as the hand continuing to move closer to the seatbelt buckle, the head continuing to move closer to the window, the torso continuing to move away from the child safety seat back, or the body contour continuing to deviate from the child safety seat area.
[0220] Step S606: Determine the behavioral risk level based on the behavioral category, behavioral confidence level, and behavioral development trend.
[0221] Here, behavioral risk level is used to indicate the degree of impact of the target occupant's current behavior on passenger safety. Behavioral risk level can include low risk level, medium risk level, and high risk level, and can also be divided into more levels according to actual application needs.
[0222] When determining the risk level of a behavior, first determine the basic risk level corresponding to the behavior category. For example, slight twisting of the body or short-term shift in posture corresponds to a low-risk level. Significant forward leaning or continuous tilting of the body corresponds to a medium-risk level. Hands continuously approaching the seatbelt buckle, pulling on the seatbelt, head approaching the window opening area, or body clearly leaving the child safety seat area corresponds to a high-risk level. Then, adjust the basic risk level based on the confidence level and the trend of the behavior.
[0223] When the confidence level of a behavior is lower than the confidence level requirement for the corresponding behavior category, the behavior risk level is not increased temporarily, or subsequent image frames are collected for confirmation. When the confidence level of a behavior meets the confidence level requirement, and the trend of the behavior points towards a preset dangerous part or a preset dangerous area, the behavior risk level is increased. For example, if the key points of the hand are continuously close to the seat belt buckle and the behavior confidence level meets the requirement, the behavior risk level is determined to be high-risk; if the torso tilt angle continues to increase but has not yet touched the dangerous area, the behavior risk level is determined to be medium-risk; if the target occupant only briefly twists their body and the direction of the movement does not point towards a dangerous part, the behavior risk level is determined to be low-risk.
[0224] In some implementations, the behavioral risk level can be determined by combining vehicle status and the target occupant's historical behavior data. Vehicle status includes vehicle speed, whether the vehicle is in motion, whether the windows are open, whether the doors are locked, and whether cruise control is activated. The same behavior can correspond to different risk levels under different vehicle statuses. For example, when the vehicle is traveling at high speed and the windows are open, the risk level of the head near the window area is higher than when the vehicle is stationary. The target occupant's historical behavior data is used to determine whether the target occupant has a habit of frequently unfastening their seatbelt, frequently leaning forward, or repeatedly triggering warnings, and is used to individually adjust the risk assessment criteria.
[0225] In an optional implementation, step S606 includes the following steps S701-S705.
[0226] Step S701: Obtain the risk assessment conditions corresponding to the behavior category.
[0227] Here, the behavioral categories include at least one of the following: normal riding behavior, slight twisting behavior, signs of leaning forward, signs of tilting forward, signs of unfastening the seatbelt, signs of leaning out of the window, signs of leaving the child seat, and actions expressing needs. Different behavioral categories correspond to different sources of risk, therefore, different risk assessment conditions can be configured for different behavioral categories.
[0228] Risk assessment criteria include at least one of the following: confidence level requirement, duration requirement, number of consecutive frames requirement, movement amplitude requirement, distance requirement between key human body points and target components inside the vehicle, direction of movement requirement of key human body points, trend of movement requirement, and vehicle status requirement. For example, risk assessment criteria for warning signs of seatbelt release include continuous movement of the hand key points toward the seatbelt buckle, a distance between the hand key points and the seatbelt buckle being less than a preset distance, and the behavior confidence level meeting a preset confidence level requirement. Risk assessment criteria for warning signs of leaning out of the window include continuous movement of the head or hand key points toward the window area, a continuous decrease in the distance between the key human body points and the window boundary, and the window area being open or potentially openable. Risk assessment criteria for warning signs of leaning forward include an angle between the torso key points and the child safety seat back being greater than a preset angle, and a continuous movement of the body's center of gravity toward the front of the vehicle. Risk assessment criteria for warning signs of leaving the child seat include a lower than preset ratio of overlap between the child occupant's body contour and the child safety seat area, or a continuous deviation of the body's center of gravity from the child safety seat cushion area.
[0229] Risk assessment criteria can be pre-stored in the vehicle's central control system's local memory, or adjusted based on vehicle model, seat structure, child safety seat installation location, target occupant's age group, and guardian settings. For child occupants, risk assessment criteria can also be individually modified by incorporating the child occupant's historical behavioral data to reduce the problem of fixed rules failing to adapt to the different behavioral habits of child occupants.
[0230] Step S702: Determine whether the confidence level of the behavior meets the confidence level requirements corresponding to the risk assessment conditions.
[0231] Here, behavior confidence is used to represent the reliability of the behavior recognition model's judgment of the current behavior category. Different behavior categories can correspond to different confidence requirements. Behavior categories with lower risk levels can correspond to lower confidence requirements, while behavior categories with higher risk levels can correspond to higher confidence requirements, in order to avoid false triggering of high-risk warnings.
[0232] For example, the confidence level requirement for slight twisting behavior can be relatively low. Behaviors such as signs of seatbelt unfastening, signs of leaning out of the window, and signs of leaving a child seat may trigger stronger warnings or vehicle safety assistance controls, and therefore the confidence level requirement can be relatively high. For signs of leaning forward and tilting forward, the duration and amplitude of the movement should be considered together, rather than relying solely on the confidence level of a single behavior.
[0233] When determining the confidence level of a behavior, single-frame confidence, average confidence across multiple consecutive frames, highest confidence within a sliding time window, or weighted confidence are used. Weighted confidence is determined by combining image clarity, the completeness of human body key points, the reliability of in-vehicle target component recognition, and the output of the behavior recognition model. If the behavior confidence level does not meet the requirements, the behavior category may be temporarily excluded as a valid risk behavior, or subsequent image data may be collected for secondary confirmation. Confidence level assessment can reduce the probability of false alarms caused by image occlusion, short-term jitter, or changes in lighting.
[0234] Step S703: If the confidence level of the behavior meets the confidence level requirements, determine whether the trend of behavior development points to a preset dangerous component or a preset dangerous area.
[0235] Here, the pre-defined hazardous components include at least one of the following: seatbelt buckle, seatbelt latch, window, door handle, door locking area, and child safety seat boundary. The pre-defined hazardous areas include at least one of the following: window opening area, door opening area, outer area of child safety seat, seat edge area, and front passenger compartment gap area.
[0236] Behavioral trends are determined based on the positional changes of key human body points in consecutive frames of image data. If the hand key points gradually move closer to the seatbelt buckle in consecutive frames of image data, the behavioral trend is judged to be towards the seatbelt buckle. If the head or hand key points gradually move closer to the window edge, the behavioral trend is judged to be towards the window area. If the torso key points continuously move away from the child safety seat back and move forward of the vehicle, the behavioral trend is judged to be towards the danger zone in front of the vehicle. If the body contour gradually deviates from the child safety seat area, the behavioral trend is judged to be towards the outer area of the child safety seat.
[0237] When determining the trend of behavioral development, changes in direction, distance, and speed can be considered together. For example, even if a child occupant's hands are near the seatbelt, if the direction of hand movement is away from the seatbelt buckle, the behavioral trend cannot be considered to be pointing towards the seatbelt buckle. Similarly, if a child occupant's head briefly turns towards the window but the distance from the window edge does not continuously decrease, the behavioral trend cannot be considered to be pointing towards the window area. By judging the trend of behavioral development, it is possible to distinguish between normal actions and precursors to dangerous actions, thus improving the accuracy of risk assessment.
[0238] Step S704: If the behavioral trend points to a pre-set dangerous component or a pre-set dangerous area, the behavioral category is determined as a valid risky behavior.
[0239] Here, when the confidence level of a behavior meets the confidence requirements and the trend of the behavior points to a pre-set hazardous component or area, the behavior is classified as a valid risk behavior. A valid risk behavior indicates that the current behavior has met the pre-set risk assessment conditions, serving as the basis for determining the risk level of the behavior.
[0240] For example, when the behavior category is a premonition of seatbelt release, the behavior confidence level meets the corresponding confidence level requirement, and the hand key points are continuously close to the seatbelt buckle, the premonition of seatbelt release is identified as a valid risk behavior. When the behavior category is a premonition of leaning out of the window, the behavior confidence level meets the corresponding confidence level requirement, and the head key points or hand key points are continuously close to the window opening area, the premonition of leaning out of the window is identified as a valid risk behavior. When the behavior category is a premonition of leaning forward, the torso key points continuously move forward towards the vehicle and the leaning angle exceeds a preset angle, the premonition of leaning forward is identified as a valid risk behavior.
[0241] When determining valid risk behaviors, the duration of the action and the number of consecutive triggers can also be considered. If a behavior category appears only in a single frame or lasts only a very short time, it is not yet considered a valid risk behavior. If the same behavior category appears consecutively within a preset time window or repeats itself within a short period, the behavior category is determined to be a valid risk behavior. By screening for valid risk behaviors, distracting actions such as waving, brief head turns, and temporary adjustments to posture can be prevented from being mistaken for precursors to dangerous actions.
[0242] Step S705: Determine the risk level of the behavior based on the degree of risk corresponding to the effective risk behavior.
[0243] Here, the degree of risk is determined by the type of behavior corresponding to the effective risk behavior, the magnitude of the action, the duration of the action, the trend of the behavior, the distance between the target occupant and the dangerous parts or dangerous areas, and the vehicle status.
[0244] In one implementation, a basic risk level is pre-configured for different valid risk behaviors. For example, slight twisting corresponds to a low-risk level. Precursors to leaning forward and tilting correspond to a medium-risk level. Precursors to unfastening the seatbelt, leaning out of the window, and leaving the child seat correspond to a high-risk level. If multiple valid risk behaviors exist simultaneously, the final risk level can be determined based on the highest risk level, or a comprehensive score can be applied to all valid risk behaviors to determine the final risk level.
[0245] In another implementation, the risk level of the behavior is dynamically adjusted based on the development of the effective risky behavior. For example, when the child occupant's hand just begins to move towards the seatbelt buckle, it is determined to be a medium-risk level. When the child occupant's hand has already approached the seatbelt buckle and shows a pulling tendency, it is determined to be a high-risk level. When the child occupant's body leans slightly forward, it is determined to be a low-risk or medium-risk level. When the child occupant's body continues to lean forward and their center of gravity is significantly deviated from the child safety seat back, it is determined to be a high-risk level.
[0246] In some implementations, vehicle status can also be used to determine the behavioral risk level. When the vehicle is traveling at high speed, the windows are open, or the vehicle's cruise control is activated, the same valid risky behavior corresponds to a higher behavioral risk level. When the vehicle is stationary, the same valid risky behavior corresponds to a lower behavioral risk level.
[0247] In an optional implementation, step S704 identifies the behavior category as a valid risky behavior, including the following steps S801-S804.
[0248] Step S801: Within a preset time window, perform a consistency judgment on the behavior categories corresponding to multiple frames of image data.
[0249] Here, the preset time window is set based on the duration of common dangerous actions of the target occupant, the image acquisition frame rate, and the vehicle's computing power. For example, the preset time window can cover image frames within 0.5s to 3s, or it can correspond to 16, 32, or 64 consecutive frames of image data. By using the preset time window for judgment, the risk assessment is avoided from being directly triggered by misidentification results in a single frame image.
[0250] Consistency assessment is used to determine whether the behavioral categories corresponding to multiple frames of image data maintain the same or similar trends of change. For example, in multiple frames of image data, if the child occupant's hand key points all move towards the seatbelt buckle, and multiple frames of image data are all identified as signs of the seatbelt being released, then the behavioral category is considered consistent within a preset time window. If the child occupant's head key points and torso key points all continuously move towards the window area, and multiple frames of image data are all identified as signs of leaning out of the window, then the behavioral category is considered consistent within a preset time window.
[0251] Step S802: If multiple consecutive frames of image data correspond to the same behavior category within a preset time window, and the confidence levels of the behaviors corresponding to the same behavior category all meet the confidence level requirements, then the same behavior category is identified as a candidate risk behavior.
[0252] Here, "continuous frames" can refer to a number of consecutive frames or multiple frames within a preset time window that account for a preset proportion. For example, when the preset time window contains 32 frames of image data, it is required that at least 8 consecutive frames, 12 consecutive frames, or a preset proportion of image frames are identified as belonging to the same behavior category.
[0253] When multiple consecutive frames of image data correspond to the same behavior category, it is also necessary to determine whether the behavior confidence level corresponding to the same behavior category meets the confidence level requirements. Behavior confidence level can be the single-frame confidence level corresponding to each frame of image data, or it can be the average confidence level, weighted confidence level, or minimum confidence level corresponding to multiple consecutive frames of image data. When the behavior confidence level meets the confidence level requirements, it indicates that the behavior recognition result has high reliability.
[0254] For example, if multiple consecutive frames of image data consistently indicate that a child occupant's hand is close to the seatbelt buckle, and the confidence level of the behaviors corresponding to the premonitions of seatbelt release meets the confidence level requirements, then the premonitions of seatbelt release are identified as candidate risk behaviors. As another example, if multiple consecutive frames of image data consistently indicate that a child occupant's head or hand is close to the window area, and the confidence level of the behaviors corresponding to the premonitions of leaning out of the window meets the confidence level requirements, then the premonitions of leaning out of the window are identified as candidate risk behaviors.
[0255] Candidate risk behaviors are used to indicate that the current behavior already has risk characteristics, but it is still necessary to combine short-term interference filtering to determine whether the current behavior is a real risk behavior.
[0256] Step S803: Perform short-term interference filtering on candidate risky behaviors.
[0257] Here, short-term disruptive behaviors include actions that should not trigger effective risk assessment, such as a child passenger briefly waving, turning their head temporarily, briefly adjusting their sitting posture, being obscured by toys or clothing, instantaneous changes in posture caused by vehicle bumps, and recognition jumps caused by abnormal image acquisition.
[0258] Short-term interference filtering is based on the duration, direction, amplitude, target area, and trend of the action. For example, if a child occupant's hand briefly passes near the seatbelt buckle but the direction of movement does not consistently point towards the buckle, or if the hand leaves the buckle area in a very short time, this action is considered a short-term interference. If a child occupant's head briefly turns towards the window but the distance between their head and the window edge does not continuously decrease, this action is also considered a short-term interference. If a child occupant slightly adjusts their sitting posture and then quickly returns to a normal sitting posture, this action is also considered a short-term interference.
[0259] Short-term interference filtering can also be combined with deduplication and low-confidence filtering. For the same candidate risk behavior that recurs within the same preset time window, candidate risk behaviors with higher risk level or higher confidence are retained to avoid the same action triggering multiple risk results. For candidate risk behaviors with behavior confidence close to the lower limit of the confidence requirement, poor image quality, or many missing human key points, we continue to wait for confirmation in subsequent image frames, or reduce the risk weight of the candidate risk behavior.
[0260] In some implementations, short-term interference filtering can also combine the target occupant's historical behavioral habits with the vehicle's status. For example, if a child occupant frequently waves their hand during short periods of interaction, and the waving motion is not directed at dangerous components such as seat belts, windows, or doors, such actions are filtered as normal interactive behavior. When the vehicle is stationary and the doors are open, the risk level of certain body movement behaviors is reduced. When the vehicle is in motion, similar body movement behaviors retain a higher level of risk concern.
[0261] Step S804: If the candidate risk behavior is not identified as a short-term disruptive behavior, the candidate risk behavior is identified as a valid risk behavior.
[0262] Here, a valid risk behavior means that the target occupant's current action has met the requirements of continuous triggering, confidence level, and interference filtering, serving as the basis for determining the behavioral risk level.
[0263] For example, if a child occupant's hand continuously approaches the seatbelt buckle in multiple consecutive image frames, and the confidence level of the behavior corresponding to the premonition of seatbelt release meets the confidence requirement, and the hand movement is not judged as a brief wave or normal adjustment movement, then the premonition of seatbelt release is identified as a valid risk behavior. As another example, if a child occupant's head and torso continuously move towards the window area in multiple consecutive image frames, and the confidence level of the behavior corresponding to the premonition of leaning out of the window meets the confidence requirement, and the trend of the movement is not judged as a brief head turn, then the premonition of leaning out of the window is identified as a valid risk behavior.
[0264] Effective risk behaviors include at least one of the following: effective behavior category, effective behavior confidence level, change trajectory of key points on the human body, behavior duration, behavior development trend, and information on the target hazardous component.
[0265] In an optional implementation, when the target occupant is a child occupant, after step S701, the method further includes the following steps S901-S904.
[0266] Step S901: Obtain historical behavior data and historical warning data corresponding to child occupants.
[0267] Here, historical behavioral data refers to records of child occupants' behavior during past rides. This data includes at least one of the following: records of typical sitting postures, limb movement trajectories, hand range of motion, head range of motion, trunk tilt range, historical expressions of needs, historical normal interactive actions, and historical risky actions. Historical warning data consists of records generated after warnings were triggered during past rides. This data includes at least one of the following: warning trigger time, warning trigger scenario, corresponding behavioral category, corresponding behavioral risk level, warning output method, warning duration, warning cancellation result, and guardian feedback.
[0268] Historical behavioral data is retrieved from the occupant profiles corresponding to the child occupants. These profiles are linked to the child occupant's facial features, child safety seat binding information, and the guardian's or vehicle user's account. Historical behavioral data can be stored chronologically or categorized by riding stage, age group, behavioral type, or risk level. For example, historical behavioral data can be categorized by the stage of vehicle stationary travel, low-speed vehicle travel, high-speed vehicle travel, and long-term riding. It can also be categorized by the child occupant's age group.
[0269] Historical warning data is used to reflect the risk trigger patterns corresponding to past risky behaviors of child occupants. For example, child occupants are more likely to lean forward after long journeys, or they are more likely to lean towards the window when near it. By acquiring historical warning data, subsequent risk assessments can be made not only based on fixed rules, but also by considering the individual behavioral habits of the child occupants.
[0270] Step S902: Based on historical behavior data and historical warning data, determine the individualized risk assessment parameters for child occupants; wherein, the individualized risk assessment parameters are used to characterize the risk assessment standards for child occupants at different stages of riding, at different ages, and with different historical behavioral habits.
[0271] Here, the individualized risk assessment parameters include at least one of the following: individualized confidence threshold, individualized distance threshold, individualized angle threshold, individualized duration threshold, individualized consecutive frame count threshold, individualized danger zone range, and individualized risk level correction parameters.
[0272] In one implementation, the typical range of movements of child occupants during historical riding is statistically analyzed and used as a reference basis for individualized risk assessment. The typical range of movements includes the range of trunk tilt angle in a normal sitting posture, the range of hand movements during normal interaction, the range of head movements during normal head turning, and the duration of normal posture adjustments. If the current behavioral status falls within the typical range of movements, the risk level corresponding to the current behavior is reduced or the continuous confirmation time is extended. If the current behavioral status significantly exceeds the typical range of movements, the level of risk attention corresponding to the current behavior is increased.
[0273] In another approach, risky movement characteristics of child occupants are statistically analyzed from historical warning data, and these characteristics are used as a basis for adjusting individualized risk assessments. These risky movement characteristics include hand trajectories corresponding to historical warning signs of seatbelt release, head movement trajectories corresponding to historical warning signs of leaning out of the window, trunk tilt changes corresponding to historical warning signs of leaning forward, and body contour shifts corresponding to historical warning signs of leaving the child seat. If the current behavioral trend is similar to historical risky movement characteristics, the sensitivity of risk assessment corresponding to the current behavior category is increased.
[0274] Individualized risk assessment parameters can also be determined based on the child occupant's age. Younger children have a weaker ability to understand safety cues, so behaviors such as warning signs of seatbelt unfastening, leaning out of the window, or leaving the child seat can be assessed using parameters with higher sensitivity. Older children may have a wider range of normal movement, so behaviors such as leaning forward or slight tilting can be assessed by considering both duration and direction of movement to avoid frequently triggering warnings.
[0275] Step S903: Based on individualized risk assessment parameters, adjust the risk assessment conditions, confidence requirements, preset hazardous components and / or preset hazardous areas corresponding to the behavior category.
[0276] Here, the risk assessment criteria include distance requirements, angle requirements, duration requirements, number of consecutive frames requirements, behavioral development trend requirements, and vehicle status requirements. By adjusting the risk assessment criteria, the risk assessment rules can be matched with the actual behavioral habits of child occupants.
[0277] For example, if historical behavioral data shows that child occupants have a wide range of hand movements during normal interactions, but these movements typically don't approach the seatbelt buckle, the normal range of hand movements should be appropriately expanded while maintaining high sensitivity around the seatbelt buckle. If historical warning data shows that child occupants repeatedly bring their hands close to the seatbelt buckle and trigger warnings, the distance threshold or the number of consecutive frames required for the seatbelt release warning should be lowered to allow for earlier identification of the warning. If historical behavioral data shows that child occupants frequently look down briefly at toys but quickly return to a normal sitting posture, the duration requirement for the forward lean warning should be appropriately increased to reduce false warnings caused by normal movements.
[0278] Individualized risk assessment parameters can also be used to adjust confidence requirements. For historically frequent behavior categories with high security risks, the confidence requirement can be appropriately lowered to make the corresponding risky behaviors easier to identify in advance. For historically frequently falsely triggered behavior categories, the confidence requirement can be appropriately increased, or a continuous frame confirmation requirement can be added. Through individualized adjustments to the confidence requirement, both the timeliness and accuracy of early warnings can be balanced.
[0279] Individualized risk assessment parameters can also be used to adjust preset hazardous components and preset hazardous areas. Different child occupants have different heights, arm spans, sitting postures, and child safety seat installation positions, so the hazardous area corresponding to the same target component inside the vehicle can also differ. For example, taller child occupants or those with longer arm spans correspond to a larger hazardous area near the window or the seatbelt buckle; shorter child occupants require a re-determined hazardous area based on the child safety seat position and actual reachability. By adjusting preset hazardous components and preset hazardous areas, the assessment of behavioral development trends is made more consistent with the actual reachability of the child occupant.
[0280] Step S904: Based on the adjusted risk assessment criteria, behavior category, behavior confidence level, and behavior development trend, determine the behavioral risk level of the child occupant.
[0281] Here, the adjusted risk assessment criteria serve as the current risk assessment rules specific to child occupants. They are used to determine whether the behavior category meets the requirements for valid risk behavior and to further determine whether it is low-risk, medium-risk, or high-risk.
[0282] In one implementation, it is first determined whether the behavior confidence level meets the adjusted confidence level requirement, and then it is determined whether the behavior trend points to the adjusted preset hazardous component or preset hazardous area. If the behavior confidence level meets the adjusted confidence level requirement, and the behavior trend points to the adjusted preset hazardous component or preset hazardous area, the behavior risk level is determined based on the risk level corresponding to the behavior category. If the behavior confidence level does not meet the adjusted confidence level requirement, or the behavior trend does not point to the adjusted preset hazardous component or preset hazardous area, a lower behavior risk level is maintained, or subsequent image data is collected for confirmation.
[0283] For example, for child occupants who have a history of frequently attempting to touch the seatbelt buckle, the adjusted risk assessment criteria now classify any hand movement near the seatbelt buckle as high-risk earlier. For child occupants who have a history of frequently triggering window warnings due to normal hand waving, the adjusted criteria require the head or torso to move simultaneously towards the window area before classifying any signs of leaning out of the window as medium or high-risk. Thus, different child occupants require different risk assessment standards.
[0284] After determining the behavioral risk level, the risk level, corresponding behavioral category, adjusted risk assessment criteria, early warning output, and subsequent feedback from the guardian are recorded in the child passenger's passenger file. When individualized risk assessment parameters are subsequently determined again, the newly added records are used to update the parameters. Through continuous accumulation and updating, individualized risk assessment parameters can be gradually optimized as the child passenger grows older, their riding habits change, and guardian feedback evolves, thereby improving the accuracy and adaptability of child passenger behavioral risk level assessment.
[0285] In an optional implementation, step S902 includes the following steps S1001-S1005.
[0286] Step S1001: Statistically analyze the routine movement trajectories and historical risk movement trajectories of child occupants during historical riding.
[0287] Here, during multiple vehicle rides by child occupants, the system continuously records their behavioral status, behavior recognition results, behavioral risk levels, and warning trigger results. This information is stored as historical behavioral data and historical warning data in the child occupant's corresponding occupant file. The occupant file is linked to the child occupant's identification, child safety seat binding information, or guardian's account, so that historical data can be retrieved when the child occupant re-enters the vehicle.
[0288] Routine movement trajectories are the normal movement patterns of child occupants when no effective risky behavior is triggered. These trajectories include head movements in a normal sitting posture, hand movements, slight trunk swaying, normal head turning, normal waving, normal looking down at toys, and brief adjustments to sitting posture. Routine movement trajectories are used to characterize the natural movement habits of child occupants in a safe riding position.
[0289] Historical risk action trajectories are the movement patterns of child occupants during past riding experiences when warnings are triggered or when their actions are identified as valid risk behaviors. These trajectories include: hands gradually moving towards the seatbelt buckle; hands pulling on the seatbelt; heads or hands gradually moving towards the window area; the torso continuously moving forward of the vehicle; the body contour gradually deviating from the child safety seat area; and the body's center of gravity continuously shifting towards the door.
[0290] When analyzing routine and historical risk-related movement trajectories, historical data should be categorized by behavior type, riding stage, vehicle status, and age group. For example, movement trajectories can be analyzed separately for stationary, low-speed, and high-speed vehicle riding, as well as for short-duration and long-duration riding stages. This categorized analysis allows for a more accurate reflection of child occupants' behavioral habits in different scenarios.
[0291] Step S1002: Determine the range of routine behaviors corresponding to the child occupant based on the routine movement trajectory.
[0292] Here, the range of normal behavior refers to the range of motion, direction of motion, area of activity, and duration of movement that a child occupant typically exhibits when safely in a riding position.
[0293] Routine behavioral ranges include at least one of the following: routine hand movement range, routine head rotation range, routine trunk tilt range, routine changes in body center of gravity range, duration of normal sitting posture adjustment range, and range of normal interactive movements. For example, based on a child occupant's historical normal hand-waving trajectory, determine the routine range of hand movement when not near seatbelt buckles, windows, or door handles. Based on a child occupant's historical normal head-turning trajectory, determine the routine range of head rotation angle and head movement distance. Based on a child occupant's historical normal sitting posture adjustment trajectory, determine the routine range of short-term forward or sideways trunk tilting.
[0294] When determining the scope of routine behaviors, abnormal historical data is excluded. For example, data generated when the image quality is poor, many key human figures are missing, the vehicle is experiencing severe vibrations, or the child occupant's image is obscured, will not be used as the primary basis for determining the scope of routine behaviors. For actions that occur repeatedly but do not trigger warnings, the weight of those actions in the scope of routine behaviors is increased. Based on this, the scope of routine behaviors can more closely reflect the actual safe activity habits of child occupants.
[0295] Step S1003: Determine the range of risky behaviors corresponding to child occupants based on historical risky action trajectories.
[0296] Here, the scope of risky behavior refers to the area, direction, magnitude, and duration of actions in which a child passenger is likely to trigger a risk during a past ride.
[0297] The risk range includes at least one of the following: seatbelt buckle risk range, window risk range, door risk range, child safety seat outer side risk range, forward leaning risk range, and body tilting risk range. For example, if historical data shows that a child occupant repeatedly reaches for the seatbelt buckle, the seatbelt buckle risk range is determined based on the movement trajectory and distance changes of the hand before approaching the seatbelt buckle. If historical data shows that a child occupant repeatedly leans towards the window, the window risk range is determined based on the historical trajectories of the head, hands, and torso approaching the window area.
[0298] The scope of risky behavior is determined based on the premonitory signs preceding the risky action, not just the final position after the dangerous action is completed. For example, in a seatbelt release scenario, the area where the hand gradually moves from the inside of the body to the middle area near the seatbelt buckle is defined as part of the scope of risky behavior. In this way, during subsequent real-time identification, even if the child occupant has not actually touched the seatbelt buckle, the risk can be judged in advance based on the movement trajectory.
[0299] When determining the scope of risky behaviors, historical warning cancellation results can also be considered. If a historical risky action quickly stops after a warning is issued, it indicates that the area and direction of the action have high warning value. If a historical warning is reported as a false alarm by the ward, the weight of the data corresponding to that historical risky action in the scope of risky behaviors should be reduced to avoid repeated misjudgments in the future.
[0300] Step S1004: Based on the difference between the scope of normal behavior and the scope of risky behavior, generate individualized risk assessment parameters.
[0301] Here, the differences include at least one of the following: differences in activity area, differences in movement direction, differences in distance change, differences in angle change, differences in action duration, and differences in action frequency.
[0302] Individualized risk assessment parameters include at least one of the following: individualized distance threshold, individualized angle threshold, individualized duration threshold, individualized consecutive frame count threshold, individualized confidence threshold, individualized hazard area range, and individualized risk level correction parameters. For example, when a large distance is maintained between the normal behavioral range and the seatbelt buckle, but historical risky action trajectories repeatedly enter the area near the seatbelt buckle, an individualized distance threshold corresponding to the pre-release warning signs is generated based on the boundary difference between the two ranges. When a child occupant's torso tilt range is small in a normal sitting posture, but the torso tilt angle significantly increases during historical risky actions, an individualized angle threshold corresponding to the pre-tilt warning signs is generated based on the angle difference.
[0303] When generating individualized risk assessment parameters, the parameters should cover both the normal activity habits of child occupants and be able to identify their historical risk trends in advance. If the boundary between the range of normal behavior and the range of risky behavior is clear, a relatively clear threshold should be used for judgment. If there is some overlap between the range of normal behavior and the range of risky behavior, a comprehensive assessment parameter should be generated by combining the direction of the action, the duration of the action, and the confidence level of the behavior.
[0304] Step S1005: Upon receiving feedback from the guardian regarding the early warning information, update the individualized risk assessment parameters based on the feedback.
[0305] Here, after issuing the warning information, feedback from the guardian is received. The feedback is input by the guardian via the in-vehicle display device, mobile terminal, voice command, or physical button. Feedback can include at least one of the following: warning accurate, warning too early, warning too late, false alarm, missed warning, no reminder needed, and need for enhanced reminder.
[0306] When the feedback result indicates that the warning is accurate, the current individualized risk assessment parameters are retained, and the weight of the corresponding behavioral trajectory in the historical risk action trajectories is increased. When the feedback result indicates that the warning is a false alarm, the weight of the corresponding behavioral trajectory in the risk behavior range is reduced, or the confidence threshold, duration threshold, or consecutive frame number threshold of the corresponding behavior category is increased. When the feedback result indicates that the warning is too late, the corresponding danger zone is expanded, or the distance threshold and duration threshold of the corresponding behavior category are reduced, so that subsequent similar risk actions can be identified earlier. When the feedback result indicates that a stronger warning is needed, the risk level correction parameter of the corresponding behavior category is increased, so that subsequent similar behaviors trigger a stronger warning.
[0307] When updating individualized risk assessment parameters, the feedback time, the identity of the person providing the feedback, the corresponding warning information, the corresponding behavioral category, and the updated parameter value should be recorded simultaneously. Multiple feedback results can form a feedback history. For example, if a guardian repeatedly marks the same type of forward leaning warning as a false alarm, the duration requirement for that type of behavior should be gradually increased. If a guardian repeatedly marks the seatbelt release warning as accurate, the warning sensitivity for that type of behavior should be maintained or enhanced.
[0308] In an optional implementation, in step S105, based on the virtual assistant's image and behavioral risk level, interactive information corresponding to the target occupant is output, including the following steps S1101-S1105.
[0309] Step S1101: Based on the behavioral state information, determine the target interaction scenario corresponding to the target occupant.
[0310] Here, the target interaction scenarios include at least one of the following: daily companionship scenarios, comforting scenarios, entertainment interaction scenarios, demand response scenarios, safety reminder scenarios, and early warning intervention scenarios.
[0311] Behavioral status information includes the target occupant's posture, body movements, facial expressions, expressions of need, warning signs of dangerous actions, and duration of the ride. For example, if the target occupant maintains a normal sitting posture and does not exhibit any risky actions, the target interaction scenario is identified as a daily companionship scenario. If the target occupant cries, frowns, or frequently sways their body, the target interaction scenario is identified as a soothing scenario. If the target occupant smiles, gazes at the display area, or actively engages in interactive actions, the target interaction scenario is identified as an entertainment interaction scenario. If the target occupant points to a water cup, toy, or a specific item inside the vehicle, the target interaction scenario is identified as a need response scenario. If the target occupant exhibits signs of unfastening their seatbelt, leaning out of the window, or leaning forward significantly, the target interaction scenario is identified as a safety reminder or early warning intervention scenario.
[0312] The target interaction scenario can be determined based on a single behavioral state information or a combination of multiple behavioral state information. For example, if the target occupant is leaning forward while maintaining a relatively calm expression, it is prioritized as a safety reminder scenario. If the target occupant is crying but shows no signs of impending danger, it is prioritized as a reassurance scenario.
[0313] Step S1102: Based on the target interaction scenario and physiological characteristic information, determine the interaction content that matches the target occupant.
[0314] Here, physiological characteristics represent the target occupant's age group, physical appearance, developmental stage, and identity information. Target occupants of different ages, developmental stages, and interaction preferences correspond to different interaction content.
[0315] When the target passenger is a child, appropriate content is selected based on the child's age group. The interactive content is also determined in conjunction with the content preferences preset by the guardian, such as nursery rhyme type, story type, voice style, and virtual assistant avatar action style.
[0316] In everyday companionship scenarios, interactive content includes nursery rhymes, stories, greetings, game-like interactions, and companion-themed animations. In soothing scenarios, interactive content includes calming music, gentle voice prompts, smiling emoticons, and slow movements. In demand response scenarios, interactive content includes demand confirmation voice prompts and guardian prompts. In safety reminder scenarios, interactive content includes gentle reminder voice prompts, safety action illustrations, and warning gestures corresponding to the virtual assistant's avatar. In early warning and intervention scenarios, interactive content includes clearer voice reminders, more prominent visual cues, and stronger early warning linkage information.
[0317] By combining the target interaction scenario with physiological characteristic information, it is possible to avoid using fixed scripts or content to output uniformly to all target passengers, so that the interaction content has both scenario adaptability and individual adaptability.
[0318] Step S1103: Control the virtual assistant image to output corresponding voice interaction information and / or visual interaction information according to the interaction content.
[0319] Here, voice interaction information is played through dedicated rear speakers, the car audio system, or other voice output devices. Visual interaction information is output through the rear display screen, the central control display screen, or a split-screen display area within the central control display screen.
[0320] Voice interaction information includes greetings, nursery rhymes, stories, soothing voices, request confirmation voices, safety reminder voices, and danger warning voices. Visual interaction information includes the virtual assistant's facial expressions, postures, actions, animations, text prompts, and safety action illustrations. The virtual assistant can display different facial expressions and actions based on the interaction content. For example, in everyday companionship scenarios, the virtual assistant displays a smiling expression and waves. In soothing scenarios, the virtual assistant displays a gentle expression and pats lightly. In safety reminder scenarios, the virtual assistant displays a worried expression, waves its hand to stop, points to the seatbelt, or gestures to sit up straight.
[0321] When outputting interactive information, the intensity of the interaction is adjusted according to the level of behavioral risk. When the level of behavioral risk is low, the virtual assistant uses a gentle tone to output reminders. When the level of behavioral risk is high, the virtual assistant uses a clearer tone of voice and visual warning actions to output interactive information. By using the virtual assistant to convey both voice and visual content, the target occupant can more intuitively understand the current interaction intent or safety reminder intent.
[0322] Step S1104: During the process of outputting voice interaction information and / or visual interaction information, continue to acquire the facial expression status information of the target occupant.
[0323] Here, facial expression status information includes at least one of the following: smiling, crying, frowning, staring, avoiding, open-mouthed, and emotionally stable. Facial expression status information is determined with the assistance of facial key points, changes in the corners of the mouth, eye state, facial orientation, head movements, and vocal input.
[0324] The purpose of continuing to acquire facial expression information is to determine whether the current interaction is effective. For example, if playing a nursery rhyme causes the target occupant to change from a crying state to a calm or smiling state, it indicates that the current interaction has a calming effect. If telling a story causes the target occupant to continue looking at the display area or smile, it indicates that the current interaction is attractive. If a safety reminder is output, the target occupant stops dangerous actions and returns to a normal sitting posture, it indicates that the reminder is effective.
[0325] In some implementations, facial expression information can be combined with voice information for joint judgment. For example, when the target occupant cries, laughs, or responds verbally, combining the voice information with facial expressions yields a more accurate emotion assessment. By continuously collecting feedback during the interaction, the interactive content can no longer be a one-way output but can be dynamically adjusted based on the target occupant's reactions.
[0326] Step S1105: Adjust the interactive content based on facial expression status information.
[0327] Here, the adjustment methods include switching content types, changing voice tone, changing playback volume, changing the virtual assistant's facial expressions, changing the virtual assistant's actions, extending the current interactive content, or stopping the current interactive content.
[0328] When facial expression information indicates that the target occupant is interested in the current interactive content, continue outputting the current interactive content, or output continuation content related to the current interactive content. When facial expression information indicates that the target occupant is not interested in the current interactive content, switch to other interactive content. When facial expression information indicates that the target occupant is emotionally unstable, adjust the interactive content to a reassuring level. When facial expression information indicates that the target occupant has calmed down or has stopped dangerous actions, reduce the intensity of the interaction or end the reminder content.
[0329] In an optional implementation, step S1105 includes the following steps S1201-S1203.
[0330] Step S1201: If the facial expression information indicates that the target passenger is in a crying state, adjust the interaction content to a soothing interaction content.
[0331] Here, when facial expression information indicates that the target occupant is crying, the current interaction content is switched to a soothing interaction content. The crying state is determined based on a comprehensive assessment of the target occupant's mouth opening, eyebrow and eye expression, facial tension, head movement, and vocal information. When the target occupant is a child, the crying state can also be determined by considering the child's travel time, ambient noise, vehicle interior temperature, and past soothing preferences.
[0332] Soothing interactive content includes soothing nursery rhymes, calming music, gentle voice prompts, short comforting phrases, slow animations, and the gentle expressions of a virtual assistant avatar. The virtual assistant avatar can display a smiling expression, a gentle wave, or soothing gestures, and play soothing content via a voice output device. If the target occupant has a historical preference for a particular type of soothing content, the content corresponding to that preference will be selected first. If the target occupant remains crying after continuous output of soothing interactive content, a prompt message will be generated and sent to the in-vehicle display device to remind the driver or guardian to pay attention to the target occupant's condition.
[0333] Step S1202: If the facial expression information indicates that the target occupant is in a pleasant state, maintain the current interaction content or output a continuation interaction content related to the current interaction content.
[0334] Here, when facial expression information indicates that the target occupant is in a pleasant state, the current interaction content can be maintained, or continuation interaction content related to the current interaction content can be output. The pleasant state can be determined based on the target occupant's smiling expression, gaze direction, relaxed body state, laughter, or active interaction actions.
[0335] For example, if the target passenger smiles while a nursery rhyme is playing, the current nursery rhyme or a similar nursery rhyme can be played. If the target passenger continues to look at the display area while listening to a story, the story can continue. If the target passenger responds positively to the virtual assistant avatar, the current interaction time can be extended, or questions and answers related to the current interaction topic can be provided. When the target passenger is a child, the content can be tailored to the age group in terms of difficulty and duration, avoiding overly long or comprehensible content.
[0336] While maintaining the current interactive content or continuing the interactive content, the behavioral status information of the target occupant can be monitored. If the target occupant exhibits warning signs of dangerous actions during pleasant interaction, the system can switch to safety reminders or warnings first.
[0337] Step S1203: If the behavioral state information indicates that the target occupant has a demand expression action, determine the demand prompt information based on the demand expression action, and send the demand prompt information to the vehicle display device.
[0338] When behavioral status information indicates that the target occupant is expressing a need, the corresponding need prompt information is determined based on this expression. Expressions of need include pointing to a water cup, pointing to a toy, pointing to the window, pointing to the display screen, waving, nodding, shaking the head, patting the seat, or expressing a need through facial expressions and gestures. When the target occupant is a child, these expressions are used to help determine whether the child needs water, a toy, comfort, a change in posture, or attention from a guardian.
[0339] When determining the required information, the system considers the target occupant's hand position, finger direction, gaze direction, target object location, and historical needs records. For example, if a child occupant points to a water cup and looks towards it, a required information message is generated indicating that the child may need water. If a child occupant points to a toy or repeatedly gazes at the toy area, a required information message is generated indicating that the child may need a toy. If a child occupant is waving their hand continuously and crying, a required information message is generated indicating that the child needs attention from a guardian.
[0340] The system sends notification messages to in-vehicle displays, including the central control screen, instrument panel display area, or rear-seat displays. These notifications can be displayed using text, icons, voice, or images. For example, the central control screen might display messages such as "Child occupant may need water," "Child occupant may need comforting," or "Child occupant is requesting help." If necessary, the notifications can also be sent to the guardian's terminal.
[0341] The interaction method provided in this application can dynamically adjust the interaction content according to the target occupant's crying state, happy state, and expression of needs, and promptly prompt the driver or guardian of the target occupant's needs. This application enables smart cockpit interaction to expand from fixed content playback to adaptive interaction based on real-time feedback from the target occupant, improving the interaction experience and care efficiency.
[0342] In an optional implementation, in step S105, based on the virtual assistant's image and behavioral risk level, a warning message corresponding to the target occupant is output, including the following steps S1301-S1303.
[0343] Step S1301: When the behavioral risk level is low, control the virtual assistant image to output the first warning information.
[0344] Here, when the behavioral risk level is low, it indicates that the target occupant's behavior has become slightly abnormal, but this behavior has not yet posed a significant safety risk. A low-risk level might correspond to slight body twisting, briefly deviating from a normal sitting posture, briefly looking down, briefly raising an arm, or briefly turning the head. In this case, the virtual assistant will output a first warning message, gently prompting the target occupant to adjust their behavior.
[0345] The first warning message includes at least one of the following: voice reminder, text prompt, facial expression prompt, and lightweight animated prompt. For example, when a child passenger slightly shifts their body, the virtual assistant displays a smiling expression and outputs a gentle prompt such as "It's safer to sit still" through a dedicated rear speaker. The tone of the first warning message is relatively soft, the output volume is lower than the high-risk warning volume, and the displayed content is presented in a way that does not affect the driver's attention. Through the first warning message, child passengers can be guided back to a safer sitting posture without causing them anxiety.
[0346] In step S1302, when the behavioral risk level is medium risk, the virtual assistant image is controlled to output a second warning message, and the in-vehicle prompting device is controlled to output auxiliary prompting information.
[0347] Here, when the behavioral risk level is medium risk, it indicates that the target occupant's behavior has deviated from normal seating posture and there is a possibility of it developing into dangerous behavior. Medium risk level corresponds to behaviors such as significant body tilting, continuous forward leaning, head continuously approaching the window, and hands continuously approaching target components inside the vehicle but not yet touching dangerous locations. At this time, a second warning message and auxiliary prompts are simultaneously output to increase the target occupant's and driver's awareness of the current risk situation.
[0348] The second warning message is output through a virtual assistant. The virtual assistant displays a worried expression, makes a reminder gesture, or indicates that the passenger should sit up straight, and plays a clearer reminder statement through a voice output device. For example, when a child passenger leans forward significantly, the virtual assistant can output a reminder such as "Please sit back in your seat," and display a gesture indicating that the passenger should look up and sit up straight on the display device.
[0349] In-vehicle alert devices include at least one of the following: interior lighting devices, in-vehicle display devices, instrument panel alert devices, or rear-seat displays. Auxiliary alert information includes flashing lights, icon alerts, text alerts, or audible alerts. Auxiliary alert information is used to remind the driver or guardian to pay attention to the status of a target occupant. For example, when a child occupant's body is continuously leaning towards the window, the in-vehicle display device may show a message such as "Abnormal sitting posture of rear-seat occupant," and the interior lighting devices may flash briefly. By combining secondary warning information and auxiliary alert information, warnings and interventions can be completed before the risk escalates further.
[0350] Step S1303: When the behavioral risk level is high, extract the specific dangerous action features corresponding to the behavioral status information, control the virtual assistant image to present a warning posture corresponding to the specific dangerous action features, output the third warning information, and control the in-vehicle alarm device to output danger warning information to the driver.
[0351] Here, when the behavioral risk level is high, it indicates that the target occupant's behavior already poses a significant safety risk, or that the behavioral trend is pointing towards a pre-set dangerous component or area. High risk levels correspond to behaviors such as signs of seatbelt release, signs of leaning out of the window, severe forward leaning, severe tilting, signs of leaving a child seat, pulling on the seatbelt, approaching the door handle, or approaching the window opening area. In this case, it is necessary to extract specific dangerous action characteristics from the behavioral status information and output stronger warning information based on these characteristics.
[0352] Specific dangerous action characteristics include at least one of the following: dangerous action category, trajectory of changes in key human body points, target dangerous component, target dangerous area, direction of action development, duration of action, and confidence level of action. For example, specific dangerous action characteristics corresponding to warning signs of seatbelt release include: hand key points continuously approaching the seatbelt buckle; the distance between the hand and the seatbelt buckle being less than a preset distance; and abnormal changes in the position of the seatbelt webbing. Specific dangerous action characteristics corresponding to warning signs of leaning out of the window include: head key points or hand key points continuously approaching the window edge; and key human body points moving towards the window opening area. Specific dangerous action characteristics corresponding to warning signs of leaning forward include: torso key points moving away from the seat back; the body's center of gravity shifting forward towards the vehicle; and the leaning angle exceeding a preset angle.
[0353] The virtual assistant's avatar displays corresponding warning gestures based on the specific characteristics of the dangerous actions. These gestures include a worried expression, a waving gesture to stop the movement, pointing towards the seatbelt, indicating to sit up straight, indicating to move away from the window, or a serious reminder expression. For example, when the specific dangerous action corresponds to a hand near the window, the virtual assistant displays a worried expression and waves to stop the movement; when the specific dangerous action corresponds to a hand near the seatbelt buckle, the virtual assistant displays a serious reminder expression and points towards the seatbelt; when the specific dangerous action corresponds to leaning forward or bending down to pick up an object, the virtual assistant displays an anxious expression and makes a gesture to sit up straight.
[0354] The third warning information includes at least one of the following: a strong voice alert, a prominent visual alert, a text description of dangerous actions, and a high-priority audio alert. The in-vehicle alarm device may include at least one of the following: a vehicle audio system, a vehicle display device, a dashboard display device, a seat vibration device, and in-vehicle lighting devices. When the in-vehicle alarm device outputs a danger warning information to the driver, it plays a danger warning voice message through the vehicle audio system, displays the text description of dangerous actions through the vehicle display device, outputs a warning icon through the dashboard display device, or transmits a vibration reminder to the driver through the seat vibration device. Through the third warning information and the danger warning information, the driver can be promptly aware of any high-risk behavior by a target occupant and intervene manually in a timely manner.
[0355] In an optional implementation, after controlling the in-vehicle alarm device to output a danger warning message to the driver in step S1303, the method further includes the following steps S1401-S1403.
[0356] Step S1401: Obtain the status change results of the target occupant's behavior status information after the output of the danger warning information.
[0357] Here, after outputting the hazard warning message, image data of the target occupant is continuously acquired through an image acquisition device, and the occupant's behavioral state information is extracted again. By comparing the behavioral state information before and after the hazard warning message output, the state change result can be obtained. The state change result is used to determine whether the warning message has an effective intervention effect.
[0358] The results of the status change include whether the behavioral risk level has decreased, whether the dangerous action has stopped, whether the key points of the body have moved away from the dangerous parts, whether the target occupant has returned to a normal sitting posture, whether the premonition of seatbelt release has disappeared, whether the premonition of leaning out of the window has disappeared, whether the degree of forward leaning has decreased, and whether the target occupant has left the danger zone. For example, after outputting the seatbelt release warning message, if the child occupant's key points of hand move away from the seatbelt buckle and the premonition of seatbelt release no longer appears continuously, the status change result indicates that the behavioral risk level has decreased. Conversely, after outputting the premonition of leaning out of the window warning message, if the child occupant's key points of head remain close to the window edge, the status change result indicates that the behavioral risk level has not decreased.
[0359] When acquiring results of status changes, an observation time window is set. The observation time window starts counting from the moment the hazard warning information is output and continues for a preset duration. The preset duration is determined based on the behavior category and warning intensity. For example, a shorter observation time window corresponds to signs of seatbelt unfastening, while a slightly longer observation time window corresponds to signs of body tilting. By using the observation time window, it is possible to avoid unstable results caused by making judgments immediately after the warning information is output.
[0360] Step S1402: If the state change result indicates a decrease in the behavioral risk level of the target occupant, reduce the output intensity of the warning information or stop outputting the warning information.
[0361] Here, when the change in status indicates a decrease in the behavioral risk level of the target occupant, the output intensity of the warning information is reduced or the warning information is stopped. A decrease in behavioral risk level indicates that the target occupant has stopped dangerous actions, has moved away from dangerous parts, has returned to a normal sitting posture, or has returned to the safe area of the child safety seat.
[0362] Reducing the intensity of warning messages includes lowering the volume of voice prompts, stopping flashing interior lights, stopping seat vibrations, switching strong warning voice prompts to gentle reminder voice prompts, switching high-risk warning icons to normal warning icons, or reducing the frequency of prompts. Stopping the output of warning messages includes stopping the playback of hazard warning voice prompts, turning off hazard warning content on the in-vehicle display, stopping the warning gesture display of the virtual assistant avatar, and restoring the virtual assistant avatar to its normal interactive state.
[0363] In some implementations, confirmation messages are output after the level of behavioral risk has decreased. For example, the virtual assistant avatar may gently prompt the target occupant to maintain a safe posture using a soft voice, or indicate that the dangerous behavior has been eliminated using a smiling emoji. By reducing or stopping warnings after the risk has decreased, the interference caused by continuous alarms to the target occupant and driver can be reduced, and the comfort of the in-vehicle interaction experience can be maintained.
[0364] In step S1403, if the risk level of the target occupant's behavior as represented by the state change result has not decreased, the output intensity of the warning information is increased, and the vehicle is controlled to perform auxiliary safety control operations.
[0365] Here, when the state change result indicates that the risk level of the target occupant's behavior has not decreased, it means that the target occupant's dangerous behavior is still ongoing, or that the trend of the target occupant's actions is still pointing towards dangerous parts or dangerous areas. In this case, the output intensity of the warning information can be increased, and the vehicle can be controlled to perform auxiliary safety control operations.
[0366] Enhancing the output of warning information includes increasing the volume of voice prompts, increasing the frequency of prompts, extending the duration of prompts, increasing the flashing frequency of interior lights, increasing the intensity of seat vibrations, displaying more prominent text prompts on in-vehicle displays, outputting high-priority warning icons through the instrument panel, or simultaneously sending the prompt information to the guardian's terminal. The virtual assistant's avatar can also switch from a normal warning posture to a more obvious deterrent action or a serious reminder action.
[0367] Assisted safety control operations include at least one of the following: lowering the volume of the in-vehicle multimedia system, controlling the output of enhanced warning lights from the interior lighting system, controlling the output of vibration warnings from the seat vibration system, controlling the output of text warnings from the in-vehicle display, sending warning messages to the guardian's terminal, locking the rear windows, restricting the unlocking of the rear doors, and suspending cruise control when it is engaged. These assisted safety control operations can be selected and executed based on the vehicle configuration and current vehicle status. For example, if a child passenger shows signs of leaning out of the window and the window is open, the rear windows can be closed or their opening restricted. If a child passenger engages in high-risk behavior and the in-vehicle multimedia system is loud, the volume can be lowered to ensure the driver can clearly receive hazard warnings. If cruise control is engaged and the child passenger's high-risk behavior persists, cruise control can be suspended to allow the driver to take over the vehicle.
[0368] After executing auxiliary safety control operations, the system continues to acquire behavioral status information of the target occupant and reassess whether the behavioral risk level has decreased. If the behavioral risk level has not decreased, a high-intensity warning is maintained or a manual intervention prompt is issued. If the behavioral risk level has decreased, the intensity of the warning information output is reduced or the warning information output is stopped. Through cyclical feedback processing, the in-vehicle warning process can be dynamically adjusted according to the subsequent changes in the target occupant's actions, rather than only issuing a prompt once when the risk is first identified.
[0369] The interactive early warning method provided in this application can continuously monitor changes in the behavior of the target occupant after outputting a danger warning message, and dynamically adjust the warning intensity and auxiliary safety control operations according to whether the risk has been eliminated. This application can improve the continuity and effectiveness of intervention for high-risk behaviors, prevent the target occupant from continuing to remain in a dangerous state when the first warning is ineffective, thereby further improving the proactive safety protection capability of the smart cockpit for the target occupant.
[0370] Based on this, the embodiments of this application provide an interactive early warning system for an intelligent cockpit, referring to... Figure 6 The interactive early warning system provided in this application includes: Image data acquisition module 21 is used to acquire image data of the target occupant in real time through an image acquisition device.
[0371] The data processing module 22 is used to extract physiological characteristic information and behavioral status information of the target occupant based on image data.
[0372] The virtual assistant generation module 23 is used to generate or update the virtual assistant image corresponding to the target occupant based on physiological characteristic information.
[0373] Risk assessment module 24 is used to determine the behavioral risk level of the target occupant based on behavioral status information.
[0374] The interactive warning module 25 is used to output interactive information and / or warning information corresponding to the target occupant based on the virtual assistant's image and behavioral risk level.
[0375] In an optional implementation, the data processing module 22 is further configured to: The occupant image region in the image data is identified to obtain the facial features, body proportion features, and posture features of the target occupant.
[0376] Physiological characteristics are determined based on facial features and body proportions.
[0377] Behavioral state information is determined based on posture features and the positional relationship between the target occupant and the target components inside the vehicle.
[0378] In an optional implementation, the target component inside the vehicle includes at least one of a seatbelt, a window, a door, and a child safety seat. When the target occupant is a child occupant, the data processing module 22 is further configured to: The system identifies key human body points, limb orientation, and positional changes of key human body points relative to target components within the vehicle for child occupants; key human body points include at least one of the hands, head, and torso.
[0379] Based on information about changes in limb orientation and position, determine whether the child occupant exhibits at least one of the following behavioral states: signs of seatbelt release, signs of leaning out of the window, signs of leaning forward, signs of tilting, and signs of leaving the child seat.
[0380] In an optional implementation, the virtual assistant generation module 23 is further configured to: Obtain the basic configuration information corresponding to the target occupant; wherein, the basic configuration information includes at least one of age information, gender information, and image preference information.
[0381] The physical appearance parameters of the target occupant are determined based on physiological characteristic information; wherein, the physical appearance parameters include at least one of facial structure parameters, head-to-body ratio parameters, height ratio parameters, skin color parameters, hairstyle parameters, and clothing adaptation parameters.
[0382] Based on basic configuration information and physical parameters, an initial virtual assistant image corresponding to the target occupant is generated.
[0383] The initial virtual assistant avatar is associated with and stored with the identity identifier of the target passenger.
[0384] In an optional implementation, the virtual assistant generation module 23 is further configured to: The physiological characteristic information of the target occupant is reacquired according to the preset update cycle or when changes in physiological characteristics are detected to exceed a preset threshold.
[0385] By comparing the newly acquired physiological characteristic information with historical physiological characteristic information, growth and change information is obtained.
[0386] If the growth and change information meets the preset growth and change conditions, update the physical parameters of the virtual assistant's image based on the growth and change information.
[0387] If the growth and change information does not meet the preset growth and change conditions, the physical parameters of the virtual assistant's image will remain unchanged.
[0388] In an optional implementation, the risk assessment module 24 is further configured to: Acquire consecutive multi-frame image data to determine behavioral state information.
[0389] Input multiple consecutive frames of image data into a pre-trained behavior recognition model.
[0390] Spatial features are extracted from each frame of image data using a behavior recognition model to obtain a spatial feature vector corresponding to each frame of image data. The spatial feature vector is used to characterize the posture features of the target occupant and its positional relationship with the target components inside the vehicle.
[0391] Multiple spatial feature vectors are combined according to the acquisition time sequence to obtain a temporal feature sequence.
[0392] Temporal change analysis is performed on the temporal feature sequence to obtain the behavior recognition results corresponding to the target occupant; the behavior recognition results include behavior category, behavior confidence level and behavior development trend.
[0393] The behavioral risk level is determined based on the behavioral category, behavioral confidence level, and behavioral development trend.
[0394] In an optional implementation, the risk assessment module 24 is further configured to: Obtain the risk assessment criteria corresponding to the behavior category.
[0395] Determine whether the confidence level of the behavior meets the confidence level requirements corresponding to the risk assessment conditions.
[0396] If the confidence level of the behavior meets the confidence level requirements, determine whether the trend of the behavior is pointing to a pre-set dangerous component or a pre-set dangerous area.
[0397] When the trend of behavior points to a pre-set dangerous component or pre-set dangerous area, the behavior category is determined to be a valid risk behavior.
[0398] The risk level of a behavior is determined based on the degree of risk corresponding to an effective risky behavior.
[0399] In an optional implementation, the risk assessment module 24 is further configured to: Within a preset time window, consistency is determined for the behavior categories corresponding to multiple frames of image data.
[0400] If multiple consecutive frames of image data within a preset time window correspond to the same behavior category, and the confidence levels of the behaviors corresponding to the same behavior category all meet the confidence level requirements, then the same behavior category will be identified as a candidate risk behavior.
[0401] Short-term interference filtering is applied to candidate risky behaviors.
[0402] If a candidate risky behavior is not identified as a short-term disruptive behavior, the candidate risky behavior will be identified as a valid risky behavior.
[0403] In an optional implementation, when the target occupant is a child occupant, the risk assessment module 24 is further configured to: Obtain historical behavioral data and historical early warning data for child occupants.
[0404] Based on historical behavioral data and historical early warning data, individualized risk assessment parameters are determined for child occupants. These individualized risk assessment parameters are used to characterize the risk assessment standards for child occupants at different stages of travel, at different ages, and with different historical behavioral habits.
[0405] Based on individualized risk assessment parameters, the risk assessment conditions, confidence requirements, preset hazardous components and / or preset hazardous areas corresponding to the behavior category are adjusted.
[0406] Based on the adjusted risk assessment criteria, behavior category, behavior confidence level, and behavior development trend, the behavioral risk level of child occupants is determined.
[0407] In an optional implementation, the risk assessment module 24 is further configured to: The study analyzed the routine and risky movement patterns of child occupants during their historical rides.
[0408] The range of typical behaviors for child occupants is determined based on their regular movement trajectories.
[0409] The range of risky behaviors corresponding to child occupants is determined based on historical risky action trajectories.
[0410] Based on the difference between the scope of routine behavior and the scope of risky behavior, individualized risk assessment parameters are generated.
[0411] Upon receiving feedback from the guardian regarding the warning information, the individualized risk assessment parameters are updated based on the feedback.
[0412] In an optional implementation, the interactive warning module 25 is further configured to: Based on behavioral state information, the target interaction scenario corresponding to the target occupant is determined.
[0413] Based on the target interaction scenario and physiological characteristics, determine the interaction content that matches the target occupant.
[0414] Control the virtual assistant avatar to output corresponding voice interaction information and / or visual interaction information according to the interaction content.
[0415] While outputting voice interaction information and / or visual interaction information, continue to acquire the target occupant's facial expression status information.
[0416] Adjust interactive content based on facial expression status information.
[0417] In an optional implementation, the interactive warning module 25 is further configured to: If the facial expression indicates that the target passenger is crying, the interaction content will be adjusted to a calming or soothing manner.
[0418] When facial expression information indicates that the target occupant is in a pleasant state, maintain the current interaction content or output continued interaction content related to the current interaction content.
[0419] When the behavioral status information indicates that the target occupant has a demand expression action, the demand prompt information is determined based on the demand expression action and sent to the vehicle display device.
[0420] In an optional implementation, the interactive warning module 25 is further configured to: When the behavioral risk level is low, control the virtual assistant avatar to output the first warning information.
[0421] When the behavioral risk level is medium risk, control the virtual assistant to output a second warning message and control the in-vehicle prompting device to output auxiliary prompt messages.
[0422] When the behavioral risk level is high, the system extracts the specific dangerous action characteristics corresponding to the behavioral status information, controls the virtual assistant image to present a warning posture corresponding to the specific dangerous action characteristics, outputs a third warning message, and controls the in-vehicle alarm device to output a danger warning message to the driver.
[0423] In an optional implementation, the interactive warning module 25 is further configured to: Obtain the behavioral status information of the target occupant and the status change results after the output of the danger warning information.
[0424] When the state change result indicates a decrease in the behavioral risk level of the target occupant, the output intensity of the warning information should be reduced or the warning information should be stopped.
[0425] If the risk level of the target occupant's behavior as a result of the state change has not decreased, the output intensity of the warning information should be increased, and the vehicle should be controlled to perform auxiliary safety control operations.
[0426] The computer program product provided in this application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0427] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0428] Furthermore, in the description of the embodiments of this application, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0429] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0430] In the description of this application, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0431] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the claims.
Claims
1. An interactive early warning method for an intelligent cockpit, characterized in that, include: Image data of the target occupants is acquired in real time using an image acquisition device; Physiological features and behavioral status information of the target occupant are extracted based on the image data; Based on the physiological characteristics information, generate or update the virtual assistant image corresponding to the target occupant; Based on the behavioral status information, the behavioral risk level of the target occupant is determined; Based on the virtual assistant's image and the behavioral risk level, output interactive information and / or warning information corresponding to the target occupant.
2. The method according to claim 1, characterized in that, The step of extracting the physiological feature information and behavioral state information of the target occupant based on the image data includes: The occupant image region in the image data is identified to obtain the facial features, body proportion features, and posture features of the target occupant; The physiological characteristic information is determined based on the facial features and the body proportion features; The behavioral state information is determined based on the posture characteristics and the positional relationship between the target occupant and the target components inside the vehicle.
3. The method according to claim 2, characterized in that, The target components inside the vehicle include at least one of seat belts, windows, doors, and child safety seats; When the target occupant is a child occupant, the step of determining the behavioral state information based on the posture characteristics and the positional relationship between the target occupant and the target component inside the vehicle includes: Identify the key human body points, limb orientation, and positional changes of the key human body points relative to the target component inside the vehicle for the child occupant; wherein the key human body points include at least one of the hands, head, and torso; Based on the limb orientation and position change information, determine whether the child occupant exhibits at least one of the following behavioral states: signs of seatbelt release, signs of leaning out of the window, signs of leaning forward, signs of tilting, and signs of leaving the child seat.
4. The method according to claim 1, characterized in that, The step of generating a virtual assistant image corresponding to the target occupant based on the physiological characteristic information includes: Obtain the basic configuration information corresponding to the target occupant; wherein, the basic configuration information includes at least one of age information, gender information, and image preference information; The physical appearance parameters of the target occupant are determined based on the physiological characteristic information; wherein, the physical appearance parameters include at least one of facial structure parameters, head-to-body ratio parameters, height-to-body ratio parameters, skin color parameters, hairstyle parameters, and clothing fit parameters; Based on the basic configuration information and the physical parameters, an initial virtual assistant image corresponding to the target occupant is generated. The initial virtual assistant image is associated with and stored with the identity identifier of the target passenger.
5. The method according to claim 4, characterized in that, The step of updating the virtual assistant image corresponding to the target occupant based on the physiological characteristic information includes: The physiological characteristic information of the target occupant is reacquired according to a preset update cycle or when a change in physiological characteristics is detected to exceed a preset threshold. The newly acquired physiological characteristic information is compared with historical physiological characteristic information to obtain growth and change information; When the growth change information meets the preset growth change conditions, the physical parameters of the virtual assistant image are updated based on the growth change information; If the growth change information does not meet the preset growth change conditions, the physical parameters of the virtual assistant image remain unchanged.
6. The method according to claim 1, characterized in that, The step of determining the behavioral risk level of the target occupant based on the behavioral state information includes: Acquire consecutive multi-frame image data for determining the behavioral state information; Input the image data from multiple consecutive frames into a pre-trained behavior recognition model; The behavior recognition model extracts spatial features from each frame of image data to obtain a spatial feature vector corresponding to each frame of image data; wherein, the spatial feature vector is used to characterize the posture features of the target occupant and its positional relationship with the target components inside the vehicle; Multiple spatial feature vectors are combined according to the acquisition time sequence to obtain a temporal feature sequence; A temporal change analysis is performed on the temporal feature sequence to obtain the behavior recognition result corresponding to the target occupant; the behavior recognition result includes behavior category, behavior confidence level, and behavior development trend. The risk level of the behavior is determined based on the behavior category, the confidence level of the behavior, and the development trend of the behavior.
7. The method according to claim 6, characterized in that, The step of determining the behavioral risk level based on the behavioral category, the behavioral confidence level, and the behavioral development trend includes: Obtain the risk assessment criteria corresponding to the behavior category; Determine whether the confidence level of the behavior meets the confidence level requirement corresponding to the risk assessment condition; If the confidence level of the behavior meets the confidence requirement, determine whether the trend of the behavior points to a preset dangerous component or a preset dangerous area; If the behavioral trend points to the preset hazardous component or the preset hazardous area, the behavioral category will be determined as a valid risky behavior. The risk level of the behavior is determined based on the degree of risk corresponding to the effective risk behavior.
8. The method according to claim 7, characterized in that, The step of identifying the behavior category as a valid risky behavior includes: Within a preset time window, a consistency judgment is made on the behavior categories corresponding to multiple frames of the image data; If multiple consecutive frames of image data within the preset time window correspond to the same behavior category, and the confidence scores of the behaviors corresponding to the same behavior category all meet the confidence score requirements, then the same behavior category is identified as a candidate risk behavior. The candidate risk behaviors are subjected to short-term interference filtering; If the candidate risk behavior is not identified as a short-term disruptive behavior, the candidate risk behavior is identified as the effective risk behavior.
9. The method according to claim 7, characterized in that, When the target occupant is a child occupant, after the step of obtaining the risk assessment conditions corresponding to the behavior category, the method further includes: Obtain the historical behavior data and historical warning data corresponding to the child occupants; Based on the historical behavior data and the historical early warning data, individualized risk assessment parameters are determined for the child occupants; wherein, the individualized risk assessment parameters are used to characterize the risk assessment standards corresponding to the child occupants at different riding stages, at different age stages, and with different historical behavioral habits; Based on the individualized risk assessment parameters, the risk assessment conditions corresponding to the behavior category, the confidence level requirements, the preset hazardous components and / or the preset hazardous areas are adjusted; The behavioral risk level of the child occupant is determined based on the adjusted risk assessment criteria, the behavior category, the behavior confidence level, and the behavior development trend.
10. The method according to claim 9, characterized in that, The step of determining the individualized risk assessment parameters for the child occupant based on the historical behavioral data and the historical early warning data includes: The routine and risky movement trajectories of the child occupants during their historical rides were statistically analyzed. The range of typical behaviors corresponding to the child occupant is determined based on the aforementioned typical movement trajectory. The range of risky behaviors corresponding to the child occupants is determined based on the historical risk action trajectories. The individualized risk assessment parameters are generated based on the difference between the range of normal behavior and the range of risky behavior. Upon receiving feedback from the guardian regarding the warning information, the individualized risk assessment parameters are updated based on the feedback.
11. The method according to claim 1, characterized in that, The step of outputting interaction information corresponding to the target occupant based on the virtual assistant avatar and the behavioral risk level includes: Based on the behavioral state information, the target interaction scenario corresponding to the target occupant is determined; Based on the target interaction scenario and the physiological characteristic information, determine the interaction content that matches the target occupant; Control the virtual assistant avatar to output corresponding voice interaction information and / or visual interaction information according to the interaction content; During the process of outputting the voice interaction information and / or the visual interaction information, the facial expression status information of the target occupant is continuously acquired. The interactive content is adjusted based on the facial expression status information.
12. The method according to claim 11, characterized in that, The step of adjusting the interactive content based on the facial expression state information includes: If the facial expression information indicates that the target occupant is crying, the interaction content will be adjusted to a soothing interaction content. When the facial expression information indicates that the target occupant is in a pleasant state, maintain the current interaction content or output continued interaction content associated with the current interaction content; When the behavioral status information indicates that the target occupant has a demand expression action, a demand prompt information is determined based on the demand expression action, and the demand prompt information is sent to the vehicle display device.
13. The method according to claim 1, characterized in that, Based on the virtual assistant's appearance and the behavioral risk level, output warning information corresponding to the target occupant, including: When the behavior risk level is low, the virtual assistant image is controlled to output a first warning message; When the behavioral risk level is medium risk, the virtual assistant image is controlled to output a second warning message, and the in-vehicle prompting device is controlled to output auxiliary prompting information. When the behavior risk level is high, the specific dangerous action features corresponding to the behavior status information are extracted, the virtual assistant image is controlled to present a warning posture corresponding to the specific dangerous action features, a third warning message is output, and the in-vehicle alarm device is controlled to output a danger warning message to the driver.
14. The method according to claim 13, characterized in that, After the step of controlling the in-vehicle alarm device to output a danger warning message to the driver, the method further includes: The result of the change in the behavioral state information of the target occupant after the hazard warning information is output is obtained; If the state change result indicates a decrease in the behavioral risk level of the target occupant, the output intensity of the warning information shall be reduced or the warning information shall be stopped. If the risk level of the target occupant's behavior does not decrease as indicated by the state change result, the output intensity of the warning information is increased, and the vehicle is controlled to perform auxiliary safety control operations.
15. An interactive early warning system for an intelligent cockpit, characterized in that, include: The image data acquisition module is used to acquire image data of the target occupant in real time through an image acquisition device; The data processing module is used to extract the physiological characteristics and behavioral state information of the target occupant based on the image data; The virtual assistant generation module is used to generate or update the virtual assistant image corresponding to the target occupant based on the physiological characteristic information. The risk assessment module is used to determine the behavioral risk level of the target occupant based on the behavioral status information. The interactive warning module is used to output interactive information and / or warning information corresponding to the target occupant based on the virtual assistant image and the behavioral risk level.
16. A vehicle, characterized in that, It includes an image acquisition device, an in-vehicle central control system, and an interactive early warning system for the smart cockpit as described in claim 15; The intelligent cockpit's interactive early warning system is integrated into the vehicle's central control system; The image acquisition device is communicatively connected to the vehicle central control system, and the image acquisition device includes a rear-seat monitoring camera installed on the top of the rear seat or at the B-pillar of the vehicle.