Passenger car intelligent cockpit dynamic interaction system based on multi-passenger behavior recognition

By collecting multimodal perception data in the passenger cabin and constructing a virtual space topology, the problem of decreased accuracy in occupant behavior recognition in existing technologies has been solved. This has enabled accurate recognition of multiple occupant behaviors and dynamic data fusion, thereby improving the synergistic optimization of driving safety and passenger experience.

CN121706034BActive Publication Date: 2026-04-14XIAMEN JINLONG CAR ACCESSORIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAMEN JINLONG CAR ACCESSORIES CO LTD
Filing Date
2026-02-13
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing intelligent cockpit systems for buses fail to fully consider the impact of passengers' actual behavioral direction and dynamic interaction scenarios during the multimodal perception data fusion process. This results in a decrease in the accuracy of behavior recognition in complex dynamic scenarios, making it impossible to effectively distinguish passenger behavior and affecting driving safety and passenger experience.

Method used

By collecting multimodal perception data such as vision, depth, sound, and pressure from multiple key points in the passenger cabin, a virtual space topology is constructed and perception subdomains are divided. By combining the perception credibility coefficient and the data fusion adjustment factor, adaptive weighted fusion of multimodal data is achieved, and the data fusion weight is dynamically adjusted to improve the reliability and adaptability of data processing.

Benefits of technology

It achieves comprehensive coverage and accurate capture of multi-occupant behavior information, and can accurately identify the number of occupants, their identity attributes, action intentions and emotional states, thereby improving the reliability of data processing in complex dynamic scenarios and ensuring the coordinated optimization of driving safety and riding experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121706034B_ABST
    Figure CN121706034B_ABST
Patent Text Reader

Abstract

The application provides a passenger car intelligent cabin dynamic interaction system based on multi-occupant behavior recognition, and relates to the technical field of passenger car intelligent cabin, and comprises: a calculation module, which is used for constructing a virtual space topological structure in the cabin based on four center points and multi-modal perception data, dividing the virtual space topological structure into multiple perception sub-domains, and generating a perception credibility coefficient for each perception sub-domain; for each center point, a corresponding three-dimensional perception geometric model is established, a space intersection of an occupant behavior direction vector and each three-dimensional perception geometric model is calculated, and a data fusion adjustment factor is generated for each perception sub-domain. Through the closed-loop mechanism of multi-modal data perception, calculation fusion and dynamic decision execution, the application realizes accurate identification and adaptive regulation and control of the multi-occupant behavior of the passenger car cabin, optimizes environmental regulation and individualized service experience, and simultaneously relies on iterative optimization to continuously improve the intelligence and adaptability of cabin interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent cockpit technology for buses, and in particular to a dynamic interactive system for intelligent cockpits for buses based on multi-occupant behavior recognition. Background Technology

[0002] In existing intelligent cockpit systems for buses, one technical solution for achieving occupant behavior recognition and interaction is to distribute multiple single-function sensor nodes within the cabin. This involves installing a visual camera in front of the driver's seat, a wide-angle camera in the center of the passenger area ceiling, and several microphones evenly distributed throughout the cabin. The data collected by these sensors (such as video and audio) is typically sent separately to a processing unit and pre-assigned to a fixed perception area (e.g., "driver's area" or "front of passenger area") based on their physical installation location. A major technical problem with this approach is that the fusion process of its multimodal perception data is based on static, pre-defined spatial partitioning, failing to fully consider the impact of the occupant's actual behavioral direction and dynamic interaction scenarios on the effectiveness of perception. Specifically, the system's perception weights are fixed after initialization and cannot adaptively adjust according to the occupant's real-time posture, line of sight, or sound source direction, potentially leading to a decrease in the accuracy of behavior recognition in complex dynamic scenarios.

[0003] For example, on a moving intercity bus, when a passenger in the middle of the carriage turns from facing forward to talking to their neighbor, their behavioral intention changes. At this moment, although the wide-angle camera on the top of the passenger area can capture their side profile, it becomes more difficult to effectively identify facial expressions and gesture details. At the same time, the fixed microphone array in the vehicle, because its beamforming direction is preset to the center of each area, may have difficulty accurately focusing the sound of the passenger shifting from their original position. However, the existing system still mechanically follows the static partitioning rule of "middle of the passenger area" and performs equal or fixed weighted fusion analysis on the data from the camera and microphone. This processing method may cause the system to be unable to effectively distinguish between the passenger's "talking sideways" behavior and "turning to face the seat armrest control panel to operate", thus causing subsequent service (such as accidental touch of media playback or lighting control) response errors. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a dynamic interactive system for intelligent cockpits of buses based on multi-occupant behavior recognition, so as to achieve synergistic optimization of driving safety and riding experience, and make up for the shortcomings of single system decision-making and difficulty in balancing safety and experience.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0006] The bus intelligent cockpit dynamic interaction system based on multi-occupant behavior recognition includes:

[0007] The perception module is used to collect multimodal perception data from the center point of the forward field of view in the passenger cabin, the center point of the driver's seat operating plane, the center point of the sound field in the passenger area, and the center point of the seat bearing plane.

[0008] The computing module is used to construct the virtual space topology within the cockpit based on four center points and multimodal perception data, and divide the virtual space topology into multiple perception subdomains, generating a perception confidence coefficient for each perception subdomain. For each center point, a corresponding three-dimensional perception geometric model is established. Based on continuous frame visual image data collected from the center point of the forward field of view and the center point of the driver's seat operating plane, the module extracts key facial and limb feature points of the occupants in adjacent frame images, calculates the displacement of each feature point in the pixel plane, obtains the spatial distance corresponding to the feature points by combining depth data, and converts the pixel displacement into a three-dimensional displacement vector in the spatial coordinate system by using the camera's in-camera installation orientation to generate an occupant behavior direction vector. The module calculates the spatial intersection points of the occupant behavior direction vector with each three-dimensional perception geometric model, and generates a data fusion adjustment factor for each perception subdomain.

[0009] The recognition module is used to combine the perception confidence coefficient of each perception subdomain with the data fusion adjustment factor to fuse and register the multimodal perception data collected from the center point, generate a unified perception data stream, and input the unified perception data stream into the pre-trained behavior recognition model for processing to obtain the behavior feature set of each occupant.

[0010] The decision-making module is used to calculate the final cabin environment adjustment strategy and personalized service strategy based on the set of behavioral characteristics and the vehicle's driving status through dynamic decision processing.

[0011] The execution module is used to execute cabin environment adjustment strategies and personalized service strategies, control the display terminals in the cabin in a coordinated manner, and collect passenger behavior feedback data and cabin environment status data after execution in order to iteratively optimize the interaction optimization data.

[0012] The above-described solution of the present invention has at least the following beneficial effects:

[0013] Breaking through the limitations of single-dimensional perception, this system collects multimodal perception data, including visual, depth, sound, and pressure data, from multiple key central points to achieve comprehensive coverage and accurate capture of multi-occupant behavior information within the bus cabin, avoiding information gaps caused by single-sensor acquisition. A dynamic data fusion mechanism is constructed, dividing perception subdomains through a virtual spatial topology and combining perception reliability coefficients with data fusion adjustment factors to achieve adaptive weighted fusion of multimodal data. This overcomes the rigid limitations of static partitioned data fusion, dynamically adjusting data fusion weights based on real-time occupant behavior direction and posture changes, improving the reliability and adaptability of data processing in complex dynamic scenarios. Based on multimodal data fusion and behavior recognition models, it achieves comprehensive and accurate identification of multi-occupant behavior characteristics, effectively distinguishing not only the number and identity attributes of occupants but also accurately judging action intentions, emotional states, and dangerous behaviors. The system addresses the issues of ambiguity and confusion in the recognition of complex behaviors, providing accurate judgment criteria for the decision-making module. It establishes a dynamic decision-making system with a dual-mode approach of safety priority and service collaboration, intelligently switching decision logic based on occupant behavior characteristics and vehicle driving status. This prioritizes driving safety when safety risks arise, while also considering group environment adaptation and individual personalized needs in normal scenarios, achieving synergistic optimization of driving safety and passenger experience. This overcomes the shortcomings of single-mode decision-making and the difficulty in balancing safety and experience. Furthermore, it constructs a closed-loop interactive mechanism encompassing execution, feedback, and iteration. Through the linkage control of cockpit equipment, it precisely implements decision strategies and continuously optimizes perception weights, recognition models, and decision logic based on behavioral feedback data and environmental status data. This enables the system to have self-learning and evolutionary capabilities, allowing it to adapt to the changing dynamic travel scenarios of multi-occupant buses over the long term, continuously improving the accuracy of interactive responses and user experience. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating the dynamic interaction system for a bus intelligent cockpit based on multi-occupant behavior recognition, provided by an embodiment of the present invention. Detailed Implementation

[0015] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0016] like Figure 1 As shown, embodiments of the present invention propose a dynamic interactive system for a bus intelligent cockpit based on multi-occupant behavior recognition, comprising:

[0017] The perception module is used to collect multimodal perception data from the center point of the forward field of view in the passenger cabin, the center point of the driver's seat operating plane, the center point of the sound field in the passenger area, and the center point of the seat bearing plane.

[0018] The computing module is used to construct the virtual space topology within the cockpit based on four center points and multimodal perception data, and divide the virtual space topology into multiple perception subdomains, generating a perception confidence coefficient for each perception subdomain. For each center point, a corresponding three-dimensional perception geometric model is established. Based on continuous frame visual image data collected from the center point of the forward field of view and the center point of the driver's seat operating plane, the module extracts key facial and limb feature points of the occupants in adjacent frame images, calculates the displacement of each feature point in the pixel plane, obtains the spatial distance corresponding to the feature points by combining depth data, and converts the pixel displacement into a three-dimensional displacement vector in the spatial coordinate system by using the camera's in-camera installation orientation to generate an occupant behavior direction vector. The module calculates the spatial intersection points of the occupant behavior direction vector with each three-dimensional perception geometric model, and generates a data fusion adjustment factor for each perception subdomain.

[0019] The recognition module is used to combine the perception confidence coefficient of each perception subdomain with the data fusion adjustment factor to fuse and register the multimodal perception data collected from the center point, generate a unified perception data stream, and input the unified perception data stream into the pre-trained behavior recognition model for processing to obtain the behavior feature set of each occupant.

[0020] The decision-making module is used to calculate the final cabin environment adjustment strategy and personalized service strategy based on the set of behavioral characteristics and the vehicle's driving status through dynamic decision processing.

[0021] The execution module is used to execute cabin environment adjustment strategies and personalized service strategies, control the display terminals in the cabin in a coordinated manner, and collect passenger behavior feedback data and cabin environment status data after execution in order to iteratively optimize the interaction optimization data.

[0022] In this embodiment of the invention, the limitations of single-dimensional perception are overcome. Multimodal perception data, including visual, depth, sound, and pressure data, are collected from multiple key central points to achieve comprehensive coverage and accurate capture of multi-occupant behavior information within the bus cabin, avoiding information loss caused by single-sensor acquisition. A dynamic data fusion mechanism is constructed, dividing the perception subdomains through a virtual spatial topology structure and combining the perception reliability coefficient with the data fusion adjustment factor to achieve adaptive weighted fusion of multimodal data. This breaks the rigidity of static partitioned data fusion and can dynamically adjust the data fusion weights according to the real-time behavioral direction and posture changes of occupants, improving the reliability and adaptability of data processing in complex dynamic scenarios. Based on multimodal data fusion and behavior recognition models, comprehensive and accurate identification of multi-occupant behavioral characteristics is achieved. This not only effectively distinguishes the number of occupants and their identity attributes but also accurately judges their action intentions, emotional states, and sense of danger. To address the issue of ambiguity and confusion in the system's recognition of complex behaviors, this approach provides accurate judgment criteria for the decision-making module. It establishes a dynamic decision-making system with a dual-mode approach prioritizing safety and coordinating service, intelligently switching decision logic based on occupant behavior characteristics and vehicle driving status. This prioritizes driving safety when safety risks arise while also considering group environment adaptation and individual personalized needs in normal scenarios, achieving synergistic optimization of driving safety and passenger experience. This overcomes the shortcomings of the system's singular decision-making and difficulty in balancing safety and experience. Furthermore, it constructs a closed-loop interactive mechanism encompassing execution, feedback, and iteration. Through the linkage control of cockpit equipment, it precisely implements decision strategies and continuously optimizes perception weights, recognition models, and decision logic based on behavioral feedback data and environmental status data. This enables the system to self-learn and evolve, adapting to the changing dynamic travel scenarios of multi-occupant buses and continuously improving the accuracy of interactive responses and user experience.

[0023] In another preferred embodiment of the present invention, multimodal perception data is collected from the center point of the forward field of view in the passenger cabin, the center point of the driver's operating plane, the center point of the sound field in the passenger area, and the center point of the seat bearing plane; the process of acquiring the multimodal perception data includes:

[0024] Visual image data and depth data are collected from the center point of the forward field of view and the center point of the driver's seat operating plane; sound signal data is collected from the center point of the passenger area sound field; and seat pressure distribution data is collected from the center point of the seat bearing plane. The visual image data, depth data, sound signal data, and seat pressure distribution data together constitute multimodal perception data. Specifically, this includes: firstly, determining the specific deployment positions of the four acquisition center points within the bus cabin, i.e., the center point of the forward field of view is deployed in the center above the dashboard directly in front of the driver's seat, specifically located at a horizontal distance of 60 to 70 centimeters from the front end of the driver's seat slide rail and a vertical height of 120 to 130 centimeters from the cabin floor, ensuring that the camera covers the entire driver's seat and the core forward area of ​​the cabin without obstruction; the driver's seat operating plane... The center point of the sound field in the passenger area is deployed at the geometric center of the steering wheel hub, with a horizontal distance of 50 to 60 centimeters from the driver's seat, within the range of arm movement of the driver in a natural operating posture, corresponding to the core area of ​​the driver's operating area; the center point of the sound field in the passenger area is deployed at the geometric center of the passenger area roof compartment, that is, at the intersection of the longitudinal centerline and the transverse midpoint of the compartment, with a vertical distance of 5 to 8 centimeters from the roof interior panel, ensuring that the sampling range uniformly covers the entire passenger area and the area within 1 meter around the driver's seat; the center point of the seat bearing plane is deployed at the geometric center of the seat cushion of each seat (including the driver's seat), specifically at the intersection of the midpoint of the front-to-back direction and the midpoint of the left-to-right direction of the seat cushion, with an embedding depth controlled at 2 to 3 centimeters, so as not to affect the normal comfort and load-bearing performance of the seat.

[0025] RGB-D cameras are installed at the center of the forward field of view and the center of the driver's seat operating plane. The camera lens at the center of the forward field of view is tilted downward at an angle of 15 to 20 degrees towards the driver's seat and the front area of ​​the cabin, which can fully capture the driver's facial expressions, upper body posture and the dynamic behavior of passengers in the front of the cabin. The camera lens at the center of the driver's seat operating plane is tilted upward at an angle of 10 to 15 degrees towards the driver's seat operating area, focusing on covering the steering wheel rotation angle, gear shift lever operation, central control button triggering and other operation behaviors. Both cameras are set to real-time continuous acquisition mode. When acquiring visual image data, the camera uses its built-in image sensor to capture visual information such as the outline, facial features, and body movements of occupants in the scene, generating a continuous sequence of image frames at a fixed frame rate. When acquiring depth data, the camera emits infrared light signals directionally towards the acquisition area, and simultaneously records the time difference between the infrared light emission and reflection back to the camera through the receiving module. The actual distance between the target and the camera is calculated precisely. Specifically, the depth distance value is equal to the speed of light multiplied by half the infrared light reflection time difference, i.e., depth distance value = speed of light × reflection time difference ÷ 2. Then, the target depth distance values ​​corresponding to each pixel in the image frame are integrated in pixel coordinate order to form a depth data matrix that corresponds one-to-one with the visual image frames.

[0026] A multi-channel microphone array is installed at the center of the sound field in the passenger area. The array consists of 6 to 8 microphones evenly distributed in a ring, providing comprehensive coverage of the entire passenger area and the area around the driver's seat. During operation, each microphone synchronously collects sound signals from within the cabin. The weak analog electrical signals are first amplified to a processable range by a signal amplification circuit, then converted into digital signals by an analog-to-digital converter to obtain the raw sound signal data. To achieve precise sound source localization, the microphone at the center of the array is selected as the reference microphone. The time difference between the time it takes for each other microphone to collect the same sound signal and the time it takes for the reference microphone to collect the same sound signal is calculated, resulting in multiple time difference data. This data is then combined with the microphone... The fixed spacing between microphones in the wind array (the spacing between adjacent microphones is 8 to 10 centimeters) is used to calculate the spatial coordinates of the sound source using triangulation. That is, a two-dimensional coordinate system is established with the reference microphone as the origin. Based on the distance between any two non-reference microphones and the reference microphone and the corresponding time difference, distance equations are listed respectively (distance from non-reference microphone to sound source = distance from reference microphone to sound source + sound speed × time difference). The planar coordinates of the sound source are obtained by solving the intersection of the two equations. Then, the three-dimensional coordinates of the sound source are determined by combining the installation height of the microphone array (vertical distance from the cabin floor). Finally, the original sound signal data and the three-dimensional coordinate data of the sound source are integrated along the time axis to form complete sound signal data.

[0027] A pressure sensor array is embedded at the center point of the seat's load-bearing surface in each seat. Each array contains 8 to 12 evenly distributed pressure sensing units, spaced 5 to 8 centimeters apart, ensuring coverage of the entire seat cushion's load-bearing area. When an occupant sits down or adjusts their posture, pressure acts on the seat cushion, causing the sensing units to elastically deform. The sensing units convert the pressure signal into a corresponding analog electrical signal. This analog signal is first filtered by a signal conditioning circuit to remove environmental interference and noise, and then converted into a digital voltage signal by an analog-to-digital converter. The actual pressure value is calculated based on the calibration parameters of the pressure sensor. The sensitivity coefficient of the pressure sensor is pre-calibrated to 3mV / N (millivolts / Newton), which represents the voltage output value corresponding to a unit pressure. The specific calculation process is as follows... Pressure value = actual voltage value corresponding to digital voltage signal ÷ sensitivity coefficient. Then, the pressure values ​​of all sensing units in the same seat are integrated, and combined with the fixed position coordinates of each sensing unit on the seat cushion, to form seat pressure distribution data that can reflect the occupant's sitting posture distribution and posture changes. Finally, the visual image data and depth data collected from the center point of the forward field of view and the center point of the driver's operating plane, the sound signal data collected from the center point of the passenger area sound field, and the seat pressure distribution data of all seats are time-stamped and synchronized by the synchronous control module in the cockpit. A unified system time base is used to mark the time frame of each type of data to ensure that each type of data corresponds one-to-one in the time dimension, and finally together they constitute comprehensive and synchronous multimodal perception data.

[0028] This embodiment achieves comprehensive coverage of multiple data types, including visual, depth, sound, and pressure, by deploying dedicated data acquisition equipment at key locations in the bus cabin. It captures not only the driver's operating behavior and status but also the behavior and sound information of the passenger area, compensating for the information gaps caused by single sensor acquisition. The acquisition process incorporates distance calculation, sound source localization calculation, and pressure value calculation to improve data accuracy and effectiveness. Various types of data complement and verify each other, effectively solving the defects of incomplete perception coverage and insufficient data accuracy.

[0029] In a preferred embodiment of the present invention, a virtual space topology within the cockpit is constructed based on four center points and multimodal perception data. This virtual space topology is then divided into multiple perception subdomains, and a perception confidence coefficient is generated for each subdomain. For each center point, a corresponding three-dimensional perception geometric model is established, and the spatial intersections of the occupant behavior direction vector with each three-dimensional perception geometric model are calculated. A data fusion adjustment factor is generated for each perception subdomain, including:

[0030] Using the forward field of view center point, driver's seat operating plane center point, passenger area sound field center point, and seat bearing plane center point as geometric vertices, a closed virtual space topology connecting these vertices is constructed based on the maximum effective sensing distance and field of view of the sensing devices corresponding to each center point. Specifically, this involves: first, defining the three-dimensional coordinates of the four acquisition center points, establishing a spatial rectangular coordinate system with the lower left corner of the bus cabin floor as the origin; and measuring the x-axis (longitudinal), y-axis (lateral), and z-axis (vertical height) coordinates of the forward field of view center point, driver's seat operating plane center point, passenger area sound field center point, and seat bearing plane center point, respectively, as the four basic geometric vertices of the virtual space topology; and then acquiring the core parameters of the sensing devices corresponding to each center point, namely the RGB-D cameras at the forward field of view center point and driver's seat operating plane center point, whose maximum effective sensing distance is 5 meters and field of view... The field angle is 60 degrees; the multi-channel microphone array at the center point of the passenger area sound field has a maximum effective sensing distance of 4 meters and a field of view (acoustic coverage angle) of 120 degrees; the pressure sensor array at the center point of the seat bearing plane has a maximum effective sensing distance of 0.5 meters and a field of view (pressure sensing coverage angle) of 90 degrees; based on the four vertices, according to the maximum effective sensing distance and field of view of each vertex sensing device, the effective sensing range boundary of each sensing device is drawn, that is, starting from each vertex, extending along the edge direction of the field of view to the maximum effective sensing distance, to obtain the effective sensing space boundary line of each sensing device; the sensing boundary lines of adjacent vertices are connected according to the actual space logic of the cabin, while ensuring that all boundary lines are closed, and finally a closed virtual space topology structure that fits the actual space shape of the bus cabin and connects the four vertices is constructed. This structure completely covers the effective sensing range of all sensing devices.

[0031] Based on the spatial envelope formed by the field-of-view boundaries of each sensing device within the virtual space topology, the virtual space topology is divided into multiple non-overlapping and continuous spatial regions. Each spatial region is defined as a sensing subdomain. Specifically, for each sensing device at its center point, combined with its installation parameters and sensing characteristics, the spatial shape and specific coordinate range of the field-of-view boundary are determined. The field-of-view boundary of the RGB-D camera is defined by the three-dimensional coordinates of the lens center. , , The cone-shaped surface with a vertex and a field of view of 60 degrees is the angle between the two generatrices; the camera mounting orientation is taken as the reference direction of the generatrices (e.g., the orientation of the forward-facing camera is 15 degrees off the positive x-axis and negative z-axis), and the two edge generatrices are symmetrically offset by 30 degrees to both sides of the reference direction, extending along the generatrices to the maximum effective sensing distance of 5 meters, and the coordinates of the endpoints of each generatrice are calculated. +5×cosθ, , +5×sinθ, where θ is the angle between the generatrix and the x-axis), and the endpoints are connected by a continuous curve to form a complete conical field of view boundary; the field of view boundary of the multi-channel microphone array is based on the three-dimensional coordinates of the array center ( , , A conical surface with a vertex and an acoustic coverage angle of 120 degrees, defined by the angle between the two generatrices; the reference direction of the generatrices is the negative z-axis (vertically downwards), and the two edge generatrices are symmetrically offset by 60 degrees to both sides of the y-axis, extending to a maximum effective sensing distance of 4 meters. Calculate the coordinates of the generatrice endpoints. , +4×sin60°, -4×cos60°), connecting the endpoints to form the acoustic sensing field of view boundary; the pressure sensor array field of view boundary is based on the three-dimensional coordinates of the sensor array center ( , , A hemispherical surface with a center and a pressure-sensing coverage angle of 90 degrees as a constraint; the radius of the hemisphere is 0.5 meters, and the flat part is flush with the seat cushion plane (z= The spherical part overlaps completely, with the spherical part facing into the cockpit (z> ) convexity, by calculating the coordinates of key nodes on the spherical surface ( +0.5×sinα×cosβ, +0.5×sinα×sinβ, +0.5×cosα, where α is from 0 to 90 degrees and β is from 0 to 360 degrees), connecting nodes to form a hemispherical field of view boundary; these field of view boundaries are precisely mapped onto the constructed virtual space topology. Based on the three-dimensional coordinates of the installation position of each sensing device, the field of view angle, and the maximum effective sensing distance, the three-dimensional coordinates of key nodes on the field of view boundary are calculated through spatial geometric operations. Then, the key nodes are connected using a cubic spline interpolation algorithm to form a continuous and closed spatial envelope surface; the field of view boundaries of each sensing device intersect and enclose each other in the virtual space, forming... The virtual space is divided into multiple independent continuous spatial regions. Using these spatial envelopes as the basis for segmentation, a spatial point attribution method is employed to divide the virtual space topology. Specifically, for any point (x, y, z) within the virtual space, the positional relationship between that point and each spatial envelope is calculated (e.g., to determine if a point is inside a cone, calculate the angle between the line connecting the point to the cone's vertex and the generatrix's reference direction; if the angle is ≤30 degrees and the distance from the point to the vertex is ≤ the maximum effective sensing distance, then it is inside; to determine if a point is inside a hemisphere, calculate the distance from the point to the center of the hemisphere ≤0.5 meters and z ≥ 0.5 meters). (If the area is inside, then the enclosed area to which the point belongs is determined; during the division process, the principle of non-overlapping and continuous coverage is strictly followed to ensure that each divided spatial area is independent and complete, and all areas together cover the entire virtual space without blank areas. Each independent spatial area is defined as a sensing subdomain, and each sensing subdomain belongs only to the effective sensing range of the corresponding sensing device, thus clarifying the sole source of data collection.

[0032] For each perceptual subdomain, the average edge gradient magnitude, point cloud density of depth data, signal-to-noise ratio of sound signal data, and spatial continuity of seat pressure distribution data of visual image data collected within its coverage area are statistically analyzed. Based on the numerical values ​​of average edge gradient magnitude, point cloud density, signal-to-noise ratio, and spatial continuity, a quantized perceptual confidence coefficient is generated. Specifically, this includes: for each perceptual subdomain, extracting all collected visual image data, depth data, sound signal data, and seat pressure distribution data within its coverage area, and performing feature statistics and quantization calculations respectively; for each frame of visual image within the perceptual subdomain, using the Sobel operator to calculate the edge gradient values ​​in the x and y directions respectively; for each pixel in the image, the gradient value in the x direction is calculated by the difference between the gray values ​​of the pixels to the right and left of the pixel, and the gradient value in the y direction is calculated by the difference between the gray values ​​of the pixels below and above the pixel, and then the single-pixel edge gradient magnitude is calculated by the square root of the sum of squares; the edge gradient magnitudes of all pixels in all image frames within the subdomain are statistically analyzed, and all magnitudes are summed and divided by the total number of pixels to obtain the average edge gradient magnitude. The number of effective depth values ​​in the depth data matrix within the statistical sensing subdomain is counted. The effective depth value is defined as the depth data within the maximum effective sensing distance of the corresponding sensing device. The spatial volume of the sensing subdomain is measured and calculated. The spatial volume is obtained by multiplying the length range of the subdomain on the x, y, and z coordinate axes. The point cloud density is calculated as: Point cloud density = Number of effective depth values ​​÷ Subdomain spatial volume.

[0033] From the sound signal data within the sensing subdomain, the amplitudes of the useful signal and noise signal are extracted. The useful signal amplitude is the peak amplitude of the sound signal, i.e., the amplitude corresponding to the highest point in the signal waveform. The noise signal amplitude is the average amplitude of the signal when there is no effective sound input, obtained by collecting the signal during a period of no effective sound and calculating its average amplitude. The signal-to-noise ratio (SNR) is calculated as: useful signal amplitude ÷ noise signal amplitude. The SNR is obtained through this division operation. For the seat pressure distribution data within the sensing subdomain, the pressure value differences between adjacent pressure sensing units on the same seat are statistically analyzed, and all adjacent units... The absolute values ​​of the pressure differences are added together to obtain the total pressure difference value. This total pressure difference value is then divided by the total number of adjacent units to obtain the average pressure difference value. A reference pressure value is set (e.g., the average human sitting pressure reference value of 300N). Spatial continuity = 1 ÷ (1 + average pressure difference value / reference pressure value). The closer this calculation result is to 1, the better the spatial continuity of the seat pressure distribution. The above four quantitative indicators are then normalized so that the values ​​of each indicator are between 0 and 1. The normalization formula is: Normalized value = (Original value - Minimum value of the indicator) ÷ (Maximum value of the indicator) The values ​​are defined as follows: (minimum value of the indicator), where the maximum and minimum values ​​of each indicator are preset industry-standard effective range extreme values. Specifically, the effective range extreme values ​​for the average edge gradient magnitude are a minimum of 0 and a maximum of 100; the effective range extreme values ​​for point cloud density are a minimum of 0.01 points / cubic meter and a maximum of 10 points / cubic meter; the effective range extreme values ​​for signal-to-noise ratio are a minimum of 10 and a maximum of 50; and the effective range extreme values ​​for spatial continuity are a minimum of 0.1 and a maximum of 1. This ensures the reasonableness and comparability of the normalized results. A weighted summation method is used to calculate the perceived credibility coefficient. Because the four indicators—average edge gradient magnitude, point cloud density, signal-to-noise ratio (SNR), and spatial continuity—are all core dimensions reflecting the reliability of sensing data and have equal weight in influencing data quality, the weight of each indicator is set to 0.25. By weighted summation, the performance of each dimension can be comprehensively analyzed, objectively quantifying the overall reliability of the sensing data. The sensing reliability coefficient is calculated as follows: (normalized average edge gradient magnitude × 0.25) + (normalized point cloud density × 0.25) + (normalized signal-to-noise ratio × 0.25) + (normalized spatial continuity × 0.25). The core function of this coefficient is to provide a clear weighting basis for subsequent multimodal data fusion, prioritizing the fusion of sensing subdomains with high reliability coefficients to improve the accuracy and effectiveness of data fusion. Through this weighted summation, the quantified sensing reliability coefficient for each sensing subdomain is obtained, ranging from 0 to 1. The closer the coefficient is to 1, the higher the reliability of the sensing data in that subdomain.

[0034] For each center point, a three-dimensional geometric body representing its actual sensing space is constructed as a three-dimensional sensing geometric model based on the installation position, orientation, and sensing range parameters of its corresponding sensing device. Specifically, this includes: for each acquisition center point, collecting the three-dimensional coordinates of its corresponding sensing device's installation position, installation orientation angle (angle with each axis of the spatial coordinate system), and sensing range parameters (maximum effective sensing distance, field of view angle), ensuring the accuracy and completeness of the parameters; for the RGB-D camera at the center point of the forward field of view and the center point of the driver's seat operating plane, its installation position three-dimensional coordinates ( , , Using the vertex as an example, determine the generatrix direction of the conical surface based on the installation orientation angle, ensuring the generatrix direction aligns with the camera's installation orientation. Extend along the generatrix direction to the maximum effective sensing distance H (5 meters), constructing a conical three-dimensional geometric body. The base of the cone is a circular plane perpendicular to the generatrix, with a base radius r = H × tan(field of view ÷ 2) = 5 × tan30°. The equation of this cone is... + + = × ,in The unit vector in the direction of the generatrix. The three-dimensional coordinates of any point on the cone surface are given. This cone is the three-dimensional perceptual geometric model of the RGB-D camera, which can accurately represent its actual visual perception space; for the multi-channel microphone array at the center point of the passenger area sound field, its installation position three-dimensional coordinates are given. Using ) as the vertex, determine the generatrix direction of the conical surface based on the installation orientation angle, and extend along the generatrix direction to the maximum effective sensing distance L (4 meters) to construct a conical three-dimensional geometric body. The base radius R = L × tan(acoustic coverage angle ÷ 2) = 4 × tan60°; the equation of the cone is... + + = × ,in The unit vector in the direction of the generatrix. The three-dimensional coordinates of any point on the conical surface are used, and this cone serves as the three-dimensional perceptual geometric model of the microphone array, accurately representing its acoustic perception space; for the pressure sensor array at the center point of the seat bearing plane, its three-dimensional coordinates at its installation position are used. ( ) is the center of the sphere, and the maximum effective sensing distance is... Using a radius of 0.5 meters and incorporating the pressure-sensing coverage angle, construct a hemispherical three-dimensional geometry. The planar portion of the hemisphere intersects with the seat cushion plane (z= The two spaces completely overlap, leaving only the space facing the interior of the cockpit (z≥ The effective sensing area is the hemisphere of a circle, and the equation of the sphere is: , Given the three-dimensional coordinates of any point on the sphere, this hemisphere is the three-dimensional sensing geometric model of the pressure sensor array, accurately representing its pressure sensing space.

[0035] For continuous frame visual image data acquired from the center point of the forward field of view and the center point of the driver's seat operating plane, the occupant behavior direction vector, representing the occupant's movement trend, is obtained through pixel displacement calculation. The set of intersection points between each occupant behavior direction vector and the surface of each three-dimensional perceptual geometric model is calculated, and the spatial distribution characteristics of the intersection point sets belonging to the same perceptual subdomain are statistically analyzed. Specifically, this includes: extracting continuous frame visual image data acquired from the center point of the forward field of view and the center point of the driver's seat operating plane, selecting two adjacent frames as the analysis object; extracting key feature points of the occupant in the previous frame image using the SIFT algorithm, such as facial features and limb joint features; finding matching feature points corresponding to these feature points in subsequent frames; recording the pixel coordinate changes of each feature point in the two frames; and obtaining the pixel displacement of each feature point, such as from pixel coordinates... Become The pixel displacement is Starting from the pixel coordinates of the feature point in the first frame image, and using the pixel displacement as the direction vector, the pixel displacement is converted into a spatial displacement vector by combining the camera's intrinsic parameters (focal length f, pixel size p). The specific calculation process is as follows: the physical displacement corresponding to the pixel displacement. That is, the actual physical length corresponding to each pixel displacement is calculated by converting the pixel size; according to the principle of similar triangles, the actual length L of the spatial displacement vector = (magnitude of the physical displacement) × (depth value of the target distance from the camera) ÷ f, where the depth value of the target distance from the camera is obtained from the depth data; the direction of the spatial displacement vector is consistent with the direction of the pixel displacement, and combined with the camera mounting orientation, the displacement direction in the pixel plane is converted into a direction vector in the spatial coordinate system. ),in They are spatial vectors and The included angle of the axis is obtained through the camera's internal installation angle calibration, ultimately resulting in a spatial displacement vector. The direction of the occupant's behavior is the occupant's behavior direction vector, which can accurately represent the occupant's movement trend.

[0036] Substitute each occupant behavior direction vector into the 3D perception geometric model corresponding to all center points, and solve for the intersection points of the vectors with the surface of the 3D geometric body through spatial geometric calculations; for the conical model, the parametric equations of the occupant behavior direction vectors ( ),in( ) is the coordinate of the starting point of the vector, ( () is a unit vector in the direction of the vector. Substituting the non-negative real number (into the corresponding conical equation) and expanding it, we get the following about... Solving the quadratic equation in one variable yields... Positive solution ( >0 and ≤ vector length), then Substituting the values ​​into the parametric equations yields the coordinates of the intersection points; for the hemispherical model, substituting the vector parametric equations into the spherical equations and solving for the results... Find the positive solution, while ensuring that the coordinates of the intersection point satisfy z≥ The hemispherical constraint conditions are used to obtain the coordinates of the intersection points; all the coordinates obtained from the solution constitute the intersection point set; based on the three-dimensional coordinates of each intersection point, its corresponding receptive subdomain is determined, that is, by comparing the intersection point coordinates with the spatial coordinate range of each receptive subdomain (e.g., to determine whether it is within the conical subdomain, verify the included angle and distance conditions; to determine whether it is within the hemispherical subdomain, verify the distance and z-axis conditions), to determine which receptive subdomain the intersection point falls into; all intersection points within the same receptive subdomain are classified to form the intersection point set corresponding to that receptive subdomain; statistical analysis is performed on each intersection point set, recording the number of intersection points, the three-dimensional coordinates, and the range of the intersection points. Information such as the distribution range on the axis and the distance between any two intersection points allows for a comprehensive understanding of the spatial distribution characteristics of the intersection point set.

[0037] Based on the density and geometric regularity of spatial distribution characteristics, a data fusion adjustment factor for weighted fusion is generated. This includes: calculating the standard deviation of the 3D coordinates of all spatial intersection points belonging to the same perceptual subdomain, using the reciprocal of the standard deviation as a density index to quantify density; calculating the directional consistency of the vector defined by the starting point of each spatial intersection point and its corresponding occupant behavior direction vector, using the directional consistency value as a regularity index to quantify geometric regularity; and normalizing the product of the density index and the regularity index to generate the data fusion adjustment factor. Specifically, this involves: extracting the 3D coordinates of all intersection points in the same perceptual subdomain; calculating the average values ​​of the x-axis, y-axis, and z-axis coordinates respectively, then calculating the sum of squared deviations of each coordinate from the corresponding axis average value, summing the sums of squared deviations of the three axes and dividing by the number of intersection points to obtain the variance; the square root of the variance is the standard deviation, and the density index = The larger the standard deviation, the smaller the density index, and vice versa. This calculation quantifies the density of intersection points. For each intersection point in the set, a local coordinate system is established with the origin of the corresponding occupant behavior direction vector. The vector of the intersection point relative to the origin is obtained, i.e., the vector pointing from the origin to the intersection point. The angle θ between this vector and the occupant behavior direction vector is calculated. The direction consistency is calculated using the vector dot product formula: Direction consistency = ... This result is equivalent to cosθ. The smaller the angle θ, the closer cosθ is to 1, and the better the directional consistency. The directional consistency of all vectors in the intersection set is statistically analyzed. The arithmetic mean of all directional consistency values ​​is obtained by summing all the values ​​and dividing by the number of intersection points. This arithmetic mean serves as the regularity index for the intersection set. The density index is multiplied by the regularity index to obtain the initial fusion factor. Multiplication is chosen because the density index reflects the concentration of intersection points within the perception subdomain. The more concentrated the intersection points, the more focused the perception of occupant behavior is in that subdomain, and the stronger the data correlation. The regularity index reflects the consistency between the vector corresponding to the intersection point and the occupant behavior direction vector. The higher the consistency, the more accurately the data collected in that subdomain can characterize the trend of occupant behavior, and the stronger the data effectiveness. Multiplying the two can comprehensively reflect the focus and effectiveness of the data within the perception subdomain, providing a reasonable priority basis for data weighted fusion. That is, the initial fusion factor = density index × regularity index. The initial fusion factors of all perception subdomains are collected, and the maximum and minimum values ​​are determined. The initial fusion factors are then normalized. After normalization, the fusion factor = The normalization operation makes the fusion factor value between 0 and 1. This value is the data fusion adjustment factor. The closer it is to 1, the higher the data fusion priority of this perceptual subdomain.

[0038] This embodiment achieves a refined division of the cockpit's perception space by constructing a virtual space topology based on four acquisition center points and combining the parameters of the sensing devices, thus dividing the space into perception subdomains. This avoids the limitations of static partitioning and ensures that the perception range accurately matches the actual cockpit space and device performance. The calculation of the perception reliability coefficient integrates the core features of multiple types of data, objectively reflecting the data reliability of each perception subdomain through quantitative indicators. This provides a scientific weight reference for data fusion, effectively improving the rationality and accuracy of data fusion. The construction of the three-dimensional perception geometric model accurately restores the actual perception space of each sensing device. By combining the intersection point set with the occupant behavior direction vector, the association between occupant behavior and perception space can be dynamically captured, making the generation of data fusion adjustment factors more closely match the dynamic behavior characteristics of occupants. The data fusion adjustment factor comprehensively considers the density and geometric regularity of the intersection point set, realizing adaptive data fusion weight adjustment based on occupant dynamic behavior. This solves the problem of insufficient recognition accuracy of fixed weight fusion in complex scenarios and improves the system's adaptability to complex dynamic scenarios.

[0039] In a preferred embodiment of the present invention, the multimodal sensing data collected from the center point is fused and registered by combining the sensing confidence coefficients of each sensing subdomain with the data fusion adjustment factor to generate a unified sensing data stream. This unified sensing data stream is then input into a pre-trained behavior recognition model for processing to obtain the behavior feature sets of each occupant, including:

[0040] For each perception subdomain, the perception confidence coefficient of the corresponding perception subdomain is multiplied by the data fusion adjustment factor to obtain the data fusion weight of that perception subdomain. Specifically, for each perception subdomain that has been divided, the perception confidence coefficient (range 0 to 1) and the data fusion adjustment factor (range 0 to 1) of that perception subdomain are retrieved first. The data fusion weight of that perception subdomain is calculated by multiplication. The calculation formula is: Data fusion weight = Perception confidence coefficient of the perception subdomain × Data fusion adjustment factor of the perception subdomain. The larger the weight value, the higher the contribution of the multimodal data of that perception subdomain in the fusion process, and the more the final fusion result will focus on the effective data of that subdomain.

[0041] For each sensing subdomain, multimodal sensing data belonging to the same time window collected from all center points within its spatial range are acquired. Using the data fusion weight of this sensing subdomain, a weighted average is calculated on the multimodal sensing data. The weighted average visual image data, depth data, sound signal data, and seat pressure distribution data are then normalized to obtain normalized multimodal sensing data. The normalized multimodal sensing data is then timestamped and aligned with the spatial coordinate system to generate a unified sensing data stream. Specifically, for each sensing subdomain, a unified time window is first defined (e.g., 1 second as a time window, with time precision controlled at the millisecond level). All center points within the spatial range of this subdomain (forward...) are then selected. Multimodal perception data (including the center point of the field of view, the center point of the driver's seat operating plane, the center point of the passenger area sound field, and the center point of the seat bearing plane) collected within this time window includes visual image data, i.e., color image frames captured by RGB-D cameras (with a uniform resolution of 1920×1080); depth data, i.e., point cloud format depth data captured by RGB-D cameras; sound signal data, i.e., audio waveform data captured by microphone arrays (sampling rate 48kHz, bit depth 16bit); and seat pressure distribution data, i.e., pressure value matrix captured by pressure sensor arrays (sampling frequency 10Hz). It is ensured that the selected data all belong to the same time window and cover the complete spatial range of this perception subdomain, without data loss or time misalignment issues.

[0042] Based on the data fusion weights, a weighted average is applied to the multimodal perception data within the same time window. For all visual image frames in the subdomain within the same time window, a weighted average is calculated pixel-by-pixel. Assuming there are several visual image frames in the subdomain, and each frame has a corresponding grayscale value at its pixel coordinates, the weighted grayscale value of that pixel is calculated as (the sum of all corresponding pixel grayscale values ​​multiplied by the data fusion weights) ÷ the number of image frames. Similarly, for the depth point cloud data in the subdomain within the same time window, a weighted average is calculated point-by-point. Assuming there are several valid depth values ​​in the subdomain, and each depth value corresponds to one original data point, the weighted depth value is calculated as (the sum of all original depth values). The sum of the values ​​after multiplying by the data fusion weights is calculated as follows: (sum of all values ​​after multiplying by the data fusion weights) ÷ total number of effective depth values; For audio waveform data within the same time window, the weighted average is calculated for each sample point. Assuming there are several audio sample points within the time window, each sample point has a corresponding amplitude. The weighted amplitude is calculated as follows: (sum of all sample point amplitudes after multiplying by the data fusion weights) ÷ total number of sample points; For pressure value matrix within the same time window, the weighted average is calculated for each sensing unit. Assuming there are several pressure sensing units within the subdomain, each unit has a corresponding pressure value. The weighted pressure value is calculated as follows: (sum of all sensing unit pressure values ​​after multiplying by the data fusion weights) ÷ total number of sensing units.

[0043] Normalize each type of data after weighted averaging so that the data values ​​are all in the range of 0 to 1; convert the pixel grayscale values ​​(0 to 255) to the range of 0 to 1, the formula is: normalized pixel value = weighted average pixel grayscale value ÷ 255; take the maximum effective sensing distance of the corresponding sensing device as the benchmark (e.g., the maximum effective distance of an RGB-D camera is 5 meters), the formula is: normalized depth value = weighted average depth value ÷ 5; take the maximum amplitude of the audio signal (16-bit depth corresponds to a maximum amplitude of 32767) as the benchmark, the formula is: normalized amplitude = weighted average amplitude ÷ 32767; take the maximum range of the pressure sensor (e.g., 500N) as the benchmark, the formula is: normalized pressure value = weighted average pressure value ÷ 500. Using a unified high-precision timestamp (millisecond level) as a benchmark, the timestamps of normalized visual images, depth, sound, and pressure data are calibrated to the same time axis, with the error controlled within ±10ms. The spatial coordinates of all data are transformed to the constructed global spatial Cartesian coordinate system of the cockpit (with the left corner of the leading edge of the cockpit floor as the origin, the x-axis vertical, the y-axis horizontal, and the z-axis vertical). The point cloud coordinates of the depth data, the installation coordinates of the pressure sensor, and the sound source localization coordinates of the sound signal are all unified to this coordinate system to ensure that the spatial positions of different modal data correspond one-to-one. Finally, all normalized and aligned single-modal data are integrated to generate a continuous unified perception data stream, which contains complete information in the time dimension, spatial dimension, and multimodal feature dimension.

[0044] The unified perception data stream is input into a pre-trained behavior recognition model in a time series manner, outputting a set of behavior features including the number of occupants, their identity attributes, action intentions, and emotional states. This includes: inputting the time series unified perception data stream into the convolutional neural network layer of the pre-trained behavior recognition model to extract spatial features from visual image data and depth data, as well as spectral features from sound signal data; concatenating the spatial and spectral features with the temporal features of seat pressure distribution data to form a fused feature vector; inputting the fused feature vector into the long short-term memory network layer of the pre-trained behavior recognition model to obtain temporal context features; and, based on these temporal context features, obtaining multiple independent feature vectors corresponding to each occupant through feature separation processing in the pre-trained behavior recognition model; inputting each independent feature vector into the identity attribute classification layer, action intention classification layer, and emotional state classification layer of the pre-trained behavior recognition model to obtain each occupant's identity attributes, action intentions, and emotional states; counting the number of independent feature vectors to obtain the number of occupants; and finally, the set of behavior features, consisting of the number of occupants, their identity attributes, action intentions, and emotional states, specifically includes:

[0045] First, the behavior recognition model is constructed and pre-trained. The behavior recognition model adopts a hybrid architecture of convolutional neural network (CNN) + long short-term memory network (LSTM) + multi-classification head. The feature extraction layer uses an improved ResNet50 as the backbone of the CNN network and adds a spectral feature extraction branch (containing 3 layers of 1D convolution) to process the spatial features of visual image / depth data and the spectral features of sound signals, respectively. The temporal feature layer adopts a two-layer LSTM network (256 dimensions per hidden layer) to process the temporal features of seat pressure distribution and the spatial / spectral features of the CNN output. The classification output layer includes an identity attribute classification layer (Softmax classifier), an action intention classification layer (Softmax classifier), and an emotional state classification layer (Softmax classifier), all connected to the LSTM output.

[0046] The model training process involves collecting multimodal perception data from various scenarios within the bus cabin (normal driving, passenger interaction, emergency braking, etc.), labeling the number of passengers, their identities (adults / children / drivers / passengers), their intentions (getting up / sitting down / operating equipment / conversing), and their emotional states (calm / anxious / pleasant / angry) to construct a labeled dataset. First, the labeled dataset is divided into training, validation, and test sets in a 7:2:1 ratio. Second, data augmentation is performed on the training set data (random image cropping, audio noise addition, and temporal perturbation of stress data). Third, the model is trained using the Adam optimizer (learning rate 1e-4) with cross-entropy loss as the loss function, iterating 100 times, stopping training when the validation set accuracy improves by less than 0.5% in each iteration. Fourth, the model performance is validated using the test set, ensuring that the accuracy of identity attribute recognition is ≥95%, the accuracy of intention recognition is ≥92%, and the accuracy of emotional state recognition is ≥88%, meeting the requirements of practical applications.

[0047] After classifying the unified perception data stream arranged in time series by modality, it is synchronously input into the pre-trained behavior recognition model. Feature extraction, fusion, and separation are performed step by step. First, multimodal feature extraction is performed. Visual image data and depth data in the unified perception data stream are input into the CNN layer frame by frame in time sequence. The CNN layer contains 5 convolutional blocks and 3 pooling layers. Each convolutional block performs feature convolution operation through a 3×3 convolutional kernel, gradually extracting multi-scale spatial features from pixel level to semantic level. Low-level features focus on detailed information such as occupant limb contours and edge textures, while high-level features cover global information such as occupant spatial position and posture contours. The pooling layer uses max pooling operation to compress data dimensionality while retaining key features, thereby improving feature extraction efficiency. For audio signal data, the audio waveform data in the time domain is first converted into a Mel spectrogram (the Mel scale is divided into 26 frequency bands, the frame length is 20ms, and the frame shift is 10ms) through a short-time Fourier transform. Then, the Mel spectrogram is input into the spectral feature branch of the CNN. This branch contains 3 1D convolutional layers, which sequentially extract the low-frequency basic features, mid-frequency pitch features, and high-frequency detail features of the audio signal. Finally, the frequency domain features covering information such as speech pitch, sound source frequency, and sound intensity changes are output.

[0048] Subsequently, feature dimension concatenation is performed, combining the spatial features (dimension 512) output from the CNN layer, the frequency domain features (dimension 256) output from the spectral feature branch, and the temporal features of the seat pressure distribution data (obtained by extracting the trend of change from the pressure value matrix over time, dimension 256). During the concatenation process, the dimensions of the three types of features are first standardized and aligned to ensure data format consistency. Then, the three are merged by channel dimension through the feature concatenation layer to form a fused feature vector with dimension 1024. This vector comprehensively integrates key information from visual, auditory, and tactile multimodal sensing.

[0049] Next, temporal context features are extracted. The fused feature vector is input into the LSTM layer in chronological order. The LSTM layer contains two hidden layers (each with a dimension of 256). Through the synergistic effect of the input gate, forget gate, and output gate, the temporal dependencies of the fused feature vector are captured. For dynamic information such as continuous changes in occupant movements (e.g., the continuous movement of limb joints when standing up) and temporal fluctuations in pressure values ​​(e.g., the change in pressure from zero to high and from low when sitting down), the LSTM layer can effectively memorize long-term temporal features and filter out irrelevant noise interference. Finally, the output temporal context features with a dimension of 512 are obtained. These features contain both multimodal static features and dynamic change information in the time dimension.

[0050] Finally, feature separation and classification prediction are performed. The model's built-in attention-based feature separation module performs individual-level separation of temporal context features. This module first calculates the attention weight for each temporal position, focusing on feature regions related to the individual occupant. Then, through spatial attention allocation and feature masking operations, it distinguishes the features of different occupants, ultimately obtaining multiple independent feature vectors corresponding to each occupant (each vector has a dimension of 256), ensuring that each vector contains only the feature information of a single occupant without cross-interference. Each independent feature vector is then input into the identity attribute classification layer, action intention classification layer, and emotional state classification layer, respectively. Each classification layer consists of a fully connected layer and a Softmax classifier. The fully connected layer first maps the independent feature vectors to the category probability space, and then the Softmax classifier converts the mapping results into probability distributions for each category (the sum of all category probabilities is 1). For the identity attribute classification layer, the categories include adults, children, drivers, and passengers. The category corresponding to the highest probability is taken as the occupant's identity attribute (e.g., adult, child, driver, passenger). The system outputs an adult probability of 0.98, a child probability of 0.01, a driver probability of 0.005, and a passenger probability of 0.005, thus identifying the occupant as an adult. For the action intent classification layer, categories include getting up, sitting down, operating equipment, conversing, and seeking help; the occupant's current action intent is determined by the highest probability. For the emotional state classification layer, categories include calm, anxious, happy, and angry; the occupant's emotional state is determined based on the highest probability, ultimately achieving accurate identification of each occupant's multi-dimensional features. The number of independent feature vectors obtained after feature separation is the number of occupants in the current cabin. The statistically obtained number of occupants is integrated with each occupant's identity attributes, action intent, and emotional state to form a complete behavioral feature set. For example, the behavioral feature set = {Number of occupants: 3; Occupant 1: Identity attribute - Adult, Action intent - Conversation, Emotional state - Calm; Occupant 2: Identity attribute - Child, Action intent - Get up, Emotional state - Happy; Occupant 3: Identity attribute - Driver, Action intent - Operating equipment, Emotional state - Anxious}.

[0051] This embodiment obtains the fusion weight by multiplying the perceived credibility coefficient by the data fusion adjustment factor, which integrates data reliability and the correlation with occupant behavior. This makes the weight allocation more in line with the actual perception scenario, avoids fusion bias caused by single-dimensional weights, and improves the rationality of data fusion. The entire process of weighted averaging, normalization, and spatiotemporal alignment of multimodal data unifies the data's dimensions, time reference, and spatial reference, solving the problem of inconsistent formats, ranges, and spatiotemporal dimensions of different modal data. The process of building the behavior recognition model before training ensures the model's effectiveness and generalization ability. Combined with the architecture of CNN to extract spatial / spectral features and LSTM to capture temporal features, it can comprehensively mine the deep features of multimodal data. Feature separation and multi-classification layer design realize the accurate identification of individual occupant features. The output behavior feature set fully covers the dimensions of occupant number, identity, intent, and emotion, providing a comprehensive and accurate behavioral analysis basis for scenarios such as intelligent cockpit interaction and safety monitoring.

[0052] In a preferred embodiment of the present invention, based on a set of behavioral characteristics and the vehicle's driving state, a final cabin environment adjustment strategy and personalized service strategy are calculated through dynamic decision processing, including:

[0053] Based on the behavioral feature set, it is determined whether there are safety-related behaviors. These safety-related behaviors include at least one of driver fatigue, driver distraction, and dangerous passenger actions. Specifically, this involves: first, retrieving the complete behavioral feature set and extracting the identity attributes, action intentions, emotional states, and passenger location distribution information of all occupants; then, based on preset safety behavior judgment rules, verifying one by one whether there are any of the three types of safety-related behaviors: driver fatigue, driver distraction, and dangerous passenger actions. The specific judgment criteria for each type of behavior are as follows: First, driver fatigue judgment: if the driver's emotional state in the behavioral feature set is drowsy, and the action intention is characterized by a drooping head, closed eyes, and relaxed limbs, and the duration exceeds a preset 5-second threshold, then it is determined that there is a safety-related behavior of driver fatigue; second, driver distraction judgment: if the driver's emotional state in the behavioral feature set is drowsy, and the action intention is characterized by a drooping head, closed eyes, and relaxed limbs, and the duration exceeds a preset 5-second threshold, then it is determined that there is a safety-related behavior of driver fatigue; The first category is driver distraction safety-related behavior. This includes actions such as holding non-driving devices, turning the head towards the passenger area, and deviating from the driving direction, all of which are unrelated to normal vehicle operation. The second category is passenger dangerous actions. If a passenger's actions in the behavioral characteristic set involve climbing windows, pulling on doors, interfering with driving, extending limbs outside the vehicle, or engaging in actions that disrupt cabin order such as getting up or running while the vehicle is in motion, this is considered a dangerous safety-related behavior. During the judgment process, continuous time verification of the dynamic behavioral information in the behavioral characteristic set is required to ensure the accuracy of the judgment and avoid misjudgment based on a single action. If any one or more of the above three types of behavior are present in the behavioral characteristic set, it is considered a safety-related behavior; otherwise, it is considered that no safety-related behavior exists.

[0054] If safety-related behaviors exist, the system enters a safety-priority decision-making mode. Based on the vehicle's driving status, a first group adjustment strategy is generated as the final cabin environment adjustment strategy, while all personalized service strategies are suspended. This first group adjustment strategy includes reducing passenger area volume, turning off passenger area entertainment displays, and increasing driver area warning intensity. If no safety-related behaviors exist, the system enters a service collaboration decision-making mode, generating a second group adjustment strategy at the group level and a corresponding personalized service strategy. The first or second group adjustment strategy is combined with the corresponding personalized service strategy to form the final cabin environment adjustment strategy and personalized service strategy. The process of determining the second group adjustment strategy at the group level includes: matching corresponding personalized service instructions from a preset service rule base based on the identity attributes, action intentions, and emotional states of each occupant in the behavioral feature set to generate an individual-level personalized service strategy; simultaneously, based on the number of occupants and their location distribution in the behavioral feature set, calculating cabin sound field equalization parameters, temperature equalization parameters, and lighting equalization parameters to generate the second group adjustment strategy at the group level, specifically including:

[0055] Based on the judgment results, the cabin environment adjustment and service strategy generation operations are executed in two categories: safety-first decision mode and service-coordinated decision mode. The specific process is as follows: If there are safety-related behaviors, the safety-first decision mode is executed. First, the current vehicle driving status data is collected, including core parameters such as vehicle speed, road conditions, gear position, and braking status, to ensure that the generated adjustment strategy is adapted to the actual driving status of the vehicle. Then, based on this vehicle driving status data, the first group adjustment strategy is generated. This strategy is a globally unified environmental adjustment strategy for the entire cabin, designed to improve driving safety and reduce external interference. Specifically, it includes three fixed adjustment operations: first, reducing the volume in the passenger area by lowering the volume of the passenger area audio system to below a preset 20 decibels; if the system is in a muted state, this remains unchanged; second, turning off the passenger area entertainment system... The system will immediately shut down all entertainment content playback on all displays in the passenger area, keeping the screens black or displaying driving safety reminders. Thirdly, it will enhance the warning intensity in the driver's area by increasing the brightness of visual warnings (instrument panel indicator lights, head-up display reminders) by 50%, the volume of auditory warnings (buzzers, voice reminders) by 30%, and the intensity of tactile warnings (steering wheel vibration, seat vibration) by 40%. Subsequently, it will suspend the execution and push of all personalized service strategies in the cabin, including all non-safety-related services such as personalized music playback, temperature adjustment, seat adjustment, and entertainment content recommendations for occupants, until safety-related behaviors are eliminated and the cabin returns to a safe state. Finally, the generated first-group adjustment strategy will be used as the final cabin environment adjustment strategy. If there are no corresponding personalized service strategies, they will be directly issued to the various execution devices in the cabin for execution.

[0056] If no safety-related behaviors are determined, a second group regulation strategy at the group level and a personalized service strategy at the individual level are generated simultaneously. These two strategies are combined to form the final cabin environment regulation and personalized service strategies. The specific implementation process consists of three steps: personalized service strategy generation, second group regulation strategy generation, and strategy combination. Specifically, when generating the personalized service strategy at the individual level, the complete information of each occupant in the behavioral feature set is first analyzed, including identity attributes (adult, child, driver, passenger), action intentions (sitting, resting, conversing, operating equipment, etc.), emotional state (calm, happy, relaxed, etc.), and spatial location (driver's seat, front passenger seat, rear left side, rear right side, etc.). Then, a preset service rule base is retrieved. This rule base is a standardized set of rules constructed based on the correspondence between occupant characteristics and service instructions, containing different identity attributes, action intentions, and emotional states. The system first assigns appropriate service instructions to occupants based on their location, ensuring all instructions are safety-verified and do not conflict with the vehicle's driving status. Then, based on each occupant's specific characteristics, it accurately matches corresponding personalized service instructions from the service rule base. A single occupant can be matched with one or more service instructions, and the service instructions for different occupants are independent of each other without overlapping interference. Finally, the personalized service instructions for all occupants are integrated according to the individual occupant, generating a unique personalized service strategy for each occupant. This ensures that the strategy is adapted to the occupant's actual needs. For example, for an occupant in the rear right seat, whose identity is a child and whose emotional state is pleasant, the system matches service instructions such as playing children's animations and adjusting the seat to a child-friendly angle. For an occupant in the front passenger seat, whose identity is an adult and whose intention is to rest, the system matches service instructions such as adjusting the seat to a reclining angle, turning off ambient lights, and lowering the local volume.

[0057] When generating the second group adjustment strategy at the group dimension, this strategy serves as a global environmental equalization adjustment strategy for the cabin. Its core is based on the number of occupants and their positional distribution within the behavioral feature set. It achieves global equalization of the cabin environment by calculating sound field equalization parameters, temperature equalization parameters, and lighting equalization parameters, ensuring a consistent environmental experience for occupants in all positions. The specific calculation process for each parameter is as follows: When calculating the sound field equalization parameters, first, the total number of occupants in the cabin is counted, along with the number of occupants in the driver's seat, front passenger seat, rear left seat, rear right seat, and rear middle seat. Then, the installation location and coverage area of ​​each speaker in the cabin are determined. The system calculates the straight-line distance from each occupant's position to each speaker. The sound field equalization parameters are calculated by first calculating the inverse distance weight for each position: inverse distance weight = percentage of occupants at each position / average straight-line distance. This weight is then normalized to obtain the sound field adjustment weight for each position, which serves as the sound field equalization parameter. The percentage of occupants at each position is the number of occupants at that position divided by the total number of occupants. The average straight-line distance from each occupant's position to each speaker is the sum of the straight-line distances from that position to all speakers divided by the total number of speakers. Based on the calculated sound field equalization parameters, the playback volume and sound effects of each speaker are adjusted to achieve global sound field equalization within the cabin.

[0058] When calculating the temperature equilibrium parameters, the baseline ambient temperature inside the cabin is first obtained, and the temperature perception requirement coefficient for each occupant's position is calculated. This coefficient is preset based on the occupant's identity attributes and action intentions. For example, the coefficient for a driver operating normally is set to 1.0, because the driver is focused on driving, has a stable metabolic rate, and needs to maintain baseline temperature perception to stay awake, so no additional weighting is required. The coefficient for a resting adult is set to 1.1, because adults have a slower metabolism and slightly lower body temperature when resting, and have a slightly higher need for a warm environment, so the coefficient needs to be increased to adapt to comfort needs. The coefficient for a child is set to 1.2, because children have weaker thermoregulation capabilities and are more sensitive to temperature changes than adults, so a higher coefficient can prioritize their temperature comfort, thus adapting to the temperature perception needs of different occupants. Then, the temperature equilibrium parameter is calculated: Temperature equilibrium parameter = Baseline ambient temperature + (0. (5 degrees Celsius × total number of occupants) - (average of the product of the number of occupants at each position and the corresponding temperature perception demand coefficient × temperature adjustment coefficient), where 0.5 degrees Celsius is the average increase in cabin temperature caused by human metabolism and heat dissipation for each additional occupant, and the temperature adjustment coefficient is the temperature adjustment amount corresponding to a unit temperature perception demand coefficient (e.g., 0.2℃), ensuring that all units are consistent in degrees Celsius; the average of the product of the number of occupants at each position and the corresponding temperature perception demand coefficient = the sum of (number of occupants at each position × corresponding temperature perception demand coefficient) ÷ total number of positions. Finally, based on the calculated temperature balance parameters, the air outlet temperature, air volume, and air outlet direction of the cabin air conditioner are adjusted to ensure that the deviation between the actual perceived temperature at each occupant position and the temperature balance parameters does not exceed ±1℃, thus achieving global temperature balance in the cabin.

[0059] When calculating the lighting balance parameters, the current ambient light intensity is first obtained, and the lighting requirement coefficient for each occupant's position is calculated. This coefficient is preset based on the occupant's intention and spatial position. For example, the coefficient for occupants operating equipment is set to 1.3, as they need to clearly see interface details and have the highest light intensity requirement. A high coefficient can increase the light intensity of the corresponding area to ensure visual clarity. The coefficient for resting occupants is set to 0.7, as they need a dim environment to avoid light interference when resting. A low coefficient can reduce ambient light and is suitable for relaxation or sleep. The coefficient for rear-seat occupants is set to 1.0, as rear-seat occupants do not have specific operational needs and their lighting requirements are at a basic balance level. This serves as a benchmark value to adapt to most daily scenarios and accommodate different occupants. The lighting requirements are determined; then the lighting balance parameters are calculated, which are: lighting balance parameter = (ambient light intensity × 0.8) + (1 + (the sum of the number of occupants at each position × the corresponding lighting requirement coefficient ÷ the total number of occupants) × lighting adjustment coefficient). 0.8 is the adaptation coefficient for introducing ambient light into the cabin, used to reduce the glare of strong light and retain the softness of natural light. The lighting adjustment coefficient is (e.g., 0.1) to ensure that the overall unit remains the light intensity unit (lux). Finally, based on the calculated lighting balance parameters, the brightness of the main lighting in the cabin, the color and brightness of the ambient lights, and the on / off status and brightness of the local lighting in each position are adjusted to achieve global lighting balance in the cabin, taking into account both the visual needs and comfort experience of the occupants.

[0060] After calculating the three types of parameters, the calculated sound field equalization parameters, temperature equalization parameters, and lighting equalization parameters are integrated to form a second group adjustment strategy at the group level. This strategy serves as the basic environmental adjustment strategy for the entire cabin, adapting to the common environmental needs of all occupants and ensuring the overall balance of the cabin environment. In the strategy combination and implementation phase, after generating the personalized service strategy and the second group adjustment strategy, the two strategies are combined. The second group adjustment strategy at the group level is used as the final cabin environment adjustment strategy, and the personalized service strategy at the individual level is used as the corresponding personalized service strategy. The combination of the two forms a strategy system for global basic environmental adjustment and individual exclusive services, serving as the basis for the final cabin environment adjustment and personalized service execution. After the strategy combination is completed, it is distributed to each execution device in the cabin. The execution order is as follows: first, the second group adjustment strategy is executed to complete the equalization adjustment of the cabin's global sound field, temperature, and lighting. Then, based on the spatial distribution of each occupant, the corresponding personalized service strategy is executed respectively, ensuring that the overall cabin environment balance and the personalized service of each occupant are taken into account, and realizing the coordinated implementation of cabin environment adjustment and personalized services.

[0061] This embodiment, through clear safety-related behavior judgment criteria, accurately identifies the behavior of drivers and passengers, quickly capturing safety risk points within the cabin, avoiding omissions or misjudgments of safety risks, and improving cabin driving safety; it sets up a dual decision-making mode of safety priority and service coordination, achieving a dynamic balance between cabin safety and service, prioritizing driving safety and minimizing safety risks; when no safety-related behaviors exist, the service coordination mode is executed, taking into account both the overall environmental balance and individual service needs, improving the passenger cabin experience; the first group adjustment strategy is designed with three fixed adjustment operations around driving safety, targeting... With its strong targeting and high execution efficiency, it can quickly reduce interference factors in the cockpit while enhancing the warning effect in the driver's area, helping the driver quickly focus on driving operations and effectively respond to various safety-related behaviors; the personalized service strategy is generated through precise matching of behavioral feature sets and preset service rule bases, which can meet the needs of each passenger's identity, actions, emotions and other characteristics, and achieve personalized and precise service; the second group adjustment strategy is generated through quantitative calculation of three equalization parameters of sound field, temperature and lighting, to ensure the balance of the overall cockpit environment, avoid the service needs of a single passenger from affecting other passengers, and achieve synergy between individual service and group experience.

[0062] In a preferred embodiment of the present invention, a cabin environment adjustment strategy and a personalized service strategy are executed, the display terminal in the cabin is controlled in conjunction with these strategies, and occupant behavior feedback data and cabin environment status data are collected after execution to iteratively optimize the interaction optimization data. The interaction optimization data includes a perception credibility coefficient, a data fusion adjustment factor, a pre-trained behavior recognition model, and decision rules for dynamic decision processing, including:

[0063] The system executes the final cabin environment adjustment strategy, coordinating the control of the air conditioning, lighting, audio, seats, and multi-zone display terminals within the cabin. Simultaneously, it executes the final personalized service strategy, providing matched services to designated occupants. Specifically, this involves: first, retrieving the final cabin environment adjustment strategy and the final personalized service strategy; and then initiating the coordinated control of various cabin equipment based on these two strategies. The core process is divided into three parts: environmental adjustment strategy execution, personalized service strategy execution, and multi-zone display terminal coordination control. The specific implementation flow is as follows: The final cabin environment adjustment strategy is executed, and based on the strategy's settings for temperature, light, and sound pressure... Precise adjustments are made to the target settings for temperature, lighting, audio, and seats. The target temperature setting is 22-26°C; the target light intensity setting is divided by area: 500-800 lux for the driver's seat, 300-500 lux for the passenger area, and 100-200 lux for the rest area; the target sound pressure level setting is ≤60 dB for the driver's seat and 40-60 dB for the passenger area; and the target seat angle setting is 100-110° for the driver's seat backrest, 120-140° for the passenger rest area, and 95-105° for the regular passenger seat backrest. Based on these target settings, the cabin's air conditioning, lighting, audio, and seats are adjusted. The system adjusts the air outlet temperature, airflow, and airflow direction according to the target temperature of each zone; adjusts the brightness and on / off status of the main lighting, ambient lighting, and local lighting according to the target light intensity of each zone; adjusts the playback volume and sound effects of each zone according to the target sound pressure level; and adjusts the tilt angle and support of the seat back and cushion according to the target angle. Simultaneously, it executes the final personalized service strategy, pushing corresponding services to designated occupants based on the exclusive service content matched to each occupant in the strategy. This includes pushing equipment operation assistance content to occupants operating the equipment, noise reduction and environmental blurring services to resting occupants, and adaptive services to child occupants. The system provides entertainment and safety alerts, with all service content precisely matched to the passenger's spatial location. It achieves coordinated control of multiple display terminals within the cabin, adjusting the driver's area display terminal, front passenger area display terminal, and rear area display terminals according to environmental adjustment and personalized service strategies. The driver's area display terminal prioritizes displaying driving safety and environmental status data, while the passenger area display terminals push matching service content based on personalized service strategies. Simultaneously, the brightness and display mode of each display terminal are coordinated with the cabin lighting intensity, enabling collaborative operation between the display terminals and other cabin equipment.

[0064] Within a set time window after strategy execution, multimodal perception data is collected again from the center points of the forward field of vision, the driver's operating plane, the passenger area sound field, and the seat bearing plane as occupant behavior feedback data. Simultaneously, actual temperature, light intensity, sound pressure level, and seat angle data for each zone within the cabin are collected as cabin environment status data. Specifically, after the strategy execution is completed, the time window is immediately started (set to 30 seconds). Within this time window, multi-dimensional data collection is conducted, collecting occupant behavior feedback data and cabin environment status data. The specific collection process and requirements are as follows: Occupant behavior feedback data is collected by selecting four core perception points within the cabin as data collection points: the center point of the forward field of vision, the center point of the driver's operating plane, the center point of the passenger area sound field, and the center point of the seat bearing plane. Within a 30-second time window, multimodal sensing data is simultaneously collected from the four collection points mentioned above. The collected multimodal sensing data includes visual image data, depth data, sound signal data, and seat pressure distribution data. This part of the multimodal sensing data is used as occupant behavior feedback data. After the collection is completed, data preprocessing is performed to ensure the integrity and validity of the data. The cabin environmental status data is collected, that is, the cabin is divided into areas according to temperature zone, illumination zone, and sound field zone. Within a 30-second time window, the actual environmental parameters of each zone are collected simultaneously, including the actual temperature value of each temperature zone, the actual light intensity value of each illumination zone, and the actual sound pressure level data of each sound field zone. At the same time, the actual angle data of all seats in the cabin are collected. The above-mentioned actual measured environmental parameters and equipment status data are used as cabin environmental status data. After the collection is completed, the data is classified and stored according to zone.

[0065] Based on occupant behavior feedback data, the average edge gradient magnitude of visual image data, point cloud density of depth data, signal-to-noise ratio of sound signal data, and spatial continuity of seat pressure distribution data within each perception subdomain are recalculated. New perception reliability coefficients are generated based on these recalculated values. Specifically, this involves classifying the collected occupant behavior feedback data by perception subdomain, quantitatively calculating the data for each subdomain, and generating new perception reliability coefficients based on the calculation results. The perception reliability coefficients quantify the effectiveness of data in each perception subdomain, serving as crucial data for perception data fusion, weighted training samples for behavior recognition models, and subsequent dynamic decision-making. This improves the accuracy of data usage and the reliability of model training and decision-making. The specific implementation process involves dividing the occupant behavior feedback data into four perception subdomains according to perception type. The system comprises four subdomains: visual image perception, depth perception, sound signal perception, and seat pressure distribution perception. Each subdomain corresponds to visual image data, depth data, sound signal data, and seat pressure distribution data, respectively. For the visual image perception subdomain, the system calculates the average edge gradient magnitude of all visual image data within the subdomain by statistically analyzing the edge gradient magnitudes and calculating the arithmetic mean of all values. For the depth perception subdomain, the system calculates the depth data by statistically analyzing the total number of depth point clouds and the effective area of ​​the point cloud distribution, dividing the total number of point clouds by the effective area to obtain the point cloud density of the depth data. For the sound signal perception subdomain, the system calculates the effective signal power and noise power of the sound signal within the subdomain by statistically analyzing the effective signal power and dividing the noise power to obtain the signal-to-noise ratio of the sound signal data.

[0066] The seat pressure distribution data of the seat pressure distribution sensing subdomain are calculated, and core sampling points are selected. Specifically, there are six key sampling points: one each in the upper, middle, and lower regions of the seat back, corresponding to the shoulder, neck, lower back, and hip positions of the occupant; and one each in the front, middle, and rear regions of the seat cushion, corresponding to the front of the thigh, buttocks, and back of the thigh. These points cover the main pressure areas when the occupant is sitting, ensuring that the sampling data accurately reflects the overall pressure distribution. Then, the pressure value difference between adjacent sampling points is calculated, followed by the arithmetic mean of all differences, using a preset baseline. Subtracting the average value from the quasi-continuity value yields the spatial continuity of the seat pressure distribution data. The preset baseline continuity value is 50 Pascals, determined based on the pressure sampling range of typical cabin seats and occupant posture characteristics. Under normal sitting posture, the difference between adjacent sampling points in the seat pressure distribution is mostly concentrated between 10 and 30 Pascals. 50 Pascals covers the pressure difference range for most occupant postures, serving as a baseline threshold for distinguishing whether the pressure distribution is uniform and avoiding distortion of spatial continuity calculation results due to extreme differences, thus ensuring the accuracy of the seat pressure distribution status assessment. Based on the above four senses... The average edge gradient magnitude, point cloud density, signal-to-noise ratio, and spatial continuity calculated for each of the three perception subdomains are weighted and integrated according to their respective weight proportions. The weight proportions for each perception subdomain are set as follows: visual image perception subdomain 0.4, depth perception subdomain 0.25, sound signal perception subdomain 0.2, and seat pressure distribution perception subdomain 0.15, with the sum of each weight proportion being 1. The visual image perception subdomain has the highest weight at 0.4 because visual data is the main basis for recognizing core behavioral characteristics such as occupant's action intentions and emotional state, and thus contributes the most to behavioral judgment. The depth perception subdomain has the second highest weight at 0.25. The spatial location and distance information it provides can assist visual data in correcting recognition biases and improving the accuracy of posture judgment. Its importance is second only to vision. The sound signal perception subdomain has a weight of 0.2 and is mainly used to capture cabin sound behavior and environmental sound effects to assist in judging the occupant status and sound field adaptation effect. Its contribution is lower than that of vision and depth. The seat pressure distribution perception subdomain has the lowest weight (0.15) and is only used to reflect the occupant's sitting posture pressure state. Its impact on overall behavior recognition and environmental perception is relatively limited. After weighted integration based on the above weights, a new perception credibility coefficient adapted to the current cabin status is generated.

[0067] Based on the difference between the actual measured values ​​of each zone in the cabin environment status data and the target setpoints in the final cabin environment adjustment strategy, a correction coefficient for the data fusion adjustment factor is calculated, and the correction coefficient is used to update the current data fusion adjustment factor. Specifically, this involves: using the collected cabin environment status data and the executed final cabin environment adjustment strategy as a basis, calculating the difference between the actual measured values ​​and the target setpoints to obtain the correction coefficient for the data fusion adjustment factor, and then using the correction coefficient to update the current data fusion adjustment factor. The specific implementation process is as follows: extracting the actual measured values ​​of each zone in the cabin environment status data, including... This includes actual temperature, actual light intensity, actual sound pressure level, and actual seat angle data for each zone. Simultaneously, it extracts the corresponding target settings for each zone from the final cabin environment adjustment strategy, where the target temperature setting is 24℃, the target light intensity setting is 400 lux, the target sound pressure level setting is 50 dB, and the target seat angle setting is 100°. The deviation between the actual measured values ​​and the target settings for each zone is calculated according to the environmental parameter type. The deviation is calculated by subtracting the target setting from the actual measured value, yielding the temperature deviation, light intensity deviation, sound pressure level deviation, and seat angle deviation. These deviations are then calculated separately. The absolute values ​​of all zone deviations under each environmental parameter type are calculated, and then the arithmetic mean of the absolute values ​​under each parameter type is obtained to obtain the average deviations for temperature, light intensity, sound pressure level, and seat angle. The average deviation of each environmental parameter type is divided by its corresponding maximum permissible deviation to obtain the normalized deviation for each parameter. The normalized deviations are then weighted and summed according to the weight percentage of each environmental parameter type to obtain the comprehensive deviation score, ranging from 0 to 1. The weight percentages for each environmental parameter type are set as follows: temperature 0.4, light intensity 0.3, sound pressure level 0.2, and seat angle 0.1. The sum of the weight percentages is... The weighting is 1. This weighting value is determined by combining the core priority of each environmental parameter in cabin environment adjustment and its impact on passenger experience. Temperature is a fundamental core parameter of the cabin environment, directly affecting the overall comfort and physical condition of passengers, so its weighting is set to the highest at 0.4. Light intensity has a significant impact on passengers' visual experience and ease of operation, and its importance is second only to temperature, so its weighting is set to 0.3. Sound pressure level mainly affects the cabin's auditory experience and environmental quietness, and its impact on overall environmental adaptability is relatively low, so its weighting is set to 0.2. Seat angle only affects the individual passenger's sitting posture experience, and its impact range is relatively limited, so its weighting is set to 0.1. Based on the comprehensive deviation score, subtract the comprehensive deviation score from 1 to obtain the correction coefficient of the data fusion adjustment factor. The maximum allowable deviation for each environmental parameter is ±3℃ for temperature, ±100 lux for light intensity, ±10 dB for sound pressure level, and ±10° for seat angle. Multiply the current data fusion adjustment factor by the correction coefficient to obtain the updated data fusion adjustment factor. This correction coefficient is essentially a dynamic calibration coefficient generated based on the degree of environmental adjustment deviation. The smaller the deviation, the closer the correction coefficient is to 1; the larger the deviation, the further the correction coefficient deviates from 1. Its core function is to feed back the actual accuracy of environmental adjustment to the data fusion adjustment factor. Therefore, multiplying the current data fusion adjustment factor by this correction coefficient achieves proportional calibration of the adjustment factor. That is, when the adjustment effect is good and the deviation is small, the adjustment factor remains basically stable; when the adjustment deviation is large and the effect is poor, the adjustment factor is corrected synchronously, so that the updated factor can accurately adapt to the fusion requirements of the current cabin environmental data. The updated data fusion adjustment factor is obtained through this multiplication operation, completing the iterative update of the data fusion adjustment factor.

[0068] Using a new perceived credibility coefficient and an updated data fusion adjustment factor, combined with occupant behavior feedback data, the pre-trained behavior recognition model is trained with updated parameters. Based on the policy execution effects reflected by occupant behavior feedback data and cabin environment status data, the judgment logic and policy parameters of the safety-first decision-making mode and the service-coordination decision-making mode in dynamic decision processing are adaptively adjusted. Specifically, this includes: training the pre-trained behavior recognition model with updated parameters using the new perceived credibility coefficient and the updated data fusion adjustment factor; and adaptively adjusting the judgment logic and policy parameters of dynamic decision processing based on the policy execution effects. The specific process involves training the pre-trained behavior recognition model with updated parameters, which involves collecting... Occupant behavior feedback data serves as the core training samples for model updates. New perceptual confidence coefficients are used as weight coefficients for the training samples, assigned according to preset weights for each perceptual subdomain: visual image perception subdomain 0.4, depth perception subdomain 0.25, sound signal perception subdomain 0.2, and seat pressure distribution perception subdomain 0.15. The visual image perception subdomain has the highest weight at 0.4 because visual data is the primary basis for recognizing core behavioral features such as occupant intentions, emotional states, and posture changes, playing a decisive role in the accuracy of behavior judgment. The depth perception subdomain has the next highest weight at 0.25; its spatial position, distance, and three-dimensional shape information can assist visual data in correcting planar recognition biases and improving the accuracy of occupant posture judgment, thus increasing its importance. Second only to vision; the sound signal perception subdomain has a weight of 0.2, mainly used to capture in-cabin voice commands, emotional sounds, and environmental sound effects, assisting in judging the occupant's state and sound field adaptation effect. Its contribution to overall behavior recognition is lower than that of vision and depth. The seat pressure distribution perception subdomain has the lowest weight of 0.15, only used to reflect the occupant's sitting posture pressure distribution and body stability. Its influence on global behavior recognition and decision-making is relatively limited, hence its lowest weight. This weight is used to assign appropriate weight values ​​to samples in each perception subdomain, strengthening the influence of high-confidence perception data on model parameter updates and weakening the interference of low-confidence perception data, ensuring that the effectiveness of training samples matches the actual perception state of the cabin. The updated data fusion adjustment factor is used as the model training factor. The fusion coefficient coordinates the fusion ratio by quantitatively allocating the fusion weights of multimodal sensing data. Specifically, based on the fusion coefficient and combined with the weight ratio of each sensing subdomain, the fusion weights of visual images, depth, sound signals, and seat pressure distribution data are dynamically calibrated. At the same time, the actual deviation of cabin environment adjustment is incorporated into the multimodal data fusion process. If the fusion coefficient is high (indicating small environmental adjustment deviation and high data fusion accuracy), the fusion synergy of each modality data is increased according to the preset subdomain weight ratio to achieve data complementarity. If the fusion coefficient is low (indicating large environmental adjustment deviation and insufficient reliability of some data), the preset weights of each sensing subdomain are used (visual images 0.4, depth 0.25, sound signals 0.2, seat pressure distribution 0.15) Based on the base value, the preset weight of each subdomain is multiplied by the product of the fusion coefficient and the perceived credibility coefficient of the corresponding subdomain to obtain the dynamic fusion weight of each modality. For low-credibility modality data (low perceived credibility coefficient), the product of the fusion coefficient and the base value is smaller, and the corresponding dynamic fusion weight will be further reduced compared to the preset base value, thereby reducing the fusion ratio. For high-credibility modality data (high perceived credibility coefficient), the product of the fusion coefficient and the base value is larger, and the corresponding dynamic fusion weight will remain at a higher level, thus prioritizing the strengthening of the fusion weight of high-credibility modality data. Through this dynamic calibration method, To adapt multimodal data fusion to the actual environment of the current cockpit and improve the accuracy of multimodal data collaborative training, the training samples, weighted and fused with coefficients, are input into a pre-trained behavior recognition model. Gradient descent is used to iteratively update the parameters of core layers such as convolutional and fully connected layers. Through repeated training via multiple rounds of forward and backward propagation, the model's behavior recognition accuracy reaches a preset threshold. In this step, the preset threshold for behavior recognition accuracy is set to 95%. When the model's behavior recognition accuracy reaches this threshold, the parameter update training of the pre-trained behavior recognition model is complete, resulting in the optimized behavior recognition model.

[0069] Combining collected occupant behavior feedback data and cabin environment status data, the effectiveness of strategy implementation was assessed from both aspects. The target settings for cabin environment status were: actual temperature deviation from the target value in each zone ≤ ±1℃, actual light intensity deviation from the target value ≤ ±50 lux, actual sound pressure level deviation from the target value ≤ ±5 dB, and actual seat angle deviation from the target value ≤ ±5°. Positive occupant behavior feedback was characterized by no resistant actions, no environmental discomfort-related behaviors, and good service acceptance. If the actual cabin environment status... If the above target requirements are met and passenger behavior shows positive feedback, the strategy is considered to be performing well; otherwise, the strategy is considered to be performing poorly. The dynamic decision-making logic is adaptively adjusted: based on the strategy's performance, the switching logic between the safety-first decision-making mode and the service-coordination decision-making mode is adjusted, optimizing the thresholds and dimensions for safety-related behaviors. Specifically, the thresholds are: driver fatigue lasting ≥5 seconds, driver distraction lasting ≥3 seconds in a single instance or occurring twice or more within 10 seconds, and passenger dangerous actions occurring in a single instance. The system triggers a judgment, supplementing the judgment dimensions with behavioral characteristics, time, and environmental association. Behavioral characteristics include the occupant's intentions and emotional state; time includes the duration and frequency of the behavior; and environmental association includes the correlation between the behavior and the vehicle's driving status and cabin environment. If a safety-related behavior is determined to exist but dangerous occupant behavior still occurs after the strategy is executed, or if no safety-related behavior is determined to exist but the occupant experiences discomfort, further optimization is performed based on the aforementioned judgment thresholds and dimensions. Feature indicators for behavior recognition are added to improve the accuracy of mode switching judgments. Based on the strategy execution effect, the strategy parameters under the two decision-making modes are adjusted. If the driver's area warning intensity does not achieve the expected warning effect under the safety-first decision-making mode, the intensity parameters of visual, auditory, and tactile warnings are increased. If the cabin environment balance parameters do not meet expectations under the service collaboration decision-making mode, the calculation weights and baseline values ​​of sound field, temperature, and lighting balance parameters are adjusted. If the service content of the personalized service strategy has a low match with the occupant's needs, the matching relationship between occupant characteristics and service instructions in the service rule base is optimized to complete the adaptive adjustment of the dynamic decision-making processing strategy parameters.

[0070] This embodiment achieves coordinated execution of cabin environment adjustment and personalized service strategies through the coordinated control of cabin air conditioning, lighting, audio, seats, and multi-zone display terminals. This ensures both the uniformity and coordination of the overall cabin environment and provides precisely matched personalized services to designated occupants, enhancing the intelligence and precision of cabin environment adjustment and service provision. Occupant behavior feedback data is collected from core cabin sensing points, and cabin environment status data is collected from various zones. The selection of data points and data types balances comprehensiveness and specificity, accurately reflecting the actual cabin state and occupant behavior feedback after strategy execution. Quantitative calculations of data from each sensing subdomain generate new sensing credibility coefficients, enabling a quantitative assessment of the effectiveness of sensing data and ensuring the precise adaptation of the sensing credibility coefficients. The current cockpit perception environment has improved the reliability of perception data. By calculating the difference between the actual measured values ​​of the cockpit environment and the target set values, correction coefficients are obtained and the data fusion adjustment factors are updated. This allows the data fusion adjustment factors to dynamically change according to the actual effect of the cockpit environment adjustment, improving the accuracy of multi-source environmental data fusion. The behavior recognition model is trained with updated parameters using the new perception credibility coefficient and the updated data fusion adjustment factors, optimizing the model's behavior recognition capabilities and improving the accuracy and real-time performance of the model in recognizing occupant behavior. At the same time, the judgment logic and strategy parameters of dynamic decision processing are adaptively adjusted according to the strategy execution effect, making the switching between safety priority and service collaboration dual decision modes more precise and the strategy parameters more adapted to the actual state of the cockpit, realizing iterative optimization of dynamic decision rules.

[0071] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A dynamic interactive system for a bus intelligent cockpit based on multi-occupant behavior recognition, characterized in that, include: The perception module is used to collect multimodal perception data from the center point of the forward field of view in the passenger cabin, the center point of the driver's seat operating plane, the center point of the sound field in the passenger area, and the center point of the seat bearing plane. The computation module is used to construct a virtual space topology within the cockpit based on four center points and multimodal perception data. This virtual space topology is divided into multiple perception subdomains, and a perception confidence coefficient is generated for each subdomain. For each center point, a corresponding 3D perception geometric model is established, and the spatial intersections of the occupant behavior direction vector with each 3D perception geometric model are calculated. A data fusion adjustment factor is generated for each perception subdomain, including: using the forward field of view center point, the driver's operating plane center point, the passenger area sound field center point, and the seat bearing plane center point as geometric vertices, and constructing a closed virtual space topology connecting these vertices based on the maximum effective perception distance and field of view of the sensing devices corresponding to each center point. Based on the spatial envelope formed by the field of view boundaries of each sensing device within the virtual space topology, the virtual space topology is divided into multiple non-overlapping and continuous spatial regions, each defined as a perception subdomain. For each perception subdomain, a data fusion adjustment factor is generated. The average edge gradient magnitude, point cloud density of depth data, signal-to-noise ratio of sound signal data, and spatial continuity of seat pressure distribution data of visual image data collected within its coverage area are considered. A quantified perception confidence coefficient is generated based on the average edge gradient magnitude, point cloud density, signal-to-noise ratio, and spatial continuity. For each center point, a three-dimensional geometry representing its actual perception space is constructed as a three-dimensional perception geometric model according to the installation position, orientation, and perception range parameters of its corresponding sensing device. For continuous frame visual image data collected from the center point of the forward field of view and the center point of the driver's seat operating plane, the occupant behavior direction vector representing the occupant's movement trend is obtained through pixel displacement calculation. The set of intersection points between each occupant behavior direction vector and the surface of each three-dimensional perception geometric model is calculated, and the spatial distribution characteristics of the intersection point set belonging to the same perception subdomain are statistically analyzed. Based on the density and geometric regularity of the spatial distribution characteristics, a data fusion adjustment factor for weighted fusion is generated. The recognition module is used to combine the perception confidence coefficient of each perception subdomain with the data fusion adjustment factor to fuse and register the multimodal perception data collected from the center point, generate a unified perception data stream, and input the unified perception data stream into the pre-trained behavior recognition model for processing to obtain the behavior feature set of each occupant. The decision-making module is used to calculate the final cabin environment adjustment strategy and personalized service strategy based on the set of behavioral characteristics and the vehicle's driving status through dynamic decision processing. The execution module is used to execute cabin environment adjustment strategies and personalized service strategies, control the display terminals in the cabin in a coordinated manner, and collect passenger behavior feedback data and cabin environment status data after execution in order to iteratively optimize the interaction optimization data.

2. The dynamic interaction system for intelligent cockpits in buses based on multi-occupant behavior recognition according to claim 1, characterized in that, The process of acquiring the multimodal sensing data includes: Visual image data and depth data are collected from the center point of the forward field of view and the center point of the driver's seat operating plane; sound signal data is collected from the center point of the passenger area sound field; seat pressure distribution data is collected from the center point of the seat bearing plane; the visual image data, depth data, sound signal data and seat pressure distribution data together constitute multimodal perception data.

3. The dynamic interaction system for intelligent cockpits in buses based on multi-occupant behavior recognition according to claim 2, characterized in that, Based on the density and geometric regularity of spatial distribution characteristics, a data fusion adjustment factor is generated for weighted fusion, including: Calculate the standard deviation of the three-dimensional coordinates of all spatial intersection points belonging to the same perceptual subdomain, and use the reciprocal of the standard deviation as a density index to quantify the degree of density. In the set of spatial intersection points, the direction consistency of each spatial intersection point with the starting point of the corresponding occupant behavior direction vector is defined, and the value of the direction consistency is used as a regularity index to quantify the geometric regularity. The product of density and regularity indicators is normalized to generate a data fusion adjustment factor.

4. The dynamic interaction system for intelligent cockpits in buses based on multi-occupant behavior recognition according to claim 3, characterized in that, By combining the perception reliability coefficients of each perception subdomain with the data fusion adjustment factor, the multimodal perception data collected from the center point is fused and registered to generate a unified perception data stream. This unified perception data stream is then input into a pre-trained behavior recognition model for processing, yielding the behavior feature sets of each occupant, including: For each perception subdomain, the perception confidence coefficient of the corresponding perception subdomain is multiplied by the data fusion adjustment factor to obtain the data fusion weight of that perception subdomain. For each sensing subdomain, multimodal sensing data belonging to the same time window collected from the center point within its spatial range are acquired. Using the data fusion weight of this sensing subdomain, a weighted average is calculated on the multimodal sensing data. The weighted average visual image data, depth data, sound signal data, and seat pressure distribution data are then normalized to obtain normalized multimodal sensing data. The normalized multimodal sensing data is then timestamped and aligned with the spatial coordinate system to generate a unified sensing data stream. The unified perception data stream is input into the pre-trained behavior recognition model in time series, and the output is a set of behavior features including the number of occupants, identity attributes, action intentions and emotional states.

5. The dynamic interaction system for intelligent cockpits in buses based on multi-occupant behavior recognition according to claim 4, characterized in that, The unified perception data stream is input into a pre-trained behavior recognition model in time series order, and the output is a set of behavior features including the number of occupants, their identity attributes, action intentions, and emotional states, including: The unified perception data stream in the time series is input into the convolutional neural network layer of the pre-trained behavior recognition model to extract the spatial features of visual image data and depth data, as well as the spectral features of sound signal data. The spatial features and spectral features are then concatenated with the temporal features of seat pressure distribution data to form a fused feature vector. The fused feature vectors are input into the long short-term memory network layer of the pre-trained behavior recognition model to obtain temporal context features. Based on the temporal context features, multiple independent feature vectors corresponding to each occupant are obtained through feature separation processing in the pre-trained behavior recognition model. Each independent feature vector is input into the identity attribute classification layer, action intention classification layer, and emotional state classification layer of the pre-trained behavior recognition model to obtain the identity attribute, action intention, and emotional state of each occupant; the number of independent feature vectors is counted to obtain the number of occupants; the number of occupants, the identity attribute, action intention, and emotional state of each occupant together constitute the behavior feature set.

6. The dynamic interaction system for intelligent cockpit of a bus based on multi-occupant behavior recognition according to claim 5, characterized in that, Based on the set of behavioral characteristics and vehicle driving status, dynamic decision processing is used to calculate and form the final cabin environment adjustment strategy and personalized service strategy, including: Based on the set of behavioral characteristics, it is determined whether there are safety-related behaviors, which include at least one of driver fatigue, driver distraction, and dangerous actions by passengers. If safety-related behaviors are present, the system enters a safety-priority decision-making mode. Based on the vehicle's driving status, a first group adjustment strategy is generated as the final cabin environment adjustment strategy, while all personalized service strategies are suspended. The first group adjustment strategy includes reducing passenger area volume, turning off passenger area entertainment displays, and increasing driver area warning intensity. If no safety-related behaviors are present, the system enters a service collaboration decision-making mode, generating a second group adjustment strategy at the group level and a corresponding personalized service strategy. The first or second group adjustment strategy is then combined with the corresponding personalized service strategy to form the final cabin environment adjustment strategy and personalized service strategy.

7. The dynamic interaction system for intelligent cockpits in buses based on multi-occupant behavior recognition according to claim 6, characterized in that, The process of determining the second group regulation strategy at the group dimension includes: Based on the identity attributes, action intentions, and emotional states of each occupant in the behavioral feature set, corresponding personalized service instructions are matched from the preset service rule base to generate a personalized service strategy at the individual level. At the same time, based on the number of occupants and the location distribution of each occupant in the behavioral feature set, the cabin sound field equalization parameters, temperature equalization parameters, and lighting equalization parameters are calculated to generate a second group adjustment strategy at the group level.

8. The dynamic interaction system for intelligent cockpit of a bus based on multi-occupant behavior recognition according to claim 7, characterized in that, Implement cabin environment adjustment and personalized service strategies, control the display terminals in the cabin in a coordinated manner, and collect occupant behavior feedback data and cabin environment status data after implementation to iteratively optimize the interaction data, including: The system implements the final cabin environment adjustment strategy, which involves coordinated control of the cabin's air conditioning, lighting, audio, seats, and multi-zone display terminals. Simultaneously, it implements the final personalized service strategy, providing matched service content to designated occupants. Within the set time window after the strategy is executed, multimodal perception data is collected again from the center point of the forward field of vision, the center point of the driver's seat operating plane, the center point of the passenger area sound field, and the center point of the seat bearing plane as occupant behavior feedback data. Simultaneously, the actual temperature value, actual light intensity value, actual sound pressure level data, and seat angle data of each zone in the cabin are collected as cabin environment status data. Based on occupant behavior feedback data, the average edge gradient magnitude of visual image data, point cloud density of depth data, signal-to-noise ratio of sound signal data, and spatial continuity of seat pressure distribution data in each perception subdomain are recalculated, and a new perception credibility coefficient is generated based on the recalculated values. Based on the difference between the actual measured values ​​of each zone in the cabin environment status data and the target set values ​​in the final cabin environment adjustment strategy, the correction coefficient of the data fusion adjustment factor is calculated, and the correction coefficient is used to update the current data fusion adjustment factor. By utilizing the new perceived credibility coefficient and the updated data fusion adjustment factor, combined with occupant behavior feedback data, the pre-trained behavior recognition model is trained with updated parameters. Based on the policy execution effect reflected by occupant behavior feedback data and cabin environment status data, the judgment logic and policy parameters of the safety priority decision mode and service collaboration decision mode in dynamic decision processing are adaptively adjusted.

9. The dynamic interaction system for intelligent cockpit of a bus based on multi-occupant behavior recognition according to claim 8, characterized in that, The interactive optimization data includes the perceived credibility coefficient, data fusion adjustment factor, pre-trained behavior recognition model, and decision rules for dynamic decision processing.

Citation Information

Patent Citations

  • Intelligent cabin active interaction system and method, electronic equipment and storage medium

    CN114604191A

  • Multi-screen voice interaction system and method applied to automobile cabin

    CN120299457A