Intelligent seat adjustment control method, system and device based on machine vision and medium
Through the intelligent seat adjustment and control method based on machine vision, the age group of occupants is automatically identified and the seat locking or unlocking is controlled, which solves the safety hazards caused by the single seat locking method in the prior art, and improves driving safety and driving experience.
Patent Information
- Application Number
- CN202510590693.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-27
AI Technical Summary
The locking method of the existing electric adjustment function of the car seat is single, which can easily lead to misoperation of children or distracted drivers to unlock, posing safety hazards.
Using an intelligent seat adjustment and control method based on machine vision, the seat is automatically controlled to be locked or unlocked by obtaining the depth image information of the passenger compartment, detecting the occupant's body and identifying its age group.
It realizes automatic locking and unlocking of smart seats, improves driving safety, avoids safety hazards caused by misoperation of children or the elderly, and reduces the driver's operating frequency.
Smart Images

Figure CN120207176A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle control, and in particular to an intelligent seat adjustment control method, system, device and medium based on machine vision. Background Art
[0002] With the development of automotive intelligent networking, vehicle monitoring and control technologies have become increasingly intelligent, bringing richer driving experiences to passengers. The electric seat adjustment function of traditional vehicles is realized through physical buttons on the seat, and the locking method of the adjustment function is relatively single. After the intelligent upgrade of new energy vehicles, functions such as adjusting the seat through a touch screen and voice control are added, and a locking button for the seat adjustment function is added. However, there are still the following problems:
[0003] 1) The locking of the electric seat adjustment function mainly depends on the driver issuing corresponding instructions, and it is easy to forget to lock, so that children in the vehicle can adjust the seat electrically at will or accidentally trigger the seat electric adjustment, posing a safety hazard of pinching passengers;
[0004] 2) After the electric seat adjustment function is locked, if the child occupant of the corresponding seat is replaced by an adult occupant and the adult occupant wants to adjust the seat, the driver still needs to actively unlock it. Unlocking during driving also brings inconvenience to the driver, may cause the driver to be distracted, and there is also a safety hazard. Summary of the Invention
[0005] An object of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.
[0006] To this end, an object of an embodiment of the present invention is to provide an intelligent seat adjustment control method based on machine vision, which realizes automatic locking and unlocking of the intelligent seat, improving driving safety and the driving experience of users.
[0007] Another object of an embodiment of the present invention is to provide an intelligent seat adjustment control system based on machine vision.
[0008] To achieve the above technical object, the technical solutions adopted in the embodiments of the present invention include:
[0009] In a first aspect, an embodiment of the present invention provides an intelligent seat adjustment control method based on machine vision, including the following steps:
[0010] Obtain depth image information of the passenger compartment of the target vehicle, perform human detection on the depth image information to obtain a human depth image of the target occupant, and determine the target cockpit position corresponding to the target occupant;
[0011] Generate human depth image time series data based on consecutive multiple frames of the human depth images, and input the human depth image time series data into a pre-trained age group recognition model to obtain the age group of the target occupant;
[0012] When the age group of the occupant meets the preset child age group range or the elderly age group range, control the target seat corresponding to the target cockpit position to be in a locked state;
[0013] When the age group of the occupant does not meet the child age group range and the elderly age group range, control the target seat corresponding to the target cockpit position to be in an unlocked state.
[0014] Further, in an embodiment of the present invention, the obtaining the depth image information of the occupant compartment of the target vehicle, performing human body detection on the depth image information to obtain the human depth image of the target occupant, and determining the target cockpit position corresponding to the target occupant specifically includes:
[0015] Obtain the depth image information through the binocular camera of the occupant compartment;
[0016] Perform human body key point detection on the depth image information to obtain multiple human body key points and corresponding key point depth information;
[0017] Extract the human depth image from the depth image information according to the human body key points;
[0018] Perform 3D space mapping on the human body key points according to the key point depth information to obtain the human body position coordinates of the target occupant;
[0019] Obtain the preset cockpit position distribution information, and determine the target cockpit position corresponding to the target occupant according to the human body position coordinates and the cockpit position distribution information.
[0020] Further, in an embodiment of the present invention, the age group recognition model is trained through the following steps:
[0021] Obtain the human depth image time series samples of multiple testers in the test vehicle, and determine the age group labels corresponding to each human depth image time series sample through manual annotation;
[0022] Input the human depth image time series samples into a pre-constructed CNN-LSTM hybrid neural network to obtain an age group recognition result;
[0023] Determine the loss value according to the age group recognition result and the age group label;
[0024] Update the parameters of the CNN-LSTM hybrid neural network according to the loss value through the backpropagation algorithm to obtain the trained age group recognition model.
[0025] Further, in an embodiment of the present invention, the CNN-LSTM hybrid neural network includes an input layer, a 3D CNN convolutional layer, a 2D CNN convolutional layer, a feature fusion layer, an LSTM layer, an attention layer, and an output layer. The input layer is used to split the human depth image time series sample to obtain 3D depth image time series data and RGB image time series data. The 3D CNN convolutional layer is used to extract features from the 3D depth image time series data to obtain a depth space feature sequence. The 2D CNN convolutional layer is used to extract features from the RGB image time series data to obtain a local area feature sequence. The feature fusion layer is used to perform feature fusion on the depth space feature sequence and the local area feature sequence to obtain a spatio-temporal feature sequence. The LSTM layer is used to generate a hidden state sequence according to the spatio-temporal feature sequence. The attention layer is used to perform dynamic weight allocation on each dimension of the hidden state sequence based on the multi-head self-attention mechanism. The output layer is used to map the hidden state sequence after dynamic weight allocation to the age group recognition result.
[0026] Further, in an embodiment of the present invention, inputting the human depth image time series sample into a pre-constructed CNN-LSTM hybrid neural network to obtain an age group recognition result specifically includes:
[0027] Split the human depth image time series sample through the input layer to obtain the 3D depth image time series data and the RGB image time series data;
[0028] Extract features from the 3D depth image time series data through the 3D CNN convolutional layer to obtain the depth space feature sequence;
[0029] Extract features from the RGB image time series data through the 2D CNN convolutional layer to obtain the local area feature sequence;
[0030] Perform spatio-temporal alignment and feature fusion on the depth space feature sequence and the local area feature sequence through the feature fusion layer to obtain the spatio-temporal feature sequence;
[0031] Determine the spatio-temporal feature vectors of each time step according to the spatio-temporal feature sequence, calculate the hidden state vectors corresponding to each spatio-temporal feature vector through the forget gate, input gate, and output gate of the LSTM layer, and generate the hidden state sequence according to the hidden state vectors;
[0032] Dynamically allocate weights to each dimension of the hidden state sequence through the attention layer based on the multi-head self-attention mechanism;
[0033] Map the hidden state sequence after dynamic weight allocation to the age group recognition result through the output layer.
[0034] Further, in an embodiment of the present invention, controlling the target seat corresponding to the target cockpit position to be in a locked state specifically includes:
[0035] When the target cockpit position is the left rear seat, determine the left rear seat as the target seat, and send a control command to the left domain controller through the central controller, so that the left domain controller controls the left rear seat to be locked;
[0036] When the target cockpit position is the right rear seat, determine the right rear seat as the target seat, and send a control command to the right domain controller through the central controller, so that the right domain controller controls the right rear seat to be locked.
[0037] Further, in an embodiment of the present invention, the intelligent seat adjustment control method further includes the following steps:
[0038] In response to the driver's seat unlocking instruction, unlock the target seat in the locked state;
[0039] In response to the driver's seat locking instruction, lock the target seat in the unlocked state;
[0040] Wherein, the seat unlocking instruction and the seat locking instruction are one of a key instruction, a touch instruction, a voice instruction, and a gesture instruction.
[0041] In a second aspect, an embodiment of the present invention provides an intelligent seat adjustment control system based on machine vision, including:
[0042] An image acquisition module, configured to obtain depth image information of the passenger compartment of the target vehicle, perform human detection on the depth image information to obtain a human depth image of the target occupant, and determine the target cockpit position corresponding to the target occupant;
[0043] An age group recognition module, configured to generate human depth image time series data according to a plurality of consecutive frames of the human depth images, and input the human depth image time series data into a pre-trained age group recognition model to obtain the occupant age group of the target occupant;
[0044] A locking control module, configured to control the target seat corresponding to the target cockpit position to be in a locked state when the occupant age group meets a preset child age group interval or an elderly age group interval;
[0045] An unlocking control module, configured to control the target seat corresponding to the target cockpit position to be in an unlocked state when the age range of the occupant does not conform to the child age range and the elderly age range.
[0046] In a third aspect, an embodiment of the present invention provides an intelligent seat adjustment control device based on machine vision, including:
[0047] At least one processor;
[0048] At least one memory, configured to store at least one program;
[0049] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned intelligent seat adjustment control method based on machine vision.
[0050] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to execute the above-mentioned intelligent seat adjustment control method based on machine vision when executed by the processor.
[0051] The advantages and beneficial effects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention:
[0052] In an embodiment of the present invention, depth image information of the occupant compartment of a target vehicle is acquired, human detection is performed on the depth image information to obtain a human depth image of a target occupant, and a target cockpit position corresponding to the target occupant is determined. Human depth image time series data is generated according to a continuous multi-frame human depth image, and the human depth image time series data is input into a pre-trained age range recognition model to obtain the age range of the target occupant. When the age range of the occupant conforms to a preset child age range or elderly age range, the target seat corresponding to the target cockpit position is controlled to be in a locked state. When the age range of the occupant does not conform to the child age range and the elderly age range, the target seat corresponding to the target cockpit position is controlled to be in an unlocked state. In an embodiment of the present invention, the target occupant is detected based on the depth image information of the occupant compartment and the age range of the occupant is recognized. For occupants in the child age range or the elderly age range, the seat adjustment function corresponding to the corresponding cockpit position is controlled to be in a locked state. For occupants not in the child age range and the elderly age range, the seat adjustment function corresponding to the corresponding cockpit position is controlled to be in an unlocked state, realizing automatic locking and unlocking of the intelligent seat, avoiding potential safety hazards caused by child occupants or elderly occupants accidentally operating the seat adjustment function, and at the same time facilitating other occupants to adjust the seat according to their own needs, without the driver manually performing frequent operations of locking and unlocking the seat, improving driving safety and the driving experience of users. Brief Description of the Drawings
[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following introduces the drawings required to be used in the embodiments of the present invention. It should be understood that the drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0054] Figure 1 It is a flowchart of the steps of an intelligent seat adjustment control method based on machine vision provided by an embodiment of the present invention;
[0055] Figure 2 It is a schematic structural diagram of a CNN-LSTM hybrid neural network provided by an embodiment of the present invention;
[0056] Figure 3 It is a block diagram of the structure of an intelligent seat adjustment control system based on machine vision provided by an embodiment of the present invention;
[0057] Figure 4 It is a block diagram of the structure of an intelligent seat adjustment control device based on machine vision provided by an embodiment of the present invention. Detailed Embodiments
[0058] The following details the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0059] In the description of the present invention, the meaning of "a plurality of" is two or more. If the first and second are described, it is only for the purpose of distinguishing technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention.
[0060] Referring to Figure 1 , an embodiment of the present invention provides an intelligent seat adjustment control method based on machine vision, which specifically includes the following steps:
[0061] S101. Obtain the depth image information of the passenger compartment of the target vehicle, perform human detection on the depth image information to obtain the human depth image of the target occupant, and determine the target cockpit position corresponding to the target occupant;
[0062] S102. Generate human depth image time-series data based on consecutive multi-frame human depth images, and input the human depth image time-series data into a pre-trained age group recognition model to obtain the age group of the target occupant;
[0063] S103. When the age group of the occupant meets the preset children's age group range or the elderly age group range, control the target seat corresponding to the target cockpit position to be in a locked state;
[0064] S104. When the age group of the occupant does not meet the children's age group range and the elderly age group range, control the target seat corresponding to the target cockpit position to be in an unlocked state.
[0065] In the embodiment of the present invention, the target occupant is detected based on the depth image information of the passenger compartment and the age group of the occupant is recognized. For the occupant in the children's age group or the elderly age group, the seat adjustment function corresponding to the corresponding cockpit position is controlled to be in a locked state. For the occupant not in the children's age group and the elderly age group, the seat adjustment function corresponding to the corresponding cockpit position is controlled to be in an unlocked state, realizing the automatic locking and unlocking of the intelligent seat, avoiding the safety hazards brought by the misoperation of the seat adjustment function by child occupants or elderly occupants, and at the same time facilitating other occupants to adjust the seat according to their own needs, without the driver manually performing frequent operations of locking and unlocking the seat, improving the driving safety and the driving experience of users.
[0066] Further as an optional implementation manner, obtaining the depth image information of the passenger compartment of the target vehicle, performing human detection on the depth image information to obtain the human depth image of the target occupant, and determining the target cockpit position corresponding to the target occupant specifically includes:
[0067] S1011. Obtain the depth image information through the binocular camera in the passenger compartment;
[0068] S1012. Perform human key point detection on the depth image information to obtain multiple human key points and the corresponding key point depth information;
[0069] S1013. Extract the human depth image from the depth image information according to the human key points;
[0070] S1014. Perform 3D space mapping on the human key points according to the key point depth information to obtain the human position coordinates of the target occupant;
[0071] S1015. Obtain the preset cockpit position distribution information, and determine the target cockpit position corresponding to the target occupant according to the human body position coordinates and the cockpit position distribution information.
[0072] Specifically, in the embodiment of the present invention, the binocular camera is used to obtain the depth image information of the occupant cabin, the human body depth image of the occupant is extracted based on the detection of human body key points, and the cockpit position where the occupant is located is determined in combination with the depth information. The specific process is as follows:
[0073] 1) Acquisition of binocular camera depth image
[0074] (1) Principle: Based on the binocular stereo vision technology, the depth is calculated through the baseline distance (Baseline) and disparity (Disparity) of the left and right cameras. The formula is:
[0075]
[0076] Among them, f is the focal length, T is the baseline length, and d is the disparity (horizontal offset of corresponding pixel points in the left and right images).
[0077] (2) Implementation steps
[0078] Binocular calibration: Obtain the internal parameters (focal length, distortion coefficient) and external parameters (baseline distance, relative pose) of the camera through the calibration board.
[0079] Stereo matching: Use algorithms such as BM and SGBM to match the corresponding pixel points of the left and right images to generate a disparity map (DisparityMap).
[0080] Depth conversion: Calculate the depth value of each pixel according to the disparity map to generate a depth image (Depth Map).
[0081] 2) Detection and depth association of human body key points
[0082] Key point detection: Adopt a deep learning model (such as OpenPose, MediaPipe) to detect the key points of the human body skeleton (such as shoulders, elbows, wrists, hips, etc.) and output 2D coordinates.
[0083] Depth information association: Map the 2D coordinates of the key points into the depth image and extract the depth values at the corresponding positions.
[0084] Optimization strategy: Perform weighted averaging on the area around the key points to reduce the influence of noise.
[0085] 3) Extraction of human body depth image
[0086] Region segmentation: Generate a human body mask (Mask) based on the key points, and segment the human body contour in combination with edge detection (such as the Canny algorithm).
[0087] Depth map cropping: Retain the depth information within the mask area, eliminate background noise, and generate a depth image that only contains the target occupant, which is the human body depth image.
[0088] 4) 3D space mapping of human body key points
[0089] Coordinate transformation: According to the camera internal parameter matrix and depth value, convert the 2D image coordinates into 3D space coordinates. The formula is:
[0090]
[0091] where (c x , c y ) is the coordinate of the principal point of the image.
[0092] Coordinate system alignment: Convert the local 3D coordinates to the global coordinate system of the passenger compartment through the camera external parameters (rotation matrix, translation vector).
[0093] 5) Cockpit position matching and target positioning
[0094] Preset cockpit position distribution: Define the spatial range (such as coordinate thresholds or 3D bounding boxes) of each area within the cockpit (such as the driver's seat, co-driver's seat, and rear seats).
[0095] Position matching logic:
[0096] Cluster analysis: Cluster according to the 3D coordinates of the key points to judge the occupant's posture (such as sitting posture, lying posture).
[0097] Spatial mapping: Match the human body center point or hip joint coordinates with the preset area to determine the target cockpit position.
[0098] Dynamic update: Real-time track the coordinate changes to adapt to the scenario of occupant movement.
[0099] Furthermore, as an optional implementation manner, the age group recognition model is trained through the following steps:
[0100] S201. Obtain the time series samples of the human body depth images of multiple testers in the test vehicle, and determine the age group labels corresponding to each time series sample of the human body depth image through manual annotation;
[0101] S202. Input the time series samples of the human body depth images into the pre-constructed CNN-LSTM hybrid neural network to obtain the age group recognition result;
[0102] S203. Determine the loss value according to the age group recognition result and the age group label;
[0103] S204. Update the parameters of the CNN-LSTM hybrid neural network through the backpropagation algorithm according to the loss value to obtain a trained age group recognition model.
[0104] Specifically, collect depth image sequences (such as 30-second video streams) of different testers in the passenger compartment through an in-vehicle binocular camera, covering different age groups, sitting postures, and lighting scenarios; divide the collected samples into at least 3 age groups (such as 0-10 years old, 11-59 years old, and over 60 years old), and have at least 3 annotators label them independently. Use the majority voting method to determine the final label; input the collected depth image sequences into the CNN-LSTM hybrid neural network, extract spatial features through the CNN convolutional layer, capture temporal dimension dependencies through the LSTM layer, and finally map the output of the LSTM layer to a multi-dimensional age group probability distribution through the output layer, and generate an age group recognition result through the SoftMax activation function; calculate the loss value according to the age group recognition result and the age group label using the weighted cross-entropy loss function, use the Adam optimizer (initial learning rate 3e-4), and cooperate with the cosine annealing scheduler to dynamically adjust the learning rate to update the parameters of the CNN-LSTM hybrid neural network. When the preset convergence condition is reached, a trained age group recognition model can be obtained.
[0105] Further as an optional implementation manner, the CNN-LSTM hybrid neural network includes an input layer, a 3D CNN convolutional layer, a 2D CNN convolutional layer, a feature fusion layer, an LSTM layer, an attention layer, and an output layer. The input layer is used to split the data of the human body depth image time series sample to obtain 3D depth image time series data and RGB image time series data. The 3D CNN convolutional layer is used to extract features from the 3D depth image time series data to obtain a depth spatial feature sequence. The 2D CNN convolutional layer is used to extract features from the RGB image time series data to obtain a local area feature sequence. The feature fusion layer is used to fuse the depth spatial feature sequence and the local area feature sequence to obtain a spatio-temporal feature sequence. The LSTM layer is used to generate a hidden state sequence according to the spatio-temporal feature sequence. The attention layer is used to dynamically allocate weights to each dimension of the hidden state sequence based on the multi-head self-attention mechanism. The output layer is used to map the hidden state sequence after dynamic weight allocation to an age group recognition result.
[0106] Specifically, such as Figure 2The following is a schematic diagram of the structure of the CNN-LSTM hybrid neural network provided by the embodiment of the present invention. In the embodiment of the present invention, the time-series samples of human depth images are split into 3D depth image time-series data and RGB image time-series data. The 3D CNN convolutional layer and the 2D CNN convolutional layer are respectively used for feature extraction. Then, the sequences formed by the 3D depth space feature maps and the sequences formed by the local region feature maps extracted are subjected to feature fusion to obtain a spatio-temporal feature sequence. Then, the LSTM layer is used to capture the time-dimensional dependence relationship of the spatio-temporal feature sequence to generate a hidden state sequence. The multi-head self-attention mechanism is used to dynamically allocate weights to each dimension of the hidden state sequence, and then it is mapped to a multi-dimensional age group probability distribution to obtain the age group recognition result.
[0107] Further, as an optional implementation manner, the time-series samples of human depth images are input into a pre-constructed CNN-LSTM hybrid neural network to obtain an age group recognition result, which specifically includes:
[0108] S2021. The time-series samples of human depth images are split through the input layer to obtain 3D depth image time-series data and RGB image time-series data;
[0109] S2022. The 3D depth image time-series data is subjected to feature extraction through the 3D CNN convolutional layer to obtain a depth space feature sequence;
[0110] S2023. The RGB image time-series data is subjected to feature extraction through the 2D CNN convolutional layer to obtain a local region feature sequence;
[0111] S2024. The depth space feature sequence and the local region feature sequence are subjected to spatio-temporal alignment and feature fusion through the feature fusion layer to obtain a spatio-temporal feature sequence;
[0112] S2025. The spatio-temporal feature vectors at each time step are determined according to the spatio-temporal feature sequence. The hidden state vectors corresponding to each spatio-temporal feature vector are calculated through the forget gate, input gate, and output gate of the LSTM layer, and a hidden state sequence is generated according to the hidden state vectors;
[0113] S2026. The multi-head self-attention mechanism is used in the attention layer to dynamically allocate weights to each dimension of the hidden state sequence;
[0114] S2027. The hidden state sequence after dynamic weight allocation is mapped through the output layer to obtain an age group recognition result.
[0115] Specifically, the process of determining the age group recognition result of the time-series samples of human depth images through the CNN-LSTM hybrid neural network is as follows:
[0116] 1) Data splitting
[0117] Data splitting operation: Use programming tools (such as Python combined with the OpenCV library) to split the input time-series samples of human depth images into 3D depth image time-series data and RGB image time-series data. Specifically, extract the part related to depth information from the original data as 3D depth image time-series data, and use the part of color information as RGB image time-series data.
[0118] 2) Feature extraction of 3D depth image time-series data
[0119] Construct a 3D CNN convolutional layer: Use a deep learning framework (such as TensorFlow or PyTorch) to construct a 3D CNN convolutional layer. This layer usually contains multiple 3D convolutional kernels that slide on the 3D depth image time-series data to extract depth space features.
[0120] Feature extraction process: Input the 3D depth image time-series data into the 3D CNN convolutional layer, and perform feature extraction on the data through convolutional operations. The convolutional operation performs weighted summation on the input data to obtain a series of feature maps. After multiple convolutional and pooling operations (such as max pooling), a depth space feature sequence is finally obtained.
[0121] 3) Feature extraction of RGB image time-series data
[0122] Construct a 2D CNN convolutional layer: Similarly, use a deep learning framework to construct a 2D CNN convolutional layer. This layer contains multiple 2D convolutional kernels for feature extraction of RGB image time-series data.
[0123] Feature extraction operation: Input the RGB image time-series data into the 2D CNN convolutional layer, and extract local region features through 2D convolutional operations. The convolutional operation is performed at each time step of the RGB image to obtain local region feature maps for each time step. After multiple convolutional and pooling operations, a local region feature sequence is obtained.
[0124] 4) Feature fusion
[0125] Spatio-temporal alignment: Since there may be differences in the time and space dimensions between the depth space feature sequence and the local region feature sequence, spatio-temporal alignment is required. This can be achieved through interpolation or adjusting the dimensions of the feature sequence.
[0126] Feature fusion operation: Use a feature fusion layer (such as concatenation or weighted summation) to fuse the aligned depth space feature sequence and the local region feature sequence. The fused result is the spatio-temporal feature sequence, which contains the comprehensive features of depth information and RGB information.
[0127] 5) Processing by the LSTM layer
[0128] Determine spatio-temporal feature vectors: Based on the spatio-temporal feature sequence, determine the spatio-temporal feature vectors for each time step. The spatio-temporal feature sequence can be segmented into feature vectors for each time step through simple slicing operations.
[0129] Calculate hidden state vectors: Input the spatio-temporal feature vectors into the LSTM layer. The LSTM layer contains a forget gate, an input gate, and an output gate, and updates the hidden state vectors through the calculations of these gates. The forget gate determines how much information from the previous hidden state vector needs to be forgotten, the input gate determines how much information from the current input needs to be added to the hidden state vector, and the output gate determines how much information from the current hidden state vector needs to be output.
[0130] Generate hidden state sequences: Generate hidden state sequences based on the hidden state vectors for each time step. This sequence contains the feature information of the human body at different time steps.
[0131] 6) Attention layer processing
[0132] Multi-head self-attention mechanism: Use the attention layer to perform dynamic weight allocation for each dimension of the hidden state sequence based on the multi-head self-attention mechanism. The multi-head self-attention mechanism can simultaneously focus on different parts of the hidden state sequence, thereby better capturing important information in the sequence.
[0133] Weight allocation operation: Input the hidden state sequence into the attention layer, and determine the weights for each dimension by calculating attention scores. The attention scores represent the correlation between each dimension and other dimensions. Then, perform weighted summation on the hidden state sequence according to the attention scores to obtain the hidden state sequence after dynamic weight allocation.
[0134] 7) Output the age group recognition result
[0135] Mapping operation: Input the hidden state sequence after dynamic weight allocation into the output layer. The output layer is usually a fully connected layer, which maps the hidden state sequence to the age group recognition result through linear transformation.
[0136] Result output: The output of the output layer can be a probability distribution, indicating the probabilities of the human body belonging to different age groups. The age group with the highest probability can be selected as the final recognition result.
[0137] Through the above steps, the age group recognition result corresponding to the time series sample of the human depth image can be obtained. Then, combine the age group labels to calculate the loss value, and further update the parameters of the CNN-LSTM hybrid neural network. After iterating several times or when the loss value reaches the preset threshold condition, a trained age group recognition model can be obtained. Input the time series data of the human depth image into the pre-trained age group recognition model, and the age group of the target occupant can be recognized.
[0138] As a further optional implementation, the target seat corresponding to the target cockpit position is controlled to be in a locked state, which specifically includes:
[0139] S1031. When the target cockpit position is the left rear seat, determine the left rear seat as the target seat, and send a control instruction to the left domain controller through the central controller, so that the left domain controller controls the left rear seat to be locked;
[0140] S1032. When the target cockpit position is the right rear seat, determine the right rear seat as the target seat, and send a control instruction to the right domain controller through the central controller, so that the right domain controller controls the right rear seat to be locked.
[0141] As a further optional implementation, the intelligent seat adjustment control method further includes the following steps:
[0142] S301. In response to the driver's seat unlocking instruction, unlock the target seat in the locked state;
[0143] S302. In response to the driver's seat locking instruction, lock the target seat in the unlocked state;
[0144] Among them, the seat unlocking instruction and the seat locking instruction are one of a key instruction, a touch instruction, a voice instruction, and a gesture instruction.
[0145] Specifically, the driver can send a seat unlocking instruction for the target seat through methods such as a key instruction, a touch instruction, a voice instruction, and a gesture instruction. At this time, even if the age range of the target occupant on the target seat meets the preset child age range or the elderly age range, the target seat in the locked state will be unlocked to facilitate the driver's subsequent adjustment of the target seat; similarly, the driver can send a seat locking instruction for the target seat through methods such as a key instruction, a touch instruction, a voice instruction, and a gesture instruction. At this time, even if the age range of the target occupant on the target seat does not meet the preset child age range and the elderly age range, the target seat in the unlocked state will be locked. It should be noted that this operation is generally performed after the seat adjustment is completed to prevent the occupant from accidentally adjusting the seat to an uncomfortable state.
[0146] The method steps of the embodiments of the present invention are described above. It can be recognized that the embodiments of the present invention detect the target occupant based on the depth image information of the occupant compartment and identify the age range of the occupant. For the occupant in the child age range or the elderly age range, the seat adjustment function corresponding to the cockpit position is controlled to be in the locked state. For the occupant not in the child age range and the elderly age range, the seat adjustment function corresponding to the cockpit position is controlled to be in the unlocked state, realizing the automatic locking and unlocking of the intelligent seat, avoiding the safety hazards caused by the incorrect operation of the seat adjustment function by child occupants or elderly occupants, and at the same time facilitating other occupants to adjust the seat according to their own needs without the driver manually performing frequent operations of locking and unlocking the seat, improving the driving safety and the user's driving experience.
[0147] Referring to Figure 3 , the embodiments of the present invention provide an intelligent seat adjustment control system based on machine vision, including:
[0148] An image acquisition module, configured to acquire the depth image information of the occupant compartment of the target vehicle, perform human detection on the depth image information to obtain the human depth image of the target occupant, and determine the target cockpit position corresponding to the target occupant;
[0149] An age range recognition module, configured to generate human depth image time series data according to consecutive multiple frames of human depth images, input the human depth image time series data into a pre-trained age range recognition model, and obtain the age range of the target occupant;
[0150] A locking control module, configured to control the target seat corresponding to the target cockpit position to be in the locked state when the age range of the occupant meets the preset child age range interval or elderly age range interval;
[0151] An unlocking control module, configured to control the target seat corresponding to the target cockpit position to be in the unlocked state when the age range of the occupant does not meet the child age range interval and the elderly age range interval.
[0152] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0153] Referring to Figure 4 , the embodiments of the present invention provide an intelligent seat adjustment control device based on machine vision, including:
[0154] At least one processor;
[0155] At least one memory, configured to store at least one program;
[0156] When at least one of the above programs is executed by at least one of the above processors, the at least one processor is caused to implement the above-described intelligent seat adjustment control method based on machine vision.
[0157] The content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0158] The embodiments of the present invention also provide a computer-readable storage medium, in which a program executable by a processor is stored. The program executable by the processor is used to execute the above-described intelligent seat adjustment control method based on machine vision when executed by the processor.
[0159] A computer-readable storage medium of the embodiments of the present invention can execute an intelligent seat adjustment control method based on machine vision provided by the method embodiments of the present invention, can execute any combination of implementation steps of the method embodiments, and has the corresponding functions and beneficial effects of the method.
[0160] The embodiments of the present invention also disclose a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.
[0161] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the above blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.
[0162] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the above-described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0163] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0164] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0165] More specific examples (nonexhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the above programs can be printed, because the above programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or, if necessary, other suitable processing, and then storing them in a computer memory.
[0166] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0167] In the above description of this specification, the descriptions referring to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0168] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
[0169] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.
Claims
1. An intelligent seat adjustment control method based on machine vision, characterized in that: The following steps are involved: Acquire depth image information of a passenger compartment of a target vehicle, perform human body detection on the depth image information to obtain a human body depth image of a target passenger, and determine a target cabin position corresponding to the target passenger; Generating human body depth image time series data according to a plurality of consecutive frames of human body depth images, and inputting the human body depth image time series data into a pre-trained age group recognition model to obtain the occupant age group of the target occupant; When the age group of the passenger meets the preset child age group interval or the elderly age group interval, controlling the target seat corresponding to the target cabin position to be in a locked state; When the age group of the occupant does not conform to the children's age group interval and the elderly's age group interval, the target seat corresponding to the target cabin position is controlled to be in an unlocked state.
2. The intelligent seat adjustment control method based on machine vision according to claim 1 is characterized in that: The step of acquiring depth image information of a passenger compartment of a target vehicle, performing human body detection on the depth image information to obtain a human body depth image of a target passenger, and determining a target cabin position corresponding to the target passenger, specifically includes: Acquiring the depth image information through a binocular camera of the passenger compartment; Performing human key point detection on the depth image information to obtain a plurality of human key points and corresponding key point depth information; Extracting the human body depth image from the depth image information according to the human body key points; Performing 3D spatial mapping on the key points of the human body according to the key point depth information to obtain the human body position coordinates of the target occupant; Preset cockpit position distribution information is obtained, and the target cockpit position corresponding to the target occupant is determined according to the human body position coordinates and the cockpit position distribution information.
3. The intelligent seat adjustment control method based on machine vision according to claim 1 is characterized in that: The age group recognition model is trained by the following steps: Obtaining time-series samples of human body depth images of multiple test persons in a test vehicle, and determining age group labels corresponding to each of the time-series samples of human body depth images through manual annotation; Input the human body depth image time series samples into a pre-built CNN-LSTM hybrid neural network to obtain age group recognition results; determining a loss value according to the age group recognition result and the age group label; The parameters of the CNN-LSTM hybrid neural network are updated according to the loss value through a back propagation algorithm to obtain the trained age group recognition model.
4. The intelligent seat adjustment control method based on machine vision according to claim 3 is characterized in that: The CNN-LSTM hybrid neural network includes an input layer, a 3D CNN convolution layer, a 2D CNN convolution layer, a feature fusion layer, an LSTM layer, an attention layer and an output layer. The input layer is used to perform data splitting on the human body depth image time series samples to obtain 3D depth image time series data and RGB image time series data. The 3D CNN convolution layer is used to perform feature extraction on the 3D depth image time series data to obtain a deep space feature sequence. The 2D CNN convolution layer is used to perform feature extraction on the RGB image time series data to obtain a local area feature sequence. The feature fusion layer is used to perform feature fusion on the deep space feature sequence and the local area feature sequence to obtain a spatiotemporal feature sequence. The LSTM layer is used to generate a hidden state sequence according to the spatiotemporal feature sequence. The attention layer is used to dynamically assign weights to each dimension of the hidden state sequence based on a multi-head self-attention mechanism. The output layer is used to map the hidden state sequence after dynamic weight assignment to the age group recognition result.
5. The intelligent seat adjustment control method based on machine vision according to claim 4 is characterized in that: The method of inputting the human body depth image time series samples into a pre-built CNN-LSTM hybrid neural network to obtain an age group recognition result specifically includes: Performing data splitting on the human body depth image time series samples through the input layer to obtain the 3D depth image time series data and the RGB image time series data; Extracting features from the 3D depth image time series data through the 3D CNN convolutional layer to obtain the depth space feature sequence; Extracting features from the RGB image time series data through the 2D CNN convolutional layer to obtain the local area feature sequence; Performing spatiotemporal alignment and feature fusion on the deep spatial feature sequence and the local area feature sequence through the feature fusion layer to obtain the spatiotemporal feature sequence; Determine the spatiotemporal feature vectors of each time step according to the spatiotemporal feature sequence, calculate the hidden state vectors corresponding to each of the spatiotemporal feature vectors through the forget gate, input gate and output gate of the LSTM layer, and generate the hidden state sequence according to the hidden state vectors; Dynamically weighting the dimensions of the hidden state sequence based on a multi-head self-attention mechanism through the attention layer; The hidden state sequence after dynamic weight allocation is mapped to the age group recognition result through the output layer.
6. The intelligent seat adjustment control method based on machine vision according to claim 1 is characterized in that: The controlling the target seat corresponding to the target cabin position to be in a locked state specifically includes: When the target cabin position is the rear left seat, the rear left seat is determined to be the target seat, and a control instruction is sent to the left domain controller through the central controller, so that the left domain controller controls the rear left seat to be locked; When the target cabin position is the rear right seat, the rear right seat is determined to be the target seat, and a control instruction is sent to the right domain controller through the central controller, so that the right domain controller controls the rear right seat to be locked.
7. An intelligent seat adjustment control method based on machine vision according to any one of claims 1 to 6, characterized in that: The intelligent seat adjustment control method further comprises the following steps: In response to a seat unlocking instruction from a driver, unlocking the target seat in a locked state; In response to a seat locking instruction from a driver, locking the target seat in an unlocked state; Wherein, the seat unlocking instruction and the seat locking instruction are one of a button instruction, a touch instruction, a voice instruction and a gesture instruction.
8. An intelligent seat adjustment control system based on machine vision, characterized in that: include: An image acquisition module is used to obtain depth image information of a passenger compartment of a target vehicle, perform human body detection on the depth image information to obtain a human body depth image of a target passenger, and determine a target cabin position corresponding to the target passenger; An age group recognition module is used to generate human body depth image time series data according to a plurality of consecutive frames of human body depth images, and input the human body depth image time series data into a pre-trained age group recognition model to obtain the occupant age group of the target occupant; A locking control module, configured to control the target seat corresponding to the target cabin position to be in a locked state when the age group of the occupant meets the preset child age group interval or the elderly age group interval; The unlocking control module is used to control the target seat corresponding to the target cabin position to be in an unlocked state when the age group of the occupant does not conform to the children's age group interval and the elderly age group interval.
9. An intelligent seat adjustment control device based on machine vision, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the intelligent seat adjustment control method based on machine vision as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to execute the intelligent seat adjustment control method based on machine vision as described in any one of claims 1 to 7 when executed by the processor.
Citation Information
Cited By
Single-picture multi-target temperature analysis method based on big data acquisition
CN120893125A