Vehicle tailgate control method, device, equipment, storage medium and program product

By adaptively adjusting the feature extraction weights using a visual sensor to detect key human points, the platform compatibility issue of foot-activated sensing technology across different vehicle models has been resolved, enabling cross-vehicle tailgate control and reducing R&D costs and user learning costs.

CN120649762BActive Publication Date: 2025-11-18ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511148911.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-18
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

Existing foot-activated sensor technology has poor platform compatibility across different vehicle models, requiring the separate development of sensor brackets and detection algorithms, which increases R&D costs and time.

Method used

By acquiring target image data through a visual sensor, and adaptively adjusting the feature extraction weights based on sensor parameters to detect key points of the human body, the system can recognize foot kicking behavior and control the vehicle's tailgate.

Benefits of technology

It enables cross-vehicle platform deployment, reduces R&D costs and time, and improves the accuracy of foot kick recognition and the convenience of user operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120649762B_ABST
    Figure CN120649762B_ABST
Patent Text Reader

Abstract

The application provides a vehicle tail door control method, device, equipment, storage medium and program product, relates to the vehicle technical field, and the method comprises the steps that target image data is acquired; the target image data is collected through a visual sensor associated with a vehicle tail door; human body key point data in the target image data is obtained by performing human body key point detection on the target image data based on sensor parameters of the visual sensor; the sensor parameters are used for adaptively adjusting feature extraction weights of human body key point detection on the target image data; an identification result is obtained by performing a kicking behavior recognition based on the human body key point data, and the vehicle tail door is controlled based on the identification result. The application can improve the platform adaptation capability of the technology of kicking sensing control tail door, so that the sensor support and detection algorithm do not need to be developed separately for different vehicle models, and the overall vehicle development cost and period are saved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicles, in particular to a vehicle tailgate control method, device, equipment, storage medium and program product. BACKGROUND

[0002] At present, with the rapid development of automobile intelligence, the automobile industry is evolving towards more convenient and humanized direction. As an important configuration to improve the convenience of vehicle use, the electric tailgate has gradually become a standard function of medium and high-end vehicles.

[0003] In related technologies, users can control the opening and closing of the tailgate through various ways such as remote key, in-vehicle button or tailgate button. These control methods meet the daily use requirements to a certain extent and promote the automation development of automobile tailgate operation. In order to solve the inconvenience of traditional operation method in the scene where the user's hands are occupied (such as carrying goods, holding a baby, etc.), the technology of foot kick sensing control tailgate emerges as the times require. However, since the foot kick sensing generally needs to fix the sensor to be installed at a specific position of the bumper, it is necessary to develop the sensor bracket and detection algorithm separately for different vehicle models. SUMMARY

[0004] The main purpose of the embodiments of the present application is to propose a vehicle tailgate control method, device, equipment, storage medium and program product, which aims to improve the platform adaptation ability of the technology of foot kick sensing control tailgate, so as to save the overall research and development cost and period of the vehicle without developing the sensor bracket and detection algorithm separately for different vehicle models.

[0005] To achieve the above purpose, a first aspect of the embodiments of the present application proposes a vehicle tailgate control method, which comprises:

[0006] obtaining target image data; the target image data is collected by a visual sensor associated with a vehicle tailgate;

[0007] performing human key point detection on the target image data based on sensor parameters of the visual sensor to obtain human key point data in the target image data; the sensor parameters are used to adaptively adjust the feature extraction weight of the human key point detection on the target image data;

[0008] performing foot kick behavior recognition based on the human key point data to obtain a recognition result, and controlling the vehicle tailgate based on the recognition result.

[0009] In some embodiments, the human key point detection on the target image data based on the sensor parameters of the visual sensor comprises:

[0010] The preset initial feature extraction weights are adaptively adjusted based on the sensor parameters of the visual sensor to obtain the adjusted target feature extraction weights; the initial feature extraction weights are the initial feature extraction weights for human key point detection of the target image data.

[0011] Human key point detection is performed on the target image data based on the target feature extraction weights.

[0012] In some embodiments, the human key point detection of the target image data based on the sensor parameters of the visual sensor includes:

[0013] Human key points are detected in the target image data using a preset adaptive detection model, and the feature extraction weights for human key point detection in the target image data are adaptively adjusted based on the sensor parameters using the adaptive detection model.

[0014] In some embodiments, the step of adaptively adjusting the feature extraction weights for human keypoint detection in the target image data based on the sensor parameters using the adaptive detection model includes:

[0015] The adaptive detection model generates a scaling factor and bias term for the convolutional block based on the sensor parameters, and adaptively adjusts the feature extraction weights for human key point detection of the target image data based on the scaling factor and the bias term.

[0016] The convolutional block can be any convolutional block in the multi-layer convolutional blocks of the adaptive detection model.

[0017] In some embodiments, the human body key point data includes planar human body key point data, and the step of obtaining a recognition result by performing kicking behavior recognition based on the human body key point data includes:

[0018] Spatial mapping processing is performed on the planar human body key point data to obtain spatial human body key point data;

[0019] The recognition result is obtained by recognizing kicking behavior based on the aforementioned spatial human body key point data.

[0020] In some embodiments, the spatial mapping processing of the planar human body key point data to obtain spatial human body key point data includes:

[0021] The planar human keypoint data is processed by multi-scale temporal feature fusion using a pre-defined dilated temporal convolution model to obtain spatial human keypoint data output by the dilated temporal convolution model.

[0022] In some embodiments, the process of obtaining a recognition result by recognizing kicking behavior based on the spatial human body key point data includes:

[0023] Extract the spatial foot key point data from the spatial human body key point data;

[0024] Spatiotemporal feature encoding is performed on the spatial foot key point data to obtain a feature vector corresponding to the spatial foot key point data;

[0025] The feature vector is subjected to behavior action decoding processing to obtain the recognition result of kicking behavior.

[0026] In some embodiments, the step of performing behavior action decoding processing on the feature vector to obtain the recognition result of the kicking behavior includes:

[0027] The feature vector is processed by a preset classifier to decode the action, and the predicted probability of the kicking action output by the classifier is obtained.

[0028] If the predicted probability is greater than or equal to a preset probability threshold, the recognition result of the kicking behavior is determined as kicking behavior being recognized;

[0029] If the predicted probability is less than a preset probability threshold, the recognition result of the kicking behavior is determined to be that the kicking behavior was not recognized.

[0030] To achieve the above objectives, a second aspect of this application provides a vehicle tailgate control device, the device comprising:

[0031] A data acquisition module is used to acquire target image data; the target image data is acquired through a vision sensor associated with the vehicle's tailgate.

[0032] An adaptive detection module is used to detect human key points in the target image data based on the sensor parameters of the visual sensor, and obtain human key point data in the target image data; the sensor parameters are used to adaptively adjust the feature extraction weights for human key point detection in the target image data.

[0033] The control module is used to identify kicking behavior based on the human body key point data, obtain the identification result, and control the tailgate of the vehicle based on the identification result.

[0034] To achieve the above objectives, a third aspect of this application provides a vehicle tailgate control device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the vehicle tailgate control method described in the first aspect.

[0035] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the vehicle tailgate control method described in the first aspect.

[0036] To achieve the above objectives, a fifth aspect of the present application provides a computer program product, which includes a computer program that, when executed by a processor, implements the vehicle tailgate control method provided in the first aspect above.

[0037] The vehicle tailgate control method, apparatus, device, computer-readable storage medium, and computer program product proposed in this application acquire target image data; the target image data is acquired by a vision sensor associated with the vehicle tailgate; human key point detection is performed on the target image data based on the sensor parameters of the vision sensor to obtain human key point data in the target image data; the sensor parameters are used to adaptively adjust the feature extraction weights for human key point detection on the target image data; kicking behavior recognition is performed based on the human key point data to obtain a recognition result, and the vehicle tailgate is controlled based on the recognition result.

[0038] Compared to traditional foot-activated tailgate control, this embodiment acquires target image data via a vision sensor associated with the tailgate. Then, based on the sensor parameters of the vision sensor, it performs human key point detection on the target image data. Specifically, it adaptively adjusts the feature extraction weights for human key point detection based on the sensor parameters, thereby obtaining human key point data from the target image data. Finally, it performs foot-kicking behavior recognition processing based on this human key point data to obtain the corresponding recognition result and control the tailgate.

[0039] Thus, this embodiment of the application, based on the adaptive adjustment of the sensor parameters of the vision sensor to perform human key point detection on the target image, can optimize the human key point detection effect of image data acquired by the vision sensor under different installation environments. That is, regardless of the vehicle model or location of the vision sensor, the accuracy of human key point detection on the image data acquired by the vision sensor can be ensured by adaptively adjusting the feature extraction weights. In other words, this embodiment of the application can enable the same human key point detection algorithm to be applied to image data acquired by vision sensors with different parameters, that is, to make the same algorithm adaptable to vision sensors of different vehicle models. This effectively solves the technical bottleneck of cross-vehicle platform deployment of foot-activated tailgate control technology, improves the platform adaptability of foot-activated tailgate control technology, and eliminates the need to develop separate detection algorithms and / or sensor brackets for each vehicle model. By avoiding the need to develop separate sensor brackets and detection algorithms for different vehicle models, the cross-vehicle development cycle can be greatly shortened and R&D costs reduced.

[0040] Furthermore, this application embodiment adaptively adjusts the feature extraction weights when detecting human key points in target images based on the sensor parameters of the visual sensor, thereby optimizing the human key point detection effect of image data collected by the visual sensor under different installation environments. This can further improve the accuracy of subsequent foot kicking action recognition based on human key data, thereby solving the problem of high trigger failure rate caused by non-standard actions in the traditional foot kicking sensor control of vehicle tailgates. This allows different users to successfully trigger tailgate opening under various foot kicking postures, thereby reducing the learning cost of user operation. Attached Figure Description

[0041] Figure 1 A flowchart illustrating the steps of the vehicle tailgate control method provided in some embodiments of this application;

[0042] Figure 2 for Figure 1 A detailed flowchart of step S102;

[0043] Figure 3 for Figure 1 A schematic diagram of another detailed step in step S102;

[0044] Figure 4 for Figure 3 A detailed flowchart of step S301;

[0045] Figure 5 for Figure 1 A detailed flowchart of step S103;

[0046] Figure 6 forFigure 5 A detailed flowchart of step S502;

[0047] Figure 7 for Figure 6 A detailed flowchart of step S603;

[0048] Figure 8 A flowchart illustrating a complete embodiment of the vehicle tailgate control method provided in this application.

[0049] Figure 9 This is a schematic diagram of the structure of the vehicle tailgate control device provided in the embodiments of this application;

[0050] Figure 10 This is a schematic diagram of the hardware structure of the vehicle tailgate control device provided in the embodiments of this application. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0052] It should be noted that although functional modules are divided in the device / system schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device / system or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0054] First, the overall concept of the vehicle tailgate control method provided in the embodiments of this application will be explained.

[0055] With the rapid development of automotive intelligence, the automotive industry is evolving towards greater convenience and user-friendliness. Among these advancements, the power tailgate, as a crucial feature enhancing vehicle usability, has gradually become a standard feature in mid-to-high-end models. Currently, users can control the tailgate's opening and closing through various methods, such as remote keys, in-car buttons, or tailgate buttons. These control methods, to a certain extent, meet daily usage needs and promote the automation of automotive tailgate operations.

[0056] To address the inconvenience of traditional operating methods when users' hands are occupied (such as carrying items or holding a baby), foot-activated sensor technology has emerged. This technology allows users to trigger the automatic opening of the tailgate with a leg movement (such as kicking or sweeping), eliminating the need for manual operation and improving convenience to some extent. However, existing foot-activated sensor technology suffers from a significant drawback: poor platform adaptability. Because the sensor needs to be fixedly installed in a specific position on the bumper, different vehicle models (such as multi-purpose vehicles (MPVs), sport utility vehicles (SUVs), and sedans) require separate development of sensor brackets and detection algorithms, increasing R&D costs and time.

[0057] In summary, while the existing foot-activated tailgate technology for electric vehicles has solved the operation problem to some extent in scenarios where both hands are occupied, it still suffers from poor platform compatibility, making it difficult to meet the usage needs of different user groups and the adaptation requirements of different vehicle models. Further improvement and optimization are urgently needed.

[0058] To address this, embodiments of this application provide a vehicle tailgate control method, device, equipment, computer-readable storage medium, and computer program product, aiming to overcome the shortcomings of the aforementioned related technologies, improve the platform adaptability of the foot-activated tailgate control technology, thereby eliminating the need to develop separate sensor brackets and detection algorithms for different vehicle models, and saving overall vehicle R&D costs and time.

[0059] The vehicle tailgate control method, apparatus, device, computer-readable storage medium, and computer program product proposed in this application acquire target image data; the target image data is acquired by a vision sensor associated with the vehicle tailgate; human key point detection is performed on the target image data based on the sensor parameters of the vision sensor to obtain human key point data in the target image data; the sensor parameters are used to adaptively adjust the feature extraction weights for human key point detection on the target image data; kicking behavior recognition is performed based on the human key point data to obtain a recognition result, and the vehicle tailgate is controlled based on the recognition result.

[0060] Compared to traditional foot-activated tailgate control, this embodiment acquires target image data via a vision sensor associated with the tailgate. Then, based on the sensor parameters of the vision sensor, it performs human key point detection on the target image data. Specifically, it adaptively adjusts the feature extraction weights for human key point detection based on the sensor parameters, thereby obtaining human key point data from the target image data. Finally, it performs foot-kicking behavior recognition processing based on this human key point data to obtain the corresponding recognition result and control the tailgate.

[0061] Thus, this embodiment of the application, based on the adaptive adjustment of the sensor parameters of the vision sensor to perform human key point detection on the target image, can optimize the human key point detection effect of image data acquired by the vision sensor under different installation environments. That is, regardless of the vehicle model or location of the vision sensor, the accuracy of human key point detection on the image data acquired by the vision sensor can be ensured by adaptively adjusting the feature extraction weights. In other words, this embodiment of the application can enable the same human key point detection algorithm to be applied to image data acquired by vision sensors with different parameters, that is, to make the same algorithm adaptable to vision sensors of different vehicle models. This effectively solves the technical bottleneck of cross-vehicle platform deployment of foot-activated tailgate control technology, improves the platform adaptability of foot-activated tailgate control technology, and eliminates the need to develop separate detection algorithms and / or sensor brackets for each vehicle model. By avoiding the need to develop separate sensor brackets and detection algorithms for different vehicle models, the cross-vehicle development cycle can be greatly shortened and R&D costs reduced.

[0062] Furthermore, this application embodiment adaptively adjusts the feature extraction weights when detecting human key points in target images based on the sensor parameters of the visual sensor, thereby optimizing the human key point detection effect of image data collected by the visual sensor under different installation environments. This can further improve the accuracy of subsequent foot kicking action recognition based on human key data, thereby solving the problem of high trigger failure rate caused by non-standard actions in the traditional foot kicking sensor control of vehicle tailgates. This allows different users to successfully trigger tailgate opening under various foot kicking postures, thereby reducing the learning cost of user operation.

[0063] Furthermore, in recognizing kicking behavior based on key human body data, this embodiment of the application can also utilize spatial perception technology to map the key human body points detected in the 2D plane to 3D space, constructing a 3D human body key point sequence with temporal information. Subsequently, based on spatiotemporal feature encoding and a behavior action decoding module, the human body actions in 3D space are analyzed and recognized to accurately determine the kicking behavior. This solves the problem of high failure rates for non-standard actions (such as small-amplitude movements or actions of special groups) caused by the lack of 3D spatial feature analysis capabilities in traditional kick-sensor control of vehicle tailgates. It improves the adaptability and robustness of action recognition to different user groups and different action forms, thus further reducing the learning cost for users.

[0064] Next, the vehicle tailgate control method, device, equipment, computer-readable storage medium, and computer program product provided in this application will be specifically described through the following embodiments, and the vehicle tailgate control method provided in this application will be described in detail first.

[0065] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0066] It should be noted that the vehicle tailgate control method provided in this application embodiment can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a computer device such as a smartphone, tablet, laptop, or desktop computer, or it can be an in-vehicle terminal (e.g., an in-vehicle computing platform) or a terminal associated with the vehicle. Associating with the vehicle means that the terminal can communicate and interact with the vehicle via a network. The server can be a backend server terminal device, which can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The software can be an application implementing the vehicle tailgate control method, a computer program, and a storage medium carrying the computer program. It should be understood that, based on different design needs of practical applications, the terminal, server and software of the vehicle tailgate control method provided in this application may also be other forms not listed here, and the vehicle tailgate control method provided in this application does not specifically limit these.

[0067] Furthermore, this application can also be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, personal computers (PCs), minicomputers, mainframe computers, vehicles, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0068] For ease of understanding and explanation, the following text will use the vehicle tailgate control method provided in the embodiments of this application applied to a terminal device as an example to describe the various specific embodiments of this application in detail. The implementation of the vehicle tailgate control method provided in the embodiments of this application by any other subject as described above can refer to the process of applying the vehicle tailgate control method to a terminal device as described below.

[0069] Please refer to Figure 1 , Figure 1 The flowchart illustrates the steps of the vehicle tailgate control method provided in some embodiments of this application. It should be understood that, although... Figure 1 The flowcharts illustrating subsequent steps show the execution order of some method steps. However, based on different design needs in practical applications, the vehicle tailgate control method provided in this application embodiment can, of course, employ a different execution order of method steps than shown in the figures. That is, Figure 1 The order of the method steps shown does not constitute a limitation on the execution logic order of the vehicle tailgate control method provided in the embodiments of this application. Any other logic based on... Figure 1 Reasonable changes to the sequence of steps shown should be included within the protection scope of the vehicle tailgate control method provided in the embodiments of this application.

[0070] like Figure 1 As shown, in some embodiments, the vehicle tailgate control method provided in this application may include steps S101 to S103 as shown below.

[0071] Step S101: Acquire target image data; the target image data is acquired by a vision sensor associated with the vehicle tailgate.

[0072] It should be noted that the visual sensor associated with the vehicle tailgate can be a camera mounted on the tailgate. The terminal device establishes a communication connection with the camera; for example, the terminal device can utilize the vehicle's internal Ethernet network to establish a communication connection with the camera, thereby transmitting data via Ethernet cable. Alternatively, the visual sensor associated with the tailgate can also be a camera mounted in other locations on the vehicle. This camera establishes a wired or wireless communication connection with the terminal device and can capture images of the environment at the tailgate, inputting the captured image data to the terminal device. It should be understood that, based on different design requirements for practical applications, the visual sensor can be installed in different locations on the vehicle, or even externally. As long as the visual sensor can capture images of the environment at the tailgate and input the captured image data to the terminal device, it should be included within the scope of protection of this application.

[0073] When a user controls the tailgate by kicking it, a vision sensor associated with the vehicle's tailgate captures real-time image data of the user's kicking action and uploads the captured target image data to the terminal device. In this way, the terminal device can obtain the target image data input from the vision sensor.

[0074] In some embodiments, the terminal device can continuously receive environmental image data collected in real time by a vision sensor, and then filter out target image data from the environmental image data that contains image elements of the user performing a kicking action to control the tailgate switch.

[0075] In some embodiments, the terminal device can continuously receive video data acquired in real time by the vision sensor, and preprocess the video data, including image noise reduction to eliminate the impact of environmental noise on image quality, and then extract target image data containing image elements of the user performing a kicking action to control the tailgate switch from the preprocessed video data.

[0076] Step S102: Based on the sensor parameters of the visual sensor, perform human key point detection on the target image data to obtain human key point data in the target image data; the sensor parameters are used to adaptively adjust the feature extraction weights for human key point detection on the target image data.

[0077] After acquiring the target image data input from the vision sensor, the terminal device further processes the target image data for human keypoint detection. Furthermore, during this human keypoint detection process, the feature extraction weights for human keypoint detection are adaptively adjusted based on the sensor parameters of the vision sensor. Thus, after completing the human keypoint detection of the target image data by adaptively adjusting the feature extraction weights, the terminal device can obtain the user's human keypoint data when the user performs a kicking motion to control the tailgate opening and closing.

[0078] In some embodiments, the terminal device adaptively adjusts the feature extraction weights for human key point detection of the target image data. The human key point data obtained from human key point detection of the target image data may include, but is not limited to, the coordinates of 17 key points such as the user's ankle, knee, and hip joint in the image.

[0079] In some embodiments, when the terminal device adaptively adjusts the feature extraction weights for human key point detection in target image data based on the sensor parameters of the vision sensor, the sensor parameters used may include, but are not limited to, the installation height, viewing angle, intrinsic parameter matrix, distortion coefficient, and other parameters of the vision sensor.

[0080] In some embodiments, the terminal device can acquire the target image data input from the vision sensor while simultaneously performing parameter identification on the vision sensor to obtain its sensor parameters. For example, if the vision sensor is a camera mounted on the tailgate of a vehicle, the terminal device can obtain installation parameters through camera calibration, including: the installation height H of the rearview camera in the trunk, and the viewing angle. Camera intrinsic parameter matrix K, distortion coefficient k.

[0081] In other embodiments, before acquiring target image data from the visual sensor to perform kicking behavior recognition to control the vehicle tailgate (e.g., when the visual sensor first establishes a communication connection, and / or when the user triggers the calibration of the visual sensor installation parameters), the terminal device may first perform parameter recognition on the visual sensor to obtain the sensor parameters of the visual sensor, and store the sensor parameters locally so that when target image data is acquired later, the sensor parameters can be called from the local storage to adaptively adjust the feature extraction weights when performing human key point data on the target image data.

[0082] Step S103: Based on the human body key point data, perform kicking behavior recognition to obtain recognition results, and control the vehicle tailgate based on the recognition results.

[0083] After obtaining human key point data by detecting human key points in the target image data, the terminal device performs kicking behavior recognition based on this data to obtain the corresponding recognition result. Then, based on the kicking behavior recognition result, the terminal device performs corresponding control on the vehicle's tailgate, such as controlling the tailgate to open or close.

[0084] In some embodiments, when a terminal device processes kicking behavior recognition based on human body key point data, it can extract the coordinates of key points such as the user's ankle, knee, and hip joint from the human body key point data, and analyze the spatial position changes, movement trajectory, and temporal features of these coordinates to determine whether kicking behavior has been recognized.

[0085] In this embodiment, target image data input from a visual sensor associated with the vehicle's tailgate is acquired via a terminal device. Human key point detection is then performed on this target image data. During this detection process, the feature extraction weights are adaptively adjusted based on the sensor parameters of the visual sensor to obtain the user's human key point data when the user performs a kicking motion to control the tailgate opening and closing. Finally, the terminal device performs kicking behavior recognition based on the human key point data to obtain the corresponding recognition result. Then, the vehicle's tailgate is controlled accordingly based on the kicking behavior recognition result.

[0086] Compared to traditional foot-activated tailgate control, this embodiment acquires target image data via a vision sensor associated with the tailgate. Then, based on the sensor parameters of the vision sensor, it performs human key point detection on the target image data. Specifically, it adaptively adjusts the feature extraction weights for human key point detection based on the sensor parameters, thereby obtaining human key point data from the target image data. Finally, it performs foot-kicking behavior recognition processing based on this human key point data to obtain the corresponding recognition result and control the tailgate.

[0087] Thus, this embodiment of the application, based on the adaptive adjustment of the sensor parameters of the vision sensor to perform human key point detection on the target image, can optimize the human key point detection effect of image data acquired by the vision sensor under different installation environments. That is, regardless of the vehicle model or location of the vision sensor, the accuracy of human key point detection on the image data acquired by the vision sensor can be ensured by adaptively adjusting the feature extraction weights. In other words, this embodiment of the application can enable the same human key point detection algorithm to be applied to image data acquired by vision sensors with different parameters, that is, to make the same algorithm adaptable to vision sensors of different vehicle models. This effectively solves the technical bottleneck of cross-vehicle platform deployment of foot-activated tailgate control technology, improves the platform adaptability of foot-activated tailgate control technology, and eliminates the need to develop separate detection algorithms and / or sensor brackets for each vehicle model. By avoiding the need to develop separate sensor brackets and detection algorithms for different vehicle models, the cross-vehicle development cycle can be greatly shortened and R&D costs reduced.

[0088] Furthermore, this application embodiment adaptively adjusts the feature extraction weights when detecting human key points in target images based on the sensor parameters of the visual sensor, thereby optimizing the human key point detection effect of image data collected by the visual sensor under different installation environments. This can further improve the accuracy of subsequent foot kicking action recognition based on human key data, thereby solving the problem of high trigger failure rate caused by non-standard actions in the traditional foot kicking sensor control of vehicle tailgates. This allows different users to successfully trigger tailgate opening under various foot kicking postures, thereby reducing the learning cost of user operation.

[0089] In some embodiments, when a terminal device performs human key point detection on a target image, it can first adaptively adjust the feature extraction weights for human key point detection on the target image data based on sensor parameters, and then use the adjusted feature extraction weights to perform human key point detection on the target image data, thereby obtaining human key point data.

[0090] Please refer to Figure 2 , Figure 2 for Figure 1 A detailed flowchart of step S102.

[0091] like Figure 2 As shown, in some embodiments, the step of "detecting human key points in the target image data based on the sensor parameters of the vision sensor" in step S102 above may include steps S201 and S202 as shown below.

[0092] Step S201: Adaptively adjust the preset initial feature extraction weights based on the sensor parameters of the visual sensor to obtain the adjusted target feature extraction weights; the initial feature extraction weights are the initial feature extraction weights for human key point detection of the target image data.

[0093] Before performing human keypoint detection on the target image data input by the vision sensor, the terminal device can adaptively adjust the initial feature extraction weights for human keypoint detection on the target image data based on the sensor parameters of the vision sensor, thus obtaining the adjusted target feature extraction weights.

[0094] In some embodiments, the terminal device can first read the initial feature extraction weights for human keypoint detection in the target image data, which are pre-set (e.g., determined during training) by the detection algorithm for human keypoint detection in the target image data. Then, based on the sensor parameters of the visual sensor, the terminal device adaptively adjusts some or all of these initial feature extraction weights, for example, increasing the weight of extracting a particular keypoint feature and / or decreasing the weight of extracting a particular keypoint feature, thereby obtaining the adjusted target feature extraction weights.

[0095] Step S202: Perform human key point detection on the target image data based on the target feature extraction weights.

[0096] After the terminal device obtains the adjusted target feature extraction weights, it can perform human key point detection on the target image data based on these target feature extraction weights, thereby obtaining the human key point data in the target image data.

[0097] In some embodiments, the terminal device can also adaptively adjust the initial feature extraction weights based on the sensor parameters of the visual sensor while performing human key point detection on the target image data, thereby obtaining the adjusted target feature extraction weights, and then performing human key point detection on the target image data based on the target feature extraction weights to obtain human key point data in the target image data.

[0098] In some embodiments, the terminal device can perform human key point detection on target image data using an adaptive detection model. In this case, the terminal device can incorporate the sensor parameters of the visual sensor into the model input, thereby allowing the model to dynamically and adaptively adjust the feature extraction weights for human key point detection on the target image data based on the sensor parameters.

[0099] It should be noted that the adaptive detection model can be a 2D keypoint detection model based on multi-layer convolutional blocks. Examples include single-stage models based on heatmaps such as Hourglass Network and SimpleBaseline, regression-based models such as DeepPose and Integral Human Pose Regression, two-stage models such as Mask R-CNN and AlphaPose, and lightweight models such as MobileNet and PP-TinyPose. It should be understood that, based on different design needs in practical applications, different terminal devices may use different models to detect human keypoints in target image data. The vehicle tailgate control method provided in this application does not limit the specific type of adaptive detection model.

[0100] Please refer to Figure 3 , Figure 3 for Figure 1 A schematic diagram of another detailed step in step S102.

[0101] like Figure 3 As shown, in some embodiments, the step of "detecting human key points in the target image data based on the sensor parameters of the visual sensor" in step S102 above may also include step S301 as shown below.

[0102] Step S301: Perform human key point detection on the target image data using a preset adaptive detection model, and adaptively adjust the feature extraction weights for human key point detection on the target image data based on the sensor parameters using the adaptive detection model.

[0103] When a terminal device performs human keypoint detection on target image data using an adaptive detection model, it uses the target image data and the sensor parameters of the vision sensor as input to the adaptive detection model. This allows the adaptive detection model to adaptively adjust the feature extraction weights for human keypoint detection on the target image data based on the sensor parameters.

[0104] In some embodiments, the terminal device can also pre-input the sensor parameters of the visual sensor into the adaptive detection model. This adaptive detection model first adaptively adjusts the feature extraction weights for human keypoint detection on the target image data based on the sensor parameters, resulting in a new detection model with adjusted weights. Then, when the terminal device performs human keypoint detection on the target image data input from the visual sensor, it can use that target image data as the model input for the new detection model, thereby enabling the new detection model to perform human keypoint detection on the target image data.

[0105] Please refer to Figure 4 , Figure 4 for Figure 3 A detailed flowchart of step S301.

[0106] like Figure 4 As shown, in some embodiments, the step of "adaptively adjusting the feature extraction weights for human key point detection of the target image data based on the sensor parameters by means of the adaptive detection model" in step S301 above may include, but is not limited to, step S401 as shown below.

[0107] Step S401: The adaptive detection model generates a scaling factor and bias term for the convolutional block based on the sensor parameters, so as to adaptively adjust the feature extraction weights for human key point detection of the target image data based on the scaling factor and the bias term.

[0108] It should be noted that a convolutional block can be any one or more convolutional blocks in the multi-layer convolutional blocks that make up the adaptive detection model.

[0109] When a terminal device adaptively adjusts the feature extraction weights for human keypoint detection in target image data based on sensor parameters using an adaptive detection model, it can incorporate these sensor parameters into the model input. This allows the adaptive detection model to dynamically adjust the scaling factor γ and bias β of the convolutional block based on these sensor parameters, thereby achieving adaptive adjustment of the feature extraction weights.

[0110] For example, when a visual sensor (such as a camera) is mounted at a high height on the tailgate of a vehicle (e.g., in an SUV), its field of view is wide and the proportion of human figures in the image is relatively small. Therefore, the terminal device can incorporate the camera's mounting height parameter into the model input of an adaptive detection model. This adaptive detection model then adjusts the scaling factor γ of the convolutional blocks based on the input mounting height parameter, thereby increasing the weight for extracting features of small human keypoints and enhancing its feature extraction capabilities. Furthermore, for cameras with different distortion coefficients, the terminal device can incorporate the camera's distortion coefficient into the model input of the adaptive detection model. This adaptive detection model then adjusts the bias β based on the input distortion coefficient, thereby increasing the weight for resisting lens distortion in extracting human keypoint features and correcting keypoint position deviations caused by lens distortion.

[0111] In this embodiment, the sensor parameters of the vision sensor (such as installation height, viewing angle, intrinsic parameter matrix, distortion coefficient, etc.) are used as input conditions for the model of human key point detection of target image by the terminal device. The scaling factor γ and bias β of the convolution block are dynamically generated inside the model, so that the feature extraction weights during human key point detection can be adaptively adjusted according to the sensor parameters. This enables the same detection algorithm to be adapted to vision sensors with different installation heights and viewing angles, thus eliminating the need to develop separate algorithms for each vehicle model and solving the technical bottleneck of cross-vehicle platform deployment.

[0112] Furthermore, this embodiment optimizes the key point detection performance of image data under different installation environments by dynamically adjusting the feature extraction weights. That is, regardless of the vehicle model or location of the visual sensor, the adaptive detection model can adaptively adjust its internal parameters to ensure the accuracy of human key point detection, providing reliable basic data for subsequent kicking action recognition and further improving the overall stability and adaptability of the kick-sensor-controlled tailgate.

[0113] Furthermore, this embodiment, based on a sensor parameter adaptive model design, enhances the algorithm's versatility. This allows for easier integration of foot-activated tailgate control technology into different vehicle production lines during the automotive manufacturing process, reducing the workload of technical adaptation due to model differences. This, in turn, facilitates the rapid rollout of this intelligent configuration by automotive companies, enhancing the product's intelligent competitiveness. Moreover, because the adaptive detection model can adaptively adjust its internal calculation parameters (i.e., feature extraction weights) based on the hardware characteristics of different visual sensors (one of the sensor parameters integrated into the model input, such as resolution, frame rate, distortion, etc.), it can achieve efficient feature extraction and action recognition on automotive platforms with different hardware configurations. This avoids performance fluctuations caused by hardware differences, ensuring stable operation. In other words, this embodiment also improves the adaptability of hardware resources.

[0114] In some embodiments, the human keypoint data obtained by the terminal device from the target image data through human keypoint detection can be 2D planar human keypoint data. The terminal device can employ a spatial intelligent mapping mechanism, using spatial perception technology to map the human keypoints detected in the 2D plane to 3D space, constructing a 3D human keypoint sequence with temporal information, and then analyzing and recognizing human actions in the 3D space to accurately determine kicking behavior.

[0115] Please refer to Figure 5 , Figure 5 for Figure 1 A detailed flowchart of step S103.

[0116] like Figure 5As shown, in some embodiments, the step of "obtaining the recognition result by performing kicking behavior recognition based on the human body key point data" in step S103 above may include steps S501 and S502 as shown below.

[0117] Step S501: Perform spatial mapping processing on the planar human body key point data to obtain spatial human body key point data.

[0118] When the terminal device performs kicking behavior recognition on human body key point data, it can first perform 3D spatial mapping processing on the 2D planar human body key point data to obtain 3D spatial human body key point data.

[0119] In some embodiments, the terminal device can use RNN to model temporal human key point information to map 2D planar human key point data to 3D space, thereby obtaining 3D spatial human key point data.

[0120] In other embodiments, to further improve efficiency and accuracy, the terminal device can also use a dilated temporal convolution model to process 2D keypoint sequences for spatial mapping. Here, the dilated temporal convolution model can be... The system consists of layers of convolutional blocks, and the core design of each layer is as follows:

[0121] .

[0122] in, Indicates the current layer At time step The output; It is an activation function used to increase the nonlinear representation capability of feature vectors; For batch normalization; The kernel size is [size]. Indicates the position of the convolution kernel The weights; This represents the input of the previous layer, with index 1. , Expansion factor; This is a bias term.

[0123] Based on this, step S501 above: performing spatial mapping processing on the planar human body key point data to obtain spatial human body key point data may include the following steps:

[0124] The planar human keypoint data is processed by multi-scale temporal feature fusion using a pre-defined dilated temporal convolution model to obtain spatial human keypoint data output by the dilated temporal convolution model.

[0125] When a terminal device uses a dilated temporal convolutional model to process 2D keypoint sequences for spatial mapping, it inputs 2D planar human keypoint data into the model. The model then performs multi-scale temporal feature fusion processing on this planar human keypoint data, mapping the 2D keypoint sequence into a 3D coordinate sequence, thus obtaining and outputting 3D spatial human keypoint data. In this way, the terminal device can acquire this spatial human keypoint data.

[0126] For example, suppose the 2D keypoint coordinate sequence in the planar human body keypoint data is as follows: ,in, express time Key points of the human body The 2D coordinates are obtained by the terminal device performing multi-scale temporal feature fusion processing on the 2D keypoint coordinate sequence using a dilated temporal convolution model, resulting in a 3D coordinate sequence. .

[0127] Step S502: Based on the spatial human body key point data, perform kicking behavior recognition to obtain the recognition result.

[0128] After mapping 2D planar human body key point data into 3D spatial human body key point data, the terminal device can perform kicking behavior recognition based on this spatial human body key point data, thereby obtaining the corresponding recognition result. In this way, the terminal device can control the vehicle's tailgate accordingly based on the recognition result.

[0129] In some embodiments, the terminal device can extract the coordinates of key points such as the user's ankle, knee, and hip joints from 3D spatial human key point data, and analyze these coordinates for spatial position changes, movement trajectories, and temporal features to determine whether kicking behavior has been identified.

[0130] In this embodiment, 2D planar human keypoint data is mapped to 3D space through the terminal device, resulting in 3D spatial human keypoint data. Specifically, by converting the sequence of 2D human keypoints with temporal information (such as the coordinates of 17 key points like the ankle, knee, and hip joints in the image) to 3D spatial coordinates, the true spatial form of the kicking motion can be reconstructed and the trend of motion changes can be analyzed. Thus, compared to traditional vision-based kicking recognition methods (which typically use 2D images as input, but 2D images only present the projection of the motion, losing crucial information such as depth and spatial trajectory), this embodiment effectively captures the spatial motion state of the kicking motion by mapping 2D keypoints to 3D space, thereby improving the robustness of motion recognition. For example, due to physical limitations, the kicking amplitude of a pregnant woman may not be obvious in a 2D image, but through 3D spatial mapping, the terminal device can accurately capture the three-dimensional features of her foot's motion trajectory and displacement distance in space. This approach not only solves the problem of difficulty in recognizing kicking motions caused by special groups or objects partially obscuring the human body, but also mines the three-dimensional features of human movements through 3D spatial mapping. This enables more accurate recognition of small-amplitude movements, as well as unconventional movements such as side kicks and multi-angle kicks, thereby significantly reducing the failure rate caused by non-standard movements and covering a wider range of user groups and action scenarios.

[0131] Furthermore, although this embodiment introduces 3D spatial mapping in the foot motion recognition stage, it can still efficiently implement the overall foot-activated tailgate control algorithm through a reasonable model architecture design that dynamically adjusts convolutional block parameters to adaptively adjust feature extraction weights. Compared to computationally complex DTW algorithms and other related technologies, this embodiment optimizes the computation process, reduces computational complexity, and improves the real-time performance of foot-activated motion recognition while ensuring recognition accuracy. This enables rapid detection of foot-activated motions and timely response to control the tailgate, enhancing the user's smoothness and immediacy when using foot-activated tailgate control.

[0132] In some embodiments, the terminal device may also employ a kick recognition mechanism based on spatiotemporal feature decoding. Based on the spatiotemporal feature encoding and the behavior action decoding module, it analyzes and identifies human movements in 3D space, thereby achieving the purpose of accurately judging kicking behavior.

[0133] Please refer to Figure 6 , Figure 6 for Figure 5 A detailed flowchart of step S502.

[0134] like Figure 6As shown, in some embodiments, step S502 above: obtaining the recognition result by performing kicking behavior recognition based on the spatial human body key point data may include steps S601 to S603 as shown below.

[0135] Step S601: Extract the spatial foot key point data from the spatial human body key point data.

[0136] It should be noted that spatial foot keypoint data can be a sequence of 3D keypoints related only to foot movements from spatial human body keypoint data.

[0137] When terminal devices perform kicking behavior recognition on spatial human key point data based on spatiotemporal feature decoding, in order to further improve efficiency or ensure that the overall efficiency is not affected, spatial constraint processing can be performed on the spatial human key point data first, thereby extracting spatial foot key point data from the spatial human key point data.

[0138] For example, the terminal device can filter out irrelevant objects by setting a detection area (such as a 1.5m x 1m cube space under the rear of a vehicle) for spatial human key point data, thereby retaining only the 3D key point sequence (spatial foot key point data) related to foot movements in the spatial human key point data.

[0139] Step S602: Perform spatiotemporal feature encoding processing on the spatial foot key point data to obtain feature vectors corresponding to the spatial foot key point data.

[0140] After the terminal device extracts the spatial foot key point data from the spatial human body key point data, it can perform spatiotemporal feature encoding processing on the spatial foot key point data to obtain the feature vector corresponding to the spatial foot key point data.

[0141] In some embodiments, the terminal device may use a Long Short-Term Memory (LSTM) network to perform spatiotemporal feature encoding on the 3D keypoint sequence in the spatial foot keypoint data, thereby generating a feature vector corresponding to the spatial foot keypoint data.

[0142] For example, assume that the spatial foot keypoint data is The terminal device uses a Long Short-Term Memory (LSTM) network to encode the 3D keypoint sequence, which generates a feature vector corresponding to the foot keypoint data in that space. .

[0143] Step S603: Perform behavioral action decoding processing on the feature vector to obtain the recognition result of kicking behavior.

[0144] After performing spatiotemporal feature encoding on the spatial foot key point data, the terminal device further performs behavioral action decoding on the encoded feature vector to determine whether the user's foot action corresponding to the spatial foot key point data is a kicking action, and uses the determination result as the recognition result of the current kicking behavior recognition.

[0145] Please refer to Figure 7 , Figure 7 for Figure 6 A detailed flowchart of step S603.

[0146] like Figure 7 As shown, in some embodiments, step S603 above: performing behavior action decoding processing on the feature vector to obtain the recognition result of kicking behavior may include steps S701 to S703 as shown below.

[0147] Step S701: The feature vector is processed by a preset classifier to decode the action, and the predicted probability of the kicking action output by the classifier is obtained.

[0148] It should be noted that the classifier can be a machine learning classifier, a deep learning classifier, or an ensemble and hybrid classifier. It should be understood that, based on different design requirements of practical applications, different classifiers can be used by the terminal device to decode the behavior actions of the feature vectors in different implementations. The vehicle tailgate control method provided in this application does not limit the type of classifier.

[0149] The feature vector obtained by the terminal device through spatiotemporal feature encoding can be decoded by a classifier to determine the kicking behavior and obtain the predicted probability of the kicking behavior output by the classifier.

[0150] For example, a terminal device can use a fully connected neural network to process feature vectors. Action decoding is performed to determine whether it is a kicking motion. The decision function of this fully connected neural network can be:

[0151] .

[0152] Where p is the predicted probability. This is the Sigmoid function.

[0153] Step S702: If the predicted probability is greater than or equal to a preset probability threshold, determine that the kicking behavior has been identified.

[0154] It should be noted that the preset probability threshold can be a classifier prediction probability threshold pre-configured for the terminal device to determine whether the classifier decodes the kicking behavior. It should be understood that, based on different design needs in practical applications, different threshold values ​​can be used in different implementations of the terminal device. The vehicle tailgate control method provided in this application embodiment does not limit the specific size of the preset probability threshold.

[0155] After obtaining the predicted probability of the kicking behavior output by the classifier, the terminal device compares the preset probability with a preset probability threshold. If the comparison finds that the predicted probability is greater than or equal to the preset probability threshold, the terminal device determines that the classifier has decoded the kicking behavior, and thus determines that the kicking behavior has been identified.

[0156] Step S703: If the predicted probability is less than a preset probability threshold, determine that the recognition result of the kicking behavior is that the kicking behavior was not recognized.

[0157] After comparing the predicted probability of the kicking behavior output by the classifier with a preset probability threshold, if the comparison finds that the predicted probability is less than the preset probability threshold, the terminal device determines that the classifier has not decoded the kicking behavior, and thus determines that the kicking behavior recognition result is that the kicking behavior has not been recognized.

[0158] For example, the terminal device uses a decision function through a classifier (fully connected neural network). After decoding the feature vector V to obtain the prediction probability p, if the prediction probability... Preset probability threshold In this case, the terminal device determines that the user's foot movement corresponding to the aforementioned spatial foot key point data is a kicking motion, and if the predicted probability... <Preset probability threshold The terminal device then determines that the user's foot movement corresponding to the key foot points in the space is not a kicking motion.

[0159] In this embodiment, the terminal device analyzes and identifies human movements using 3D spatial human keypoint data based on spatiotemporal feature encoding and a behavior decoding module. This approach simultaneously considers the spatial features (position and posture of 3D keypoints) and temporal features (the changing patterns of the keypoint sequence over time) of the movement through spatiotemporal feature encoding and the behavior decoding module. Therefore, compared to related technologies that rely solely on the optical flow and shape features of two-dimensional images for kicking behavior recognition, this embodiment can more comprehensively and deeply analyze the essence of kicking movements. For example, for non-standard movements (such as small kicking angles and slow speeds), spatiotemporal feature encoding allows this embodiment to extract features such as the continuity of the movement in the temporal dimension and the motion trend in the spatial dimension. After behavior decoding, these features are accurately identified as kicking movements, thereby reducing the user's learning cost and improving the robustness of kicking behavior recognition.

[0160] Furthermore, this embodiment recognizes kicking behavior by using spatiotemporal feature encoding and behavioral action decoding. This fully utilizes the temporal sequence information and spatial morphological information of the action, enabling stable and accurate recognition of kicking actions even in complex scenes with occlusion or complex backgrounds. For example, in a scenario where a handheld object obscures the leg, this embodiment can identify the kicking action by using the action features and spatiotemporal correlation information of the unobstructed parts in 3D space. In scenarios with multiple people acting simultaneously, this embodiment can distinguish the target human body's actions based on 3D spatial location and spatiotemporal sequence.

[0161] Next, a complete embodiment of the vehicle tailgate control method provided in this application will be presented.

[0162] Please refer to Figure 8 , Figure 8 A schematic flowchart of a complete embodiment of the vehicle tailgate control method provided in this application.

[0163] like Figure 8 As shown, when the terminal device applies the vehicle tailgate control method provided in this application embodiment, it can first preprocess the video data (target image data) input from the camera (visual sensor), such as performing image noise reduction on the video data to eliminate the impact of environmental noise on image quality. Simultaneously, the terminal device also performs camera parameter recognition processing, that is, obtaining installation parameters through camera calibration, including the installation height H of the rearview camera in the trunk and the viewing angle. Camera intrinsic parameter matrix K, distortion coefficient k.

[0164] The terminal device then performs kicking behavior detection on the pre-processed video data.

[0165] In this process, the terminal device first performs 2D keypoint detection and tracking processing based on camera parameters; that is, the terminal device first converts the camera parameter vector... and the preprocessed image The input is fed into a 2D keypoint detection model (adaptive detection model), which then generates the scaling factor for the convolutional blocks through an adaptive module. and bias terms Based on this model, 17 key points of the human body (such as ankles, knees, hip joints, etc.) were detected, and 2D planar human key point data was obtained by establishing temporal correlations of key points across frames through an ID tracking algorithm.

[0166] .

[0167] After detecting kicking behavior and obtaining 2D planar human keypoint data, the terminal device further performs 3D spatial mapping and region of interest constraint on the keypoints. Specifically, the terminal device performs 3D spatial mapping on the planar human keypoint data based on a dilated temporal convolution model to obtain a 3D keypoint coordinate sequence. This allows for the acquisition of 3D spatial human body key point data. Furthermore, the terminal device filters out irrelevant objects by setting a detection area (such as a 1.5m x 1m cube space under the rear of the vehicle), retaining only the 3D key point sequence related to foot movements, thereby extracting spatial foot key point data.

[0168] Finally, the terminal device determines whether a kicking action has been detected by using spatiotemporal feature encoding and action decoding. Specifically, the terminal device uses a Long Short-Term Memory (LSTM) network to encode the 3D keypoint sequence in the spatial foot keypoint data, generating a feature vector. Then, the terminal device processes the feature vectors using a classifier (such as a fully connected neural network). Decode the action to determine whether a kicking motion has been detected.

[0169] Please see Figure 9 This application also provides a vehicle tailgate control device that can implement the above-described vehicle tailgate control method.

[0170] like Figure 9 As shown, the vehicle tailgate control device provided in this embodiment includes a data acquisition module 901, an adaptive detection module 902, and a control module 903. Among them,

[0171] The data acquisition module 901 is used to acquire target image data; the target image data is acquired by a vision sensor associated with the vehicle tailgate.

[0172] The adaptive detection module 902 is used to perform human key point detection on the target image data based on the sensor parameters of the vision sensor to obtain human key point data in the target image data; the sensor parameters are used to adaptively adjust the feature extraction weights for human key point detection on the target image data.

[0173] The control module 903 is used to perform kicking behavior recognition based on the human body key point data to obtain recognition results, and to control the tailgate of the vehicle based on the recognition results.

[0174] In some embodiments, the adaptive detection module 902 is further configured to adaptively adjust a preset initial feature extraction weight based on the sensor parameters of the visual sensor to obtain an adjusted target feature extraction weight; the initial feature extraction weight is the initial feature extraction weight for human key point detection of the target image data; and, based on the target feature extraction weight, perform human key point detection on the target image data.

[0175] In some embodiments, the adaptive detection module 902 is further configured to perform human key point detection on the target image data using a preset adaptive detection model, and to adaptively adjust the feature extraction weights for human key point detection on the target image data based on the sensor parameters using the adaptive detection model.

[0176] In some embodiments, the adaptive detection module 902 is further configured to generate a scaling factor and bias term of a convolutional block based on the sensor parameters through the adaptive detection model, so as to adaptively adjust the feature extraction weights for human key point detection of the target image data based on the scaling factor and the bias term; wherein, the convolutional block is any convolutional block in the multi-layer convolutional blocks of the adaptive detection model.

[0177] In some embodiments, the human body key point data includes planar human body key point data. The control module 903 is further configured to perform spatial mapping processing on the planar human body key point data to obtain spatial human body key point data; and to perform kicking behavior recognition based on the spatial human body key point data to obtain a recognition result.

[0178] In some embodiments, the control module 903 is further configured to perform multi-scale temporal feature fusion processing on the planar human key point data using a preset dilated temporal convolution model to obtain spatial human key point data output by the dilated temporal convolution model.

[0179] In some embodiments, the control module 903 is further configured to extract spatial foot key point data from the spatial human body key point data; perform spatiotemporal feature encoding processing on the spatial foot key point data to obtain a feature vector corresponding to the spatial foot key point data; and perform behavioral action decoding processing on the feature vector to obtain a recognition result of kicking behavior.

[0180] In some embodiments, the control module 903 is further configured to perform behavior action decoding processing on the feature vector through a preset classifier to obtain the predicted probability of the kicking behavior output by the classifier; if the predicted probability is greater than or equal to a preset probability threshold, determine that the kicking behavior is identified; and if the predicted probability is less than the preset probability threshold, determine that the kicking behavior is not identified.

[0181] It should be noted that the specific implementation of the vehicle tailgate control device provided in this application is basically the same as the specific implementation of the vehicle tailgate control method described above, and will not be repeated here.

[0182] Please see Figure 10 This application also provides a vehicle tailgate control device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-mentioned visual language model evaluation method.

[0183] In some embodiments, the vehicle tailgate control device can be any smart terminal such as an on-board hardware platform (e.g., an on-board computer), a tablet computer, a smartphone, or a wearable device.

[0184] like Figure 10 As shown, the vehicle tailgate control device provided in this application embodiment may include:

[0185] The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0186] The memory 1002 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1002 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001 to execute the vehicle tailgate control method of the embodiments of this application.

[0187] Input / output interface 1003 is used to implement information input and output;

[0188] The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0189] Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004);

[0190] The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.

[0191] This application also provides a vehicle that includes the aforementioned vehicle tailgate control device. The vehicle tailgate control device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned vehicle tailgate control method.

[0192] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described vehicle tailgate control method.

[0193] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0194] This application also provides a computer program product, including a computer program. The steps implemented by the computer program when executed by a processor are basically the same as those in the specific embodiments of the vehicle tailgate control method described above, and will not be repeated here.

[0195] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0196] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0197] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0198] Those skilled in the art will understand that all or some of the steps, apparatus / systems, and functional modules / units in the methods disclosed above can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0199] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus / system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0200] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0201] In the several embodiments provided in this application, it should be understood that the disclosed apparatus / systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0202] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0203] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0204] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0205] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for controlling a vehicle tailgate, characterized in that, The method includes: Acquire target image data; the target image data is acquired through a vision sensor associated with the vehicle's tailgate; Human key point detection is performed on the target image data based on the sensor parameters of the vision sensor to obtain human key point data in the target image data; the sensor parameters are used to adaptively adjust the feature extraction weights for human key point detection on the target image data. The kicking behavior is identified based on the human body key point data to obtain the identification result, and the tailgate of the vehicle is controlled based on the identification result; The method further includes: The model generates a scaling factor and bias term for a convolutional block based on the sensor parameters using a preset adaptive detection model. The feature extraction weights for human keypoint detection in the target image data are adaptively adjusted based on the scaling factor and the bias term. The convolutional block can be any convolutional block in the multi-layer convolutional blocks of the adaptive detection model.

2. The method according to claim 1, characterized in that, The detection of human key points in the target image data based on the sensor parameters of the visual sensor includes: The preset initial feature extraction weights are adaptively adjusted based on the sensor parameters of the visual sensor to obtain the adjusted target feature extraction weights; the initial feature extraction weights are the initial feature extraction weights for human key point detection of the target image data. Human key point detection is performed on the target image data based on the target feature extraction weights.

3. The method according to claim 1, characterized in that, The detection of human key points in the target image data based on the sensor parameters of the visual sensor includes: Human key points are detected in the target image data using a preset adaptive detection model, and the feature extraction weights for human key point detection in the target image data are adaptively adjusted based on the sensor parameters using the adaptive detection model.

4. The method according to claim 1, characterized in that, The human body key point data includes planar human body key point data, and the recognition result obtained by performing kicking behavior recognition based on the human body key point data includes: Spatial mapping processing is performed on the planar human body key point data to obtain spatial human body key point data; The recognition result is obtained by recognizing kicking behavior based on the aforementioned spatial human body key point data.

5. The method according to claim 4, characterized in that, The step of performing spatial mapping processing on the planar human body key point data to obtain spatial human body key point data includes: The planar human keypoint data is processed by multi-scale temporal feature fusion using a pre-defined dilated temporal convolution model to obtain spatial human keypoint data output by the dilated temporal convolution model.

6. The method according to claim 4, characterized in that, The recognition result obtained by performing kicking behavior recognition based on the spatial human key point data includes: Extract the spatial foot key point data from the spatial human body key point data; Spatiotemporal feature encoding is performed on the spatial foot key point data to obtain a feature vector corresponding to the spatial foot key point data; The feature vector is subjected to behavior action decoding processing to obtain the recognition result of kicking behavior.

7. The method according to claim 6, characterized in that, The step of performing behavior action decoding on the feature vector to obtain the recognition result of the kicking behavior includes: The feature vector is processed by a preset classifier to decode the action, and the predicted probability of the kicking action output by the classifier is obtained. If the predicted probability is greater than or equal to a preset probability threshold, the recognition result of the kicking behavior is determined as kicking behavior being recognized; If the predicted probability is less than a preset probability threshold, the recognition result of the kicking behavior is determined to be that the kicking behavior was not recognized.

8. A vehicle tailgate control device, characterized in that, The device includes: A data acquisition module is used to acquire target image data; the target image data is acquired through a vision sensor associated with the vehicle's tailgate. An adaptive detection module is used to detect human key points in the target image data based on the sensor parameters of the visual sensor, and obtain human key point data in the target image data; the sensor parameters are used to adaptively adjust the feature extraction weights for human key point detection in the target image data. The control module is used to identify kicking behavior based on the human body key point data, obtain the identification result, and control the tailgate of the vehicle based on the identification result; The adaptive detection module is further configured to generate a scaling factor and bias term for the convolutional block based on the sensor parameters using a preset adaptive detection model, and adaptively adjust the feature extraction weights for human key point detection of the target image data based on the scaling factor and the bias term; wherein, the convolutional block is any convolutional block in the multi-layer convolutional blocks of the adaptive detection model.

9. A vehicle tailgate control device, characterized in that, The vehicle tailgate control device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the vehicle tailgate control method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the vehicle tailgate control method as described in any one of claims 1 to 7.

11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the vehicle tailgate control method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Three-dimensional human body motion reconstruction method and device, equipment and storage medium

    CN119091052A

  • Vehicle trunk opening method and device, electronic equipment, medium and vehicle

    CN119904906A