In-vehicle temperature control method based on visual interaction of driver
Through the on-board high-definition camera and line of sight vector analysis, the in-car temperature control system is dynamically adjusted, solving the problem of insufficient perception of the driver's line of sight and gaze point in existing technologies. It achieves accurate recognition of the driver's intentions and intelligent temperature adjustment, and improves the intelligence level of in-car environment control and user experience.
Patent Information
- Application Number
- CN202510916534.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-09-19
AI Technical Summary
Existing in-vehicle temperature control systems lack accurate perception and real-time analysis of the driver's line of sight, gaze point, and intentions, resulting in an inability to dynamically adjust according to the driver's real-time needs, making it difficult to achieve intelligent adjustment and natural human-computer interaction.
The system uses an on-board high-definition camera to capture in-car images and extract the driver's facial information. The system then uses three-dimensional distance calculation and line of sight vector analysis to determine the driver's line of sight and gaze point. The system then activates the temperature control system through a sensor feedback system, dynamically adjusting the in-car temperature based on the driver's gaze behavior.
It achieves accurate recognition of the driver's intentions and intelligent adjustment of the car's temperature, improves the naturalness and efficiency of the interaction, and ensures that the car's temperature control system adapts to the driver's real-time needs.
Smart Images

Figure CN120663713A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of in-vehicle environment control, and in particular to a method for controlling in-vehicle temperature through driver visual interaction. Background Art
[0002] With the rapid development of intelligent driving technology, the automation and intelligence levels of vehicles are constantly improving. In-vehicle human-computer interaction technology has become a crucial component of intelligent driving systems. In traditional in-vehicle climate control systems, such as air conditioning temperature adjustment, drivers must manually operate buttons on the center console or touchscreen to set the temperature. This interaction method has certain limitations, especially because it distracts the driver while driving, potentially increasing driving risks. In recent years, voice recognition-based interaction technologies have gradually been introduced into in-vehicle climate control systems, enabling in-vehicle climate adjustment via voice commands. However, voice interaction is susceptible to interference in noisy environments, and the accuracy and response speed of voice commands remain limited in practical applications. Furthermore, gesture recognition-based interaction technologies are also being explored as an emerging area. However, the accuracy and reliability of gesture recognition systems for detecting hand movements depend on external conditions such as lighting and angle, and their practical application still requires further optimization. Therefore, achieving a more natural, intuitive, and efficient in-vehicle climate control method is a pressing technical challenge in the field of intelligent driving.
[0003] While existing technologies have achieved some success in the field of in-car interaction, they still face numerous shortcomings in terms of accuracy, real-time performance, and user experience. For one thing, traditional climate control systems rely heavily on active driver input, failing to perceive the driver's intent in real time and lacking the ability for proactive interaction. Furthermore, existing voice and gesture recognition technologies have limitations in accuracy and environmental adaptability. For example, voice recognition is prone to misidentification in noisy in-car environments, while gesture recognition can fail due to lighting changes or driver posture variations. Furthermore, existing technologies fail to fully utilize the driver's visual information as an interactive medium. The driver's gaze direction and gaze point are important indicators of their immediate attention and intent, but current in-car interaction systems rarely utilize this information for temperature control or other environmental adjustments. This significantly limits existing technologies in achieving more intelligent and personalized in-car climate control. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method for controlling the in-vehicle temperature through driver visual interaction, which solves the technical problem in the prior art that the in-vehicle temperature control cannot be dynamically adjusted according to the driver's real-time needs due to the lack of accurate perception and real-time analysis framework of the driver's line of sight direction, gaze point and intention, and thus it is difficult to achieve intelligent regulation of the in-vehicle temperature and improvement of the naturalness of human-computer interaction.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for controlling in-vehicle temperature through driver visual interaction, comprising:
[0008] Acquire in-car images, capture the in-car scene with an onboard high-definition camera, process the captured color images, convert them into standardized grayscale images, and extract facial information of the people in the car;
[0009] Based on the position of the facial area in the image and the spatial position relationship of fixed objects in the car, the identity of the driver and passengers is determined using a three-dimensional distance calculation method;
[0010] Obtain the driver's eye position and nose tip position, calculate the sight vector through key points, and determine the driver's sight direction and gaze point by combining reference objects in the vehicle environment;
[0011] The calculated gaze vector is compared with the position of the controller inside the car; when the driver's gaze continues to look at the controller for more than a predetermined time threshold, the in-car temperature control system is activated through the sensor feedback system, and the in-car temperature is adjusted according to the driver's gaze direction and behavior.
[0012] As a preferred embodiment of the in-vehicle temperature control method through driver visual interaction described in the present invention, the identity of the driver and the passenger is determined by calculating the three-dimensional distance between the facial area and the steering wheel in the vehicle, and a set distance threshold is used to determine the positional relationship of the facial area. If the distance is less than the threshold, the area is considered to be the driver, otherwise it is considered to be a passenger.
[0013] As a preferred embodiment of the in-vehicle temperature control method through driver visual interaction described in the present invention, the temperature control system determines whether the driver is actively adjusting the in-vehicle temperature based on the driver's line of sight and the duration of their gaze, and adjusts the air conditioning or heating system accordingly; if the driver's gaze time exceeds a set threshold, the temperature control function is activated, and the temperature is adjusted in real time through feedback from in-vehicle sensors.
[0014] As a preferred solution of the in-vehicle temperature control method through driver visual interaction of the present invention, the conversion into a standardized grayscale image includes:
[0015] Perform a comprehensive analysis of the image's brightness distribution and calculate the brightness histogram;
[0016] Redistribute the original pixel values into a wider dynamic range through a nonlinear mapping function;
[0017] The formula for the reallocation can be expressed as follows:
[0018]
[0019] HE(I(x,y)) represents the pixel value after histogram equalization, n represents the frequency of a certain pixel value, that is, the number of times the pixel value appears in the image, and H and W are the width and height of the image, respectively.
[0020] As a preferred solution of the in-car temperature control method through driver visual interaction of the present invention, extracting facial information of the occupants of the car includes:
[0021] Input image to extract features and generate feature map F;
[0022] Generate high-quality proposal regions on the feature map;
[0023] Output two key prediction values, including the classification score of the possibility that the region contains a face and the bounding box regression parameters for adjusting the position and size of the anchor box;
[0024] Calculate the foreground score of each candidate box;
[0025] Use the classifier to accurately classify each candidate box.
[0026] As a preferred solution of the in-vehicle temperature control method through driver visual interaction of the present invention, the characteristic graph is shown as follows:
[0027]
[0028] Among them, F represents the output feature map, Conv represents the convolution operation, I represents the input image or input feature map, K represents the total number of convolution kernels, and W k is the weight parameter of the kth convolution kernel, b k is the bias parameter corresponding to the kth convolution kernel.
[0029] As a preferred solution of the in-vehicle temperature control method through driver visual interaction described in the present invention, when the distance between the detected facial area and the center of the steering wheel is less than a preset threshold, the system will determine the area as the driver's position; conversely, when the distance exceeds the threshold, it will be determined as the passenger's position.
[0030] As a preferred embodiment of the in-vehicle temperature control method through driver visual interaction of the present invention, the step of determining the driver's line of sight includes:
[0031] Get the precise coordinates of pupil center and nose tip;
[0032] Calculate the initial sight direction vector using the coordinates of the pupil center and the nose tip;
[0033] Normalizing the sight line vector to convert it into a standard vector of unit length;
[0034] The gaze distance is set based on the normalized gaze vector and the gaze point position is calculated;
[0035] Map the calculated 3D gaze point coordinates to the normalized space.
[0036] In a second aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the method for controlling the in-vehicle temperature through driver visual interaction as described in the first aspect of the present invention is implemented.
[0037] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, any step of the method for controlling the in-vehicle temperature through driver visual interaction as described in the first aspect of the present invention is implemented.
[0038] The beneficial effects of the present invention are as follows: by using an on-board high-definition camera to capture images of the in-car scene, facial information of the people in the car is extracted, and the identities of the driver and passengers are dynamically judged based on the position of the facial area in the image combined with the spatial relationship of fixed objects in the car; by obtaining the position of the driver's eyes and the tip of the nose, the sight line vector is calculated and the gaze point position is determined, and the driver's sight direction and gaze behavior are perceived in real time; the driver's sight line direction is matched with the position of the in-car controller, and the in-car temperature adjustment system is activated in combination with the vehicle sensor feedback system, and the in-car temperature is dynamically adjusted according to the driver's gaze behavior, thereby realizing accurate recognition of the driver's intentions and intelligent adjustment of the in-car temperature; by introducing the sight line vector calculation, gaze point detection and time threshold judgment mechanism, it is ensured that the in-car temperature control system can adapt to the real-time needs of the driver, thereby improving the naturalness and efficiency of the interaction; an efficient in-car temperature control method is constructed from three levels of accurate perception, real-time response and intelligent adjustment, which greatly improves the intelligence level of in-car environmental control and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0040] Figure 1 The figure is a flow chart of a method for controlling in-vehicle temperature through driver visual interaction.
[0041] Figure 2 Flowchart of face detection for in-vehicle temperature control method through driver visual interaction.
[0042] Figure 3 The present invention is a flowchart of a method for controlling in-vehicle temperature through driver visual interaction, which determines whether the driver is looking at the screen.
[0043] Figure 4 Schematic diagram of the acquisition module of the in-vehicle temperature control method through driver visual interaction. DETAILED DESCRIPTION
[0044] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0045] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0046] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0047] Reference Figures 1 to 4 , is an embodiment of the present invention, which provides a method for controlling in-vehicle temperature through driver visual interaction, comprising the following steps:
[0048] Step 1: In the implementation of this invention, the core component of the data acquisition system is a high-performance, in-vehicle HD camera system. This system has been meticulously designed and laid out, with the HD camera installed within the vehicle's A-pillar. This placement, carefully researched and field-tested, not only ensures that the camera's installation and use will not interfere with the vehicle's normal operation, but also ensures continuous and complete image capture of the driver's seat area at the optimal angle. The camera's placement also fully considers the vehicle's interior spatial layout, effectively avoiding potential visual obstructions caused by the steering wheel, interior decorations, and other factors, ensuring continuous and complete image acquisition. This system utilizes a professional-grade camera with ultra-high resolution and an ultra-fast frame rate. Its exceptional performance is demonstrated by: first, it automatically adjusts exposure parameters under strong daylight conditions to avoid overexposure or overbrightening; second, it maintains excellent image quality even in low-light conditions at night, utilizing advanced photosensitive elements to ensure image clarity; and finally, it rapidly responds and adjusts to maintain consistent image quality even in complex driving environments, where lighting changes frequently and alternating between bright and dark. This all-weather, all-scene image acquisition capability provides high-quality raw data support for subsequent image processing and analysis.
[0049] Step 2:
[0050] During the image acquisition process in a vehicle-mounted environment, the original image is often interfered with by various factors and generates noise. These interference sources mainly include: mechanical vibrations caused by vehicle driving, changes in illumination at different time periods (such as changes in strong light at tunnel entrances and exits, alternating light and dark caused by mottled shadows of trees), and interference caused by factors such as air conditioning outlet and passenger movement in the car. In order to effectively deal with these complex noise problems, the present invention has developed a set of adaptive denoising filtering systems based on Gaussian functions. The system adopts a dynamic weight allocation mechanism to assign corresponding weight coefficients to pixels at different positions by accurately calculating the spatial relationship between pixels. During processing, the system automatically adjusts the weight size according to the distance relationship between the target pixel and the surrounding pixels, so that the closer the pixel is, the greater the influence, while the influence of the farther the pixel is exponentially attenuated. This processing method can not only effectively eliminate random noise, but also retain the edge features and detail information of the image to the greatest extent. The denoising filter formula of the present invention is as follows:
[0051]
[0052] After completing the initial noise processing, the system will start the intelligent contrast optimization module. This module uses an improved histogram equalization technology to adaptively adjust the contrast of the image. This technology first conducts a comprehensive analysis of the brightness distribution of the image, calculates the brightness histogram, and then redistributes the original pixel values to a wider dynamic range through a nonlinear mapping function. This processing method is particularly suitable for processing images under complex lighting conditions. For example, it can improve the visibility of dark details in strong backlighting and enhance the overall clarity of the image in low-light environments. The system also integrates a local adaptive enhancement algorithm, which can perform differentiated processing based on the characteristics of different areas of the image to ensure that while improving the overall contrast, there is no loss of local details.
[0053] The specific implementation method is to redistribute the pixel values so that the frequency of different grayscale values in the image is more uniform. The formula is:
[0054]
[0055] HE(I(x,y)) represents the pixel value after histogram equalization, n represents the frequency of a certain pixel value, that is, the number of times the pixel value appears in the image, and H and W are the width and height of the image, respectively.
[0056] Histogram equalization will increase the overall contrast of the image, making details in dark or bright areas clearer, which helps in the recognition and extraction of facial features.
[0057] In order to optimize the efficiency of subsequent processing, the system adopts an innovative intelligent grayscale conversion scheme. Unlike the traditional simple RGB weighted average, this system first converts the image into a color space that is more in line with the visual characteristics of the human eye. In this optimized color space, the system can more accurately extract brightness information, thereby generating a more natural and detail-rich grayscale image. This conversion method not only takes into account the different contributions of different color channels to the human eye's brightness perception, but also incorporates the results of psychophysics research, making the converted grayscale image more in line with the natural perception of the human eye. In addition, the system also includes an automatic color compensation mechanism, which can significantly reduce the amount of data while retaining key details, thereby improving the efficiency of subsequent processing. The grayscale formula used in the present invention is as follows:
[0058] Y=0.299×R+0.587×G+0.114×B
[0059] Among them, R, G, and B represent the pixel values of the red, green, and blue channels respectively, and Y is the brightness component of the color space. The Y value is directly taken as the grayscale value.
[0060] After grayscale conversion, the image becomes a single-channel image, with each pixel value representing brightness information, ranging from black (0) to white (255). This greatly simplifies the image data structure and saves computing resources for subsequent algorithms.
[0061] Finally, the pixel values are precisely mapped to the interval [0, 1], using a nonlinear mapping function to ensure a balanced data distribution. The normalization process not only considers the global data distribution but also introduces a local adaptive mechanism that can dynamically adjust based on the characteristics of different regions of the image. This processing method ensures data consistency while preserving the local characteristics of the image, providing high-quality input data for subsequent deep learning algorithms. The system also has a data quality monitoring mechanism designed to evaluate the normalization effect in real time to ensure that the processing results meet the preset quality standards. It should be noted that the formula for image normalization is:
[0062]
[0063] Among them, I norm is the normalized image pixel value, in the range of [0,1], and Y is the grayscale pixel value (ranging from 0 to 255).
[0064] Normalization limits the range of pixel values to a fixed interval, which helps improve the training effect and computational stability of the model, and also facilitates comparison and unified processing between different images.
[0065] Step 3: This system uses a multi-layer convolutional neural network structure for feature extraction. After the input layer receives the preprocessed image data, it extracts the visual features in the image layer by layer through a combination of multiple convolutional layers and pooling layers. The shallow network is mainly responsible for extracting basic features such as edges and textures, while the deep network can capture more abstract high-level semantic features. This hierarchical feature extraction method ensures that the system can fully obtain the key information in the image. The feature extraction network adopts an improved backbone network structure, which enhances the expressive ability of features through residual connections and feature pyramid structures while maintaining computational efficiency. The generated feature map contains rich spatial and semantic information, providing a reliable data foundation for subsequent target detection.
[0066] First, extract features from the input image and generate a feature map F, which provides rich image information for subsequent processing. This is shown in the following formula:
[0067]
[0068] Among them, F represents the output feature map, Conv represents the convolution operation, I represents the input image or input feature map, K represents the total number of convolution kernels, and W k is the weight parameter of the kth convolution kernel, b kis the bias parameter corresponding to the kth convolution kernel.
[0069] The Region Proposal Network (RPN) is one of the core components of the system, and its main responsibility is to generate high-quality candidate regions on the feature map. The network uses a sliding window approach to generate multiple anchor boxes (Anchor Box) of different scales and aspect ratios at each position in the feature map. In order to adapt to faces of different sizes and postures, the system carefully designs the size distribution of anchor boxes, including multiple basic sizes (such as 64×64, 128×128, 256×256 pixels) and multiple aspect ratios (such as 1:1, 1:2, 2:1, etc.). For each anchor box, the RPN network outputs two key prediction values: one is the classification score indicating the possibility that the region contains a face, and the other is the bounding box regression parameter used to accurately adjust the position and size of the anchor box. The formula is as follows:
[0070] Target frame position = (x+dx,y+dy,w×e dw ,h×e dh )
[0071] Among them, dx, dy, dw, dh are the predicted bounding box regression parameters, and x, y, w, h are the initial coordinates of the anchor box.
[0072] To further improve detection accuracy, the system has designed a complete candidate box optimization process. First, the foreground score of each candidate box is calculated through a dedicated scoring network. The scoring criteria comprehensively consider multiple factors such as the clarity, completeness, and confidence of the target. Then, the system uses an improved soft non-maximum suppression (NMS) algorithm to process overlapping candidate boxes. This improved algorithm can dynamically adjust the suppression strength according to the characteristics of the detected target, while retaining high-quality detection frames and avoiding the accidental deletion of valid targets. After that, the system unifies the candidate boxes of different sizes to a fixed size through the ROI pooling operation. This step not only maintains the integrity of the target features, but also significantly improves the computational efficiency of subsequent processing.
[0073] Among them, each candidate box calculates the foreground score: foreground score = sigmoid(s)
[0074] Where s is the output classification score. Non-maximum suppression (NMS) is used to remove overlapping candidate boxes and only retain boxes with high confidence.
[0075] The candidate boxes are then pooled into a uniform size and fed into the classifier. The classifier accurately classifies each candidate box and further adjusts the position of the bounding box:
[0076] Target frame coordinates = (x'+dx',y'+dy',w'×e dw' ,h'×e dh' )
[0077] Among them, dx',dy',dw'dh' are the regression parameters output by the classifier.
[0078] By combining the above operations, faces can be efficiently located under different lighting and posture conditions, providing reliable data support for gaze tracking and intelligent temperature control.
[0079] Step 4:
[0080] A complete in-vehicle 3D coordinate system was established. Within this coordinate system, the HD camera serves as the reference point, with its position defined as the origin. Through pre-calibration, the system obtains the precise coordinates of key fixed objects within the vehicle (such as the center of the steering wheel). These coordinates serve as the system's fixed reference points. Simultaneously, through real-time image processing and depth estimation algorithms, the system accurately determines the 3D spatial coordinates of detected facial regions within this coordinate system.
[0081] In practice, the system primarily focuses on the spatial distance between the facial region and the center of the steering wheel. Using the three-dimensional Euclidean distance formula, the system can calculate the precise distance between each detected facial region and the steering wheel center in real time. This distance value is a key indicator for determining identity. Based on extensive experimental data, the system has established an adaptive threshold determination mechanism. When the distance between a detected facial region and the steering wheel center is less than a preset threshold, the system identifies the area as the driver's location; conversely, when the distance exceeds the threshold, it identifies the area as a passenger's location.
[0082] For example, by detecting the facial position of the driver and passenger in the image transmitted by the camera, and the spatial position relationship relative to the fixed objects in the car (such as the steering wheel and seats), the system can effectively determine the identity. Assume that the coordinate system of the camera is (x, y, z) and the coordinate of the center of the steering wheel is P d =(x d ,y d ,z d ), and the coordinates of the facial region are P f =(x f ,y f, z f ). The relative position between the facial area and the steering wheel is calculated using the three-dimensional distance formula:
[0083] If the distance is less than a preset threshold, the facial area is determined to be within the driver's control range and thus identified as the driver; if the distance exceeds the threshold, the facial area is identified as a passenger. This positional relationship judgment is very effective in in-car scenarios, enabling rapid and accurate differentiation between drivers and passengers, providing precise identity information for subsequent gaze tracking and intelligent control.
[0084] Step 5: Project the three-dimensional coordinates of the line of sight to two-dimensional coordinates
[0085] The system establishes two key coordinate systems: a global coordinate system and a local coordinate system. The global coordinate system uses the vehicle as a reference, defines the front as the Z axis, the left and right directions as the X axis, and the up and down directions as the Y axis. This definition method conforms to the natural characteristics of the vehicle space. At the same time, a local coordinate system is established at the controller position to achieve more accurate local space positioning. The establishment of these two coordinate systems lays the foundation for subsequent space conversion and projection calculations. Perspective projection is used to convert points in three-dimensional space to a two-dimensional plane on the screen or controller. The basic formula for perspective projection of the present invention is:
[0086]
[0087] Where (x, y, z) is a point in three-dimensional space, (x', y') is a point in two-dimensional plane coordinates, and f is the focal length.
[0088] Perspective projection is used to transform a 3D point onto a 2D plane. This projection method accurately simulates the human eye's perception of the object. Adjusting the focal length parameter f allows for precise control of the projection's accuracy. The projection process considers all dimensions of the spatial point (x, y, z), ensuring that the resulting 2D coordinates accurately reflect the original spatial positional relationships.
[0089] Furthermore, in order to achieve accurate conversion between different coordinate systems, the angle alignment between the coordinate systems is achieved through the rotation matrix, and the position offset of the coordinate origin is handled through the translation matrix. This conversion mechanism ensures that the data in the local coordinate system can be accurately mapped to the global coordinate system. To perform coordinate changes and align the local coordinate system with the global coordinate system, the rotation matrix can be used:
[0090]
[0091] This matrix can be used to transform local coordinates into global coordinates.
[0092] Translation matrix:
[0093]
[0094] Where dx, dy, and dz are the translations of the controller relative to the reference point.
[0095] The obtained coordinate values are standardized to ensure the consistency of the gaze point position under different lighting, angles and distances.
[0096] Step 6: Driver's sight projection
[0097] By analyzing facial feature points such as the eye position and the nose tip, the accurate sight direction vector is calculated. This calculation process first obtains the precise coordinates of the pupil center and the nose tip, and then obtains the initial sight direction through vector calculation. In order to improve the calculation accuracy, the system normalizes the obtained sight vector and converts it into a standard vector of unit length. This processing method ensures the accuracy and consistency of subsequent calculations. The sight vector calculation calculates the sight direction based on reference points such as the eye position and the nose tip, and calculates the sight direction through the pupil center (x p, y p ,z p ) and nose tip (x n ,y n ,z n ), the present invention provides a method for calculating the sight line vector:
[0098]
[0099] The sight vector is normalized to unit length and is calculated as follows:
[0100]
[0101] The normalized view vector is:
[0102]
[0103] After obtaining the direction of sight, the system calculates the specific gaze point position by setting an appropriate gaze distance. The calculation of the gaze point takes into account multiple factors, including the characteristics of the sight vector, the actual distance between the driver and the target object, and the specific application scenario requirements. The system introduces an adjustable parameter k, which can accurately control the position of the gaze point by adjusting the k value: when k = 0, it means that the gaze point is at the pupil position, and when k = 1, it means that the gaze point is at the tip of the nose. Attention point calculation, set the gaze point distance d, the coordinates of the gaze point (x gaze ,y gaze ,z gaze ) can be expressed as:
[0104]
[0105] The value of k depends on several factors, including the length of the gaze vector, the distance between the driver and the target object (such as the A-pillar of a car), and the requirements of the application scenario. k = 0: indicates that the gaze point is completely at the pupil position.
[0106] k=1: indicates that the gaze point is at the tip of the nose.
[0107] Map the calculated three-dimensional gaze point coordinates to the standardized space. If the range of the standardized space is:
[0108]
[0109] Among them, P is the current coordinate, P min and P max are the minimum and maximum values of the coordinate range. The obtained standardized line of sight coordinates are passed to the controller to complete the control of the system.
[0110] Finally, the calculated gaze point coordinates are normalized and mapped to a predefined standard spatial range. This normalization ensures stable positioning under varying lighting conditions, viewing angles, and distances. Furthermore, the system implements a time threshold (default 1 second) to determine the driver's gaze intent. Only when the gaze duration exceeds the threshold will the system trigger the corresponding control command.
[0111] This embodiment also provides a computer device, which is suitable for the case of a method for controlling the in-vehicle temperature through visual interaction with the driver, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for controlling the in-vehicle temperature through visual interaction with the driver as proposed in the above embodiment.
[0112] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.
[0113] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for controlling the in-vehicle temperature through driver visual interaction as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0114] In summary, the present invention collects in-vehicle scene images through an on-board high-definition camera, extracts facial information of people in the car, and dynamically determines the identities of the driver and passengers based on the position of the facial area in the image and the spatial relationship of fixed objects in the car; by obtaining the position of the driver's eyes and the tip of the nose, the sight line vector is calculated and the gaze point position is determined, and the driver's sight direction and gaze behavior are perceived in real time; the driver's sight line direction is matched with the position of the in-vehicle controller, and the in-vehicle temperature adjustment system is activated in combination with the vehicle sensor feedback system, and the in-vehicle temperature is dynamically adjusted according to the driver's gaze behavior, thereby realizing accurate recognition of the driver's intentions and intelligent adjustment of the in-vehicle temperature; by introducing the sight line vector calculation, gaze point detection and time threshold judgment mechanism, it ensures that the in-vehicle temperature control system can adapt to the driver's real-time needs, and improves the naturalness and efficiency of the interaction; constructs an efficient in-vehicle temperature control method from three levels of accurate perception, real-time response and intelligent adjustment, which greatly improves the intelligence level of in-vehicle environmental control and user experience.
[0115] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for controlling in-vehicle temperature through driver visual interaction, characterized in that: include, Acquire in-car images, capture the in-car scene with an onboard high-definition camera, process the captured color images, convert them into standardized grayscale images, and extract facial information of the people in the car; Based on the position of the facial area in the image and the spatial position relationship of fixed objects in the car, the identity of the driver and passengers is determined using a three-dimensional distance calculation method; Obtain the driver's eye position and nose tip position, calculate the sight vector through key points, and determine the driver's sight direction and gaze point by combining reference objects in the vehicle environment; Compare the calculated gaze vector with the position of the controller inside the vehicle; When the driver's gaze is continuously fixed on the controller for more than a predetermined time threshold, the in-vehicle temperature control system is activated through the sensor feedback system, and the in-vehicle temperature is adjusted according to the driver's gaze direction and behavior.
2. The method for controlling in-vehicle temperature through driver visual interaction according to claim 1, wherein: The identification of the driver and the passenger is determined by calculating the three-dimensional distance between the facial area and the steering wheel in the car, and using a set distance threshold to determine the positional relationship of the facial area. If the distance is less than the threshold, the area is considered to be the driver, otherwise it is a passenger.
3. The method for controlling in-vehicle temperature through driver visual interaction according to claim 2, wherein: The temperature control system determines whether the driver is actively adjusting the temperature in the car based on the direction and duration of the driver's gaze, and makes corresponding adjustments to the air conditioning or heating system; if the driver's gaze time exceeds a set threshold, the temperature control function is activated and the temperature is adjusted in real time through feedback from the in-car sensors.
4. The method for controlling in-vehicle temperature through driver visual interaction according to claim 3, wherein: The conversion into a standardized grayscale image includes: Perform a comprehensive analysis of the image's brightness distribution and calculate the brightness histogram; Redistribute the original pixel values into a wider dynamic range through a nonlinear mapping function; The formula for the reallocation can be expressed as follows: HE(I(x,y)) represents the pixel value after histogram equalization, n represents the frequency of a certain pixel value, that is, the number of times the pixel value appears in the image, and H and W are the width and height of the image, respectively.
5. The method for controlling in-vehicle temperature through driver visual interaction according to claim 4, wherein: Extracting facial information of people in the car includes: Input image to extract features and generate feature map F; Generate high-quality proposal regions on the feature map; Output two key prediction values, including the classification score of the possibility that the region contains a face and the bounding box regression parameters for adjusting the position and size of the anchor box; Calculate the foreground score of each candidate box; Use the classifier to accurately classify each candidate box.
6. The method for controlling in-vehicle temperature through driver visual interaction according to claim 5, wherein: The feature map is shown below: Among them, F represents the output feature map, Conv represents the convolution operation, I represents the input image or input feature map, K represents the total number of convolution kernels, and W k is the weight parameter of the kth convolution kernel, b k is the bias parameter corresponding to the kth convolution kernel.
7. The method for controlling in-vehicle temperature through driver visual interaction according to claim 6, wherein: When the distance between the detected facial area and the center of the steering wheel is less than a preset threshold, the system will determine the area as the driver's position; conversely, when the distance exceeds the threshold, it will be determined as the passenger's position.
8. The method for controlling in-vehicle temperature through driver visual interaction according to claim 7, wherein: Determining the driver's sight direction includes: Get the precise coordinates of pupil center and nose tip; Calculate the initial sight direction vector using the coordinates of the pupil center and the nose tip; Normalizing the sight line vector to convert it into a standard vector of unit length; The gaze distance is set based on the normalized gaze vector and the gaze point position is calculated; Map the calculated 3D gaze point coordinates to the normalized space.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for controlling the in-vehicle temperature through driver visual interaction according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for controlling the in-vehicle temperature through driver visual interaction according to any one of claims 1 to 8 are implemented.