A non-contact gesture and voice fusion window control interaction method and system

By generating an eight-bit binary addressing mask and processing acoustic data streams within the vehicle, the problem of unutilized spatial pointing information in gesture recognition is solved, enabling efficient voice acquisition and accurate user intent recognition in complex acoustic environments, while reducing system power consumption and data processing burden.

CN122493844APending Publication Date: 2026-07-31SHANDONG YUQIANG HARDWARE PROD CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG YUQIANG HARDWARE PROD CO LTD
Filing Date
2026-05-07
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, the spatial pointing information of gesture recognition is not effectively utilized, which leads to environmental noise and interference sources affecting the accuracy of recognizing the user's true intentions during the voice signal acquisition process, especially in complex acoustic environments.

Method used

By acquiring three-dimensional voxel data inside the vehicle, an eight-bit binary addressing mask is generated. This mask is then used to cover the memory access channel of the acoustic controller, blocking the audio sampling stream in non-target areas. Combined with the temporal energy changes of the acoustic data stream, a hardware interrupt signal is generated. Visual data is monitored and captured, a multi-dimensional target constraint tensor is constructed, and control instructions for the terminal device are generated to realize the mechanical action of the vehicle window.

Benefits of technology

It effectively suppressed environmental noise and interference sources, improved the anti-interference capability of voice command acquisition, reduced system power consumption and data processing burden, and improved the accuracy of user intent recognition and system robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493844A_ABST
    Figure CN122493844A_ABST
Patent Text Reader

Abstract

This invention belongs to the technical field of in-vehicle intelligent cockpit control, and relates to a non-contact gesture and voice fusion window control interaction method and system, including the following steps: generating an eight-bit binary addressing mask to limit the acoustic receiving area; extracting the spatially gating acoustic data stream; generating a global hardware interrupt trigger signal for inversely constraining the life boundary of visual computing; separating and extracting the locked interaction frame sequence; constructing a multi-dimensional target constraint tensor containing a dual safety confirmation base of physical spatial coordinate pointing and structural morphological motion change; parsing the multi-dimensional target constraint tensor to generate irreversible terminal physical device linkage control execution instructions, completing the window mechanical action drive process in a closed-loop state. This invention solves the technical problem of how to utilize the spatial directionality of gestures to enhance the anti-interference capability of voice command acquisition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of vehicle-mounted intelligent cockpit control, and relates to a non-contact gesture and voice fusion window control interaction method and system. Background Technology

[0002] In the field of intelligent cockpit technology, contactless human-machine interaction solutions have become a research hotspot to improve driving convenience and safety. These solutions, through voice control and gesture recognition, aim to reduce the driver's reliance on physical buttons and the shift in their gaze during operation. Chinese patent application CN114564102A discloses an automotive cockpit interaction method. Such prior art typically employs a multimodal fusion strategy, such as triggering voice recognition through gestures, or combining independent voice and gesture recognition results at the decision-making level. In this mode, the system processes different modal information relatively independently.

[0003] However, in the aforementioned technical solutions, the spatial directional information contained in gesture recognition is not effectively utilized to optimize the voice signal acquisition process. Specifically, when the system receives voice commands, the audio signal acquired by the microphone still contains environmental noise and interference sources from non-directional areas. This information processing method, to some extent, affects the system's accuracy in recognizing the user's true intentions in complex acoustic environments.

[0004] Therefore, how to utilize the spatial directionality of gestures to enhance the anti-interference capability of voice command acquisition is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention

[0005] In a first aspect, the present invention provides a non-contact gesture and voice-integrated window control interaction method, comprising the following steps: S1. Acquire three-dimensional spatial voxel data inside the vehicle and construct a three-dimensional spatial direction vector to generate an eight-bit binary addressing mask for limiting the acoustic receiving area. S2. Call an eight-bit binary addressing mask to cover the memory access channel port of the acoustic controller, perform physical-level region blocking on the preset raw audio sampling stream, and extract the spatially gated acoustic data stream; S3. Monitor the temporal energy mutation nodes of the spatially gated acoustic data stream and generate a global hardware interrupt trigger signal for inversely constraining the life boundary of visual computing. S4. Based on the absolute timestamp carried by the global hardware interrupt trigger signal, perform an irreversible temporal truncation and forced cutting operation on the visual image sequence that is rolled into the front-end hardware queue, and separate and extract the locked interaction frame sequence. S5. By integrating the locked interactive frame sequence and the preset duration voice command segments in the spatially gated acoustic data stream, a multi-dimensional target constraint tensor is constructed that includes a dual security confirmation base containing physical space coordinate pointing and structural morphological motion changes. S6. Analyze the multi-dimensional target constraint tensor to generate irreversible transferable terminal physical device linkage control execution instructions, and complete the mechanical motion drive process of the car window in the closed loop state.

[0006] A further aspect of the present invention generates an eight-bit binary addressing mask for defining the acoustic receiving area, comprising the following steps: Activate the in-vehicle environment surface monitoring sensor to collect real-time continuous spatial frame sequences, and analyze the continuous spatial frame sequences to obtain interactive data of relative displacement density map characterizing the physical occupation of the user's hand; Extract continuous edge extreme points from the relative displacement density map interaction data, calculate the spatial mapping geometric relationship between the continuous edge extreme points and the fixed physical coordinate system of each independent window contained in the target vehicle, and output the real-time iterative three-dimensional spatial direction vector. Based on the preset coordinate space hard-coded forced mapping rules, the three-dimensional spatial direction vector is reduced and discretized to generate an eight-bit binary addressing mask that specifically points to the controlled window area of ​​the target.

[0007] A further aspect of the present invention outputs a real-time iterative three-dimensional spatial direction vector, comprising the following steps: Edge topology analysis was performed on the relative displacement density map interaction data, and the moving voxel clusters representing the hand were identified by the connected component labeling algorithm. Calculate the center of mass of the moving voxel cluster that is farthest from the preset user shoulder reference coordinates as the continuous edge extreme point; Extract the three-dimensional coordinates of the geometric center point corresponding to the fixed physical coordinate system of each independent window, calculate the direction vector between the three-dimensional coordinates of the continuous edge extreme points and the geometric center point, and generate a three-dimensional spatial direction vector.

[0008] A further aspect of the present invention involves extracting the spatially gated acoustic data stream, including the following steps: The microphone array pickup baseband node located in the controlled space is activated to continuously capture the ambient audio pulse code modulation stream with a preset upper limit frequency as the raw audio sampling stream, and the raw audio sampling stream is cyclically pushed into the acoustic ring buffer of the bypass architecture. The default omnidirectional digital signal reading gate of the clamping acoustic ring buffer is in a physical blocking state, cutting off the overflow path of boundless background wind noise and reverberation noise to the central processing unit. An eight-bit binary addressing mask is used as an addressing and positioning pointer and overwritten into the memory access read address pin which is in a physically blocked state. This forces the beamforming weight matrix dedicated channel of the corresponding target controlled window area to be partially turned on, so that the pickup beam can only capture the ambient sound of the physical location coordinates pointed to by the eight-bit binary addressing mask and output an exclusive spatially selected acoustic data stream.

[0009] A further aspect of the present invention generates a global hardware interrupt trigger signal for inversely constraining the life boundary of visual computing, comprising the following steps: The spatially gated acoustic data stream is guided unidirectionally into the micro digital signal processor, continuously initiating transient physical edge scanning calculations in low-power cycles; The steep offset value of the high-frequency energy mutation slope in a specified microsecond-level time-sliding channel is extracted by a differential filtering algorithm; When the steep bias value of the high-frequency energy mutation slope exceeds the fixed absolute onset voltage threshold, the transient surge of the initial speech is intercepted at the physical level, and a global hardware interrupt trigger signal with an absolute timestamp is thrown outward.

[0010] A further aspect of this invention involves extracting the steep offset value of the high-frequency energy abrupt change slope in a spatially gated acoustic data stream within a specified microsecond-level time-sliding channel using a differential filtering algorithm, comprising the following steps: High-pass digital filters are applied to the spatially gated acoustic data stream to remove low-frequency components; Calculate the short-time energy of the current window and the short-time energy of the previous window within a specified microsecond-level sliding channel; Calculate the difference between the short-time energy of the current window and the short-time energy of the previous window, and divide the difference by the window stepping time interval to generate a steep bias value for the high-frequency energy mutation slope.

[0011] A further aspect of the present invention involves separating and extracting the locked interaction frame sequence, including the following steps: The global hardware interrupt trigger signal transmitted to the system control bus is captured through an independent channel, and the global hardware interrupt trigger signal is inverted and promoted to the highest level interrupt register of the vision capture subsystem for response. The absolute timestamp carried by the global hardware interrupt trigger signal is used to forcibly freeze the lifecycle of the unlimited scrolling overwrite calculation of the full original physical images in the queue buffer storage pool of the visual capture subsystem; Using the absolute timestamp as the truncation center axis instruction pointer, the background frame outside the preset time window without related actions is hard-trimmed in both directions and discarded. Only the set of image frames that are immediately before and after the absolute timestamp are retained and combined to form a locked interactive frame sequence with exclusive properties.

[0012] A further aspect of this invention involves constructing a multidimensional target constraint tensor that includes both physical space coordinate orientation and structural morphological motion transformation for dual safety verification of the base, comprising the following steps: By penetrating and reading the locked interaction frame sequence, the point with the smallest depth value in the dynamic human point cloud cluster of the locked interaction frame sequence is located as the user's fingertip point. The pixel displacement gradient component of the user's fingertip point during the existence of the locked interaction frame sequence is analyzed to confirm the local push-pull motion vector that triggers action feedback in the controlled window area of ​​the target. Simultaneously extract audio feature island segments that immediately follow the absolute timestamp in the spatially gated acoustic data stream as short trigger confirmation recordings, and extract the corresponding audio fundamental frequency fluctuation envelope layer features; The local push-pull motion vector, which serves as the structural driving verification condition, and the audio fundamental frequency fluctuation envelope layer feature, which serves as the acoustic secondary verification condition, are merged and combined with the target window physical displacement identifier obtained by converting the eight-bit binary addressing mask. Through matrix orthogonal filling and splicing, a multi-dimensional target constraint tensor is generated to wake up the underlying electromechanical assembly switch of the device.

[0013] A further aspect of the present invention completes the mechanical action drive process of the vehicle window in a closed-loop state, including the following steps: The multidimensional target constraint tensor is fed into the edge gateway node for lookup-level lightweight decoding and matching without semantic computation, so as to obtain a deterministic execution password carrying the unique physical entity identification code of the target window and the target mechanical execution logic sequence. Translate the deterministic execution command to generate a DC pulse duty cycle wide waveform for the drive terminal anti-pinch lifting module and a synchronously bound motor forward and reverse rotation level switching signal; The DC pulse duty cycle wide waveform and the synchronously bound motor forward and reverse rotation level switching signal are sent to the base servo motor that is precisely corresponding to the unique physical entity identification code of the target window. The drive gear produces a constant power physical displacement, thus completing the mechanical action drive process of the window to achieve the expected interactive action result.

[0014] Secondly, the present invention provides a contactless gesture and voice-integrated window control interaction system, comprising the following modules: The spatial pointing mask generation module acquires three-dimensional spatial voxel data inside the vehicle and constructs a three-dimensional spatial direction vector to generate an eight-bit binary addressing mask for defining the acoustic receiving area. The acoustic data gating module calls an eight-bit binary addressing mask to cover the memory access channel port of the acoustic controller, performs physical-level region blocking on the preset raw audio sampling stream, and extracts the spatially gating acoustic data stream. The voice initiation monitoring module monitors the temporal energy mutation nodes of the spatially gated acoustic data stream and generates a global hardware interrupt trigger signal for inversely constraining the life boundary of visual computing. The visual frame locking module performs an irreversible temporal truncation and forced cutting operation on the visual image sequence that is continuously stored in the front-end hardware queue based on the absolute timestamp carried by the global hardware interrupt trigger signal, and separates and extracts the locking interaction frame sequence. The multimodal fusion tensor construction module integrates the locked interactive frame sequence and the preset duration voice command segments in the spatially gating acoustic data stream to construct a multidimensional target constraint tensor that includes a dual security confirmation base of physical space coordinate orientation and structural morphological motion change. The control command decoding and execution module parses the multi-dimensional target constraint tensor to generate irreversible transferable terminal physical device linkage control execution commands, completing the mechanical motion drive process of the car window in a closed-loop state.

[0015] In summary, the present invention has the following beneficial technical effects: 1. The spatial pointing information generated by the user's gesture is converted into a hardware-level addressing mask, which directly affects the memory access channel of the acoustic controller. Compared with signal post-processing techniques at the software level, this solution blocks memory access to non-target areas at the physical data acquisition level, causing the beamforming weights of the microphone array to prioritize the direction of the gesture. This allows for preliminary suppression of environmental noise or interfering human voices from other directions before the acoustic data enters the central processing unit, thereby helping to improve the input signal-to-noise ratio in subsequent speech processing stages. 2. A hardware interrupt signal is generated by utilizing temporal energy changes in a spatially gated acoustic data stream. Compared to conventional methods that require continuous execution of visual analysis algorithms to capture gestures, this solution uses relatively low-power acoustic monitoring as a trigger for system wake-up and data localization. When a target acoustic event is detected, the system extracts visual data corresponding to the event's time point from the front-end scrolling buffer's visual image sequence based on the timestamp carried by the interrupt signal. This approach reduces the continuous analysis and computation of large amounts of irrelevant visual data during periods without interactive intent, thereby helping to reduce the system's average power consumption and overall data processing burden, while simultaneously achieving effective temporal correlation between acoustic and visual data. 3. By constructing a multi-dimensional data structure that includes static spatial orientation of gestures, dynamic changes in gestures, and acoustic confirmation information, and requiring these dimensions to meet preset consistency conditions in time and space, an effective control command can be formed. By adding the constraint dimension required for judging the validity of the command, it is possible to better distinguish between unintentional gestures by the user or similar voices that occur accidentally in the environment, thereby reducing false triggering of commands due to interference from single-modal information and providing a more robust technical foundation for system control. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. The drawings are used to provide a further understanding of the present invention.

[0017] Figure 1 A flowchart illustrating an embodiment of this application is disclosed.

[0018] Figure 2 Structural schematic diagrams of embodiments of this application are disclosed. Detailed Implementation

[0019] The following is in conjunction with the appendix Figure 1 - Figure 2 A preferred description of the present invention is provided below.

[0020] See attached document Figure 1 This invention proposes a non-contact gesture and voice fusion window control interaction method, including the following steps: S1. Acquire three-dimensional spatial voxel data inside the vehicle and construct a three-dimensional spatial direction vector to generate an eight-bit binary addressing mask for limiting the acoustic receiving area. S2. Call an eight-bit binary addressing mask to cover the memory access channel port of the acoustic controller, perform physical-level region blocking on the preset raw audio sampling stream, and extract the spatially gated acoustic data stream; S3. Monitor the temporal energy mutation nodes of the spatially gated acoustic data stream and generate a global hardware interrupt trigger signal for inversely constraining the life boundary of visual computing. S4. Based on the absolute timestamp carried by the global hardware interrupt trigger signal, perform an irreversible temporal truncation and forced cutting operation on the visual image sequence that is rolled into the front-end hardware queue, and separate and extract the locked interaction frame sequence. S5. By integrating the locked interactive frame sequence and the preset duration voice command segments in the spatially gated acoustic data stream, a multi-dimensional target constraint tensor is constructed that includes a dual security confirmation base containing physical space coordinate pointing and structural morphological motion changes. S6. Analyze the multi-dimensional target constraint tensor to generate irreversible transferable terminal physical device linkage control execution instructions, and complete the mechanical motion drive process of the car window in the closed loop state.

[0021] In one embodiment of the present invention, step S1 includes the following steps: The system activates the in-vehicle environment surface monitoring sensor to collect real-time continuous spatial frame sequences. It then analyzes these sequences to obtain relative displacement density map interaction data that characterizes the physical occupancy of the user's hand. Continuous edge extreme points are extracted from the relative displacement density map interaction data. The system calculates the spatial mapping geometric relationship between these continuous edge extreme points and the fixed physical coordinate system of each independent window contained in the target vehicle, outputting a real-time iterative three-dimensional spatial direction vector. Based on a preset coordinate space hard-coded forced mapping rule, the three-dimensional spatial direction vector is subjected to dimensionality reduction and discretization processing to generate an eight-bit binary addressing mask that specifically points to the controlled window area of ​​the target.

[0022] Specifically, the first step of this method is executed by the central processing unit (CPU) of the in-vehicle interactive system. The goal is to acquire and parse the user's directional gestures, converting them into hardware-level directional signals for subsequent acoustic processing. Upon initiation of this step, the CPU first activates a wide-angle Time-of-Flight (ToF) sensor deployed on the top of the vehicle's interior. This sensor captures a continuous spatial frame sequence of the vehicle's three-dimensional space at a rate of 30 frames per second. The in-vehicle environmental surface monitoring sensor is preferably a Time-of-Flight (ToF) camera mounted above the rearview mirror base, with a working wavelength of 940nm, a spatial resolution of 640x480 pixels, and a frame rate between 30Hz and 60Hz. This frame rate range balances the real-time performance of gesture recognition with the computational load of the control system. Frame rates below 30Hz can cause motion blur or information loss in dynamic gesture features, such as instantaneous finger pushes and pulls, making displacement gradient calculation impossible. While frame rates above 60Hz offer higher temporal resolution, they lead to an exponential increase in point cloud data, increasing the communication burden on the CPU and data bus, and diminishing the marginal effect on improving the accuracy of macroscopic gesture recognition. Therefore, 30-60Hz is the preferred range for achieving a balance between low power consumption and high robustness. Each frame contains a set of three-dimensional coordinate point cloud data representing physical points on the scene surface.

[0023] The system performs real-time analysis on the acquired continuous spatial frame sequence, establishes a 3D voxel grid covering the interactive area, and filters out point cloud data of static objects such as seats and dashboards using a background subtraction algorithm, retaining only the point cloud data stream representing human body and hand movements. By comparing the occupancy changes of each voxel unit in two consecutive frames of point cloud data streams, the system generates real-time updated relative displacement density map interactive data representing the physical occupancy of the user's hand. In engineering implementation, the relative displacement density map interactive data is a series of dynamic 3D occupancy grid matrices, with each grid unit measuring 1cm x 1cm x 1cm, storing a value of 0 or 1, indicating whether the spatial unit is occupied by a dynamic object in the current frame.

[0024] The central processing unit performs edge topology analysis on the relative displacement density map interaction data, identifies moving voxel clusters representing the hand using a connected component labeling algorithm, and extracts continuous edge extrema points from them. It should be noted that continuous edge extrema points are defined as those in dynamically occupied voxel clusters whose velocity vector norm is greater than a dynamic velocity threshold. Furthermore, the geometric center of the isolated point or cluster of points that is spatially furthest from the center point of the driver's seat, and the dynamic speed threshold. The speed threshold is preferably set between 0.1 m / s and 0.3 m / s. This dynamic speed threshold is mainly used to filter out unintentional interference in the cabin environment, such as passive arm tremors caused by vehicle movement or slight hand drifts when the user is in a relaxed state. When the movement speed of the voxel cluster is below this threshold, the system classifies it as static or irrelevant background; only when the speed exceeds this physical lower limit is it activated as the starting point of an intentional interactive gesture.

[0025] In practice, the system calculates the centroid of one or more voxels in this voxel cluster that are farthest from the preset user shoulder reference coordinates, and uses this centroid as the endpoint of the pointing motion. The system retrieves pre-calibrated fixed physical coordinate systems of each independent window within the target vehicle from internal memory. These coordinate systems define the three-dimensional coordinates of the geometric center points of areas such as the left front window and right rear window. By calculating the vector relationship between continuous edge extreme points and the fixed physical coordinate systems of the independent windows—for example, constructing a direction vector using the coordinates of continuous edge extreme points and the sensor coordinate origin—the system outputs a real-time iterative three-dimensional spatial direction vector. .

[0026] The system applies a pre-defined coordinate space hard-coded forced mapping rule to the three-dimensional spatial direction vector. The process involves dividing the vehicle's interior space into several non-overlapping cone or pyramidal regions, each uniquely corresponding to a single window. The hard-coded coordinate space mapping rule is based on statistical data from over a hundred different vehicle model samples obtained through 3D cockpit scans. Assuming this implementation is applied to a standard five-seat sedan, the space is divided into four side window and one sunroof regions. The specific mapping rules are as follows: Azimuth Angle Within the interval [110°, 160°] and the pitch angle Mapped to the passenger-side window within the range of [-30°, 30°]; azimuth angle Within the range [-160°, -110°] and the pitch angle Mapped within the range [-30°, 30°] to the driver's side window, etc. Three-dimensional spatial direction vector. Convert to azimuth angle With pitch angle The specific conversion formula is as follows: In the above formulas for converting azimuth and elevation angles, , , These are the three-dimensional spatial direction vectors. In the vehicle-mounted Cartesian coordinate system, the unit of each of the three components is a meter (m). It is a two-parameter arctangent function, ensuring that the angle value falls within a complete 360° or Within the radian range; the output azimuth angle With pitch angle The unit is radians (rad), which can be converted to degrees (deg) for region matching.

[0027] After obtaining the aforementioned angle values, the system performs dimensionality reduction and discretization by determining the region to which these angle values ​​fall. Once an angle value matches a specific window region, the system immediately generates an eight-bit binary addressing mask that specifically points to the target controlled window region. It should be noted that the bits of the eight-bit binary addressing mask are defined as follows: from least significant bit to most significant bit, they correspond to the driver's side window, passenger side window, driver's side rear window, passenger side rear window, and sunroof, respectively. A 1 in the corresponding bit indicates the direction, while the remaining bits remain 0. If the three-dimensional spatial direction vector points to a non-window region or its magnitude is less than a preset activation threshold, a zero-based mask is generated. The final product of this step, the eight-bit binary addressing mask, will serve as the input for subsequent acoustic gating control. The activation threshold is typically set between 0.3 m and 0.6 m in physical space, depending on the cabin width of different vehicle models. The purpose of setting the activation threshold is to define an absolute "effective interaction trigger zone." Because ToF sensors have full cabin depth perception capabilities, without setting an activation threshold to constrain their movement, the actions of rear passengers or the driver's limbs extending away from the control panel area can easily lead to misjudgments. Setting an activation threshold ensures that only specific pointing actions occurring within a reasonable operating distance of the target window are captured.

[0028] For example, suppose an in-vehicle environment surface monitoring sensor captures a continuous spatial frame sequence. After parsing, the continuous edge extrema points extracted from the relative displacement density map interactive data have three-dimensional coordinates (-0.8m, 0.6m, -0.1m) in a coordinate system with the sensor as the origin. The system uses these coordinates to generate a three-dimensional spatial direction vector. The values ​​are (-0.8, 0.6, -0.1).

[0029] The system calculates its pointing azimuth angle. This translates to approximately 143.1°. Simultaneously, the pointing pitch angle was calculated. ,Right now This translates to approximately -5.7° in angle. The system matches this pair of angle values ​​(143.1°, -5.7°) with the hard-coded mandatory mapping rules in the coordinate space. Based on the aforementioned rules set for a standard five-seater sedan, this angle pair falls within the azimuth angle range [110°, 160°] and the pitch angle range [-30°, 30°]. This area is defined as pointing towards the passenger-side window. Therefore, the system ultimately generates an eight-bit binary addressing mask of 00000010 to define the acoustic receiving area.

[0030] In one embodiment of the present invention, step S2 includes the following steps: The microphone array baseband node located in the controlled space is activated to continuously capture the ambient audio pulse-code modulation stream at a preset upper limit frequency as the raw audio sampling stream. The raw audio sampling stream is then cyclically pushed into the acoustic ring buffer of the bypass architecture. The default omnidirectional digital signal reading gate of the acoustic ring buffer is physically blocked, cutting off the overflow path of the boundless background wind noise and reverberation noise to the central processing unit. An eight-bit binary addressing mask is used as an addressing and positioning pointer and overwritten into the memory access read address pin in the physically blocked state. The beamforming weight matrix dedicated channel of the beamforming weight matrix of the corresponding target controlled window area is forcibly partially opened, forcing the pickup beam to only capture the ambient sound of the physical location coordinates pointed to by the eight-bit binary addressing mask, and stably outputting an exclusive spatially selected acoustic data stream.

[0031] Specifically, after generating the eight-bit binary addressing mask in the aforementioned steps, the in-vehicle system's audio-visual entertainment controller or dedicated digital signal processing unit immediately initiates the directional acquisition process of the acoustic signal. First, this unit activates the microphone array pickup baseband nodes located in the B-pillars or headliner area of ​​the vehicle's interior. These are typically linear or circular arrays composed of four to eight digital MEMS (Micro-Electro-Mechanical Systems) microphones. The microphone array pickup baseband nodes are specifically composed of a four-unit linear digital MEMS microphone array with a microphone spacing of 4cm. The physical basis for setting the spacing to 4cm comes from the spatial sampling theorem. To avoid spatial aliasing during acoustic beamforming, the microphone spacing... Must meet The system aims to capture high-frequency environmental speech pulses, such as consonants and plosives, with an upper frequency limit of approximately 4 kHz. The speed of sound is calculated at 340 m / s, corresponding to a minimum wavelength. The wavelength is approximately 8.5cm. Therefore, the half-wavelength constraint is 4.25cm. Using a 4cm spacing not only strictly adheres to the Nyquist spatial sampling criterion, ensuring the formation of a highly directional, lobe-free beam for the target area, but also facilitates the compact structural deployment of the vehicle roof.

[0032] The microphone array's pickup baseband node is integrated into the vehicle's roof-mounted reading light control panel to achieve optimized acoustic coverage throughout the cabin. The array begins continuously capturing an ambient audio pulse-code modulation (PCM) stream with a preset upper frequency limit at a sampling rate of 48 kHz and a quantization depth of 24 bits. To achieve low-latency hardware processing, the captured PCM stream is not submitted to the central processing unit (CPU), but is instead forcibly and cyclically pushed into a dedicated 1024-byte bypass acoustic ring buffer using a first-in-first-out (FIFO) architecture. This acoustic ring buffer is a static random-access memory (SRAM) block integrated within the audio digital signal processor (ADSP). At a sampling rate of 48 kHz and a sampling depth of 24 bits (3 bytes per sample), its 1024-byte capacity can buffer approximately 7.1 ms of audio data.

[0033] To eliminate environmental noise interference and prepare for subsequent gating, the system writes a control word to a specific control register of the audio codec chip, forcibly blocking the default omnidirectional digital signal readout gate of the acoustic ring buffer. This creates a barrier at the hardware level, effectively cutting off the path for boundless background noise such as air conditioning noise, wind noise at high speeds, and excess reverberation noise inside the vehicle to spill over to the central processing unit, thus solving the problem of unnecessary computational resource consumption.

[0034] The eight-bit binary addressing mask generated in step S1 is overwritten into the read address pin mapping register of the audio memory access DMA controller, which is in a physically blocked state, via the system's internal bus in the form of a hardware address pointer. This operation utilizes the bit pattern of the eight-bit binary addressing mask to physically enable a dedicated channel of the beamforming weighting matrix that corresponds perfectly to the controlled window area pointed to by the mask. Once this channel is enabled, the DMA controller calls the specific weighting matrix bound to this channel to perform weighted summation processing on the raw signal from the microphone array, thereby forcing the pickup beam to form a narrow band in space pointing to the physical location coordinates specified by the eight-bit binary addressing mask.

[0035] Therefore, only ambient sounds from the pointed area can be effectively captured and written to the output register, and the system thus stably outputs a spatially gated acoustic data stream with high directional exclusivity for subsequent analysis.

[0036] At the underlying hardware implementation level, the physical blocking state is achieved by setting the enable position of the DMA channel to zero. Furthermore, the dedicated channel for the beamforming weight matrix actually points to a set of coefficients pre-stored in the firmware. These coefficients are calculated offline using beamforming algorithms, such as delay-and-sum, based on the geometry of the microphone array and the spatial positions of the four main windows. The total number of coefficients includes four sets of weight matrices corresponding to the four side windows.

[0037] For example, continuing from the previous steps, the central processing unit obtains an eight-bit binary addressing mask of 00000010. While the microphone array continuously captures the ambient audio pulse-code modulation stream at a preset upper frequency and pushes it into an acoustic ring buffer, the default read port of this buffer is closed. The system then writes the mask value 00000010 into the address register of the audio DMA controller. Since the second bit of this binary value is 1 and the rest are 0, the DMA controller interprets it as enabling the read channel with index binary 1.

[0038] According to the system preset, the channel with index 0 corresponds to the driver's side window, and the channel with index 1 uniquely corresponds to the passenger side window. Therefore, this operation triggers the DMA controller to load a specific beamforming weight matrix associated with the passenger side window direction from the firmware and apply it to the signal processing flow of the microphone array. The system ignores the sound from the driver and rear passenger directions, and only outputs the beamformed audio signal stream focused on the passenger side window direction as a spatially gated acoustic data stream to the subsequent processing unit, awaiting further temporal energy analysis.

[0039] In one embodiment of the present invention, step S3 includes the following steps: The spatially gated acoustic data stream is guided unidirectionally into the micro digital signal processor, continuously initiating transient physical edge scanning calculations in a low-power cycle; the high-frequency energy mutation slope steep bias value of the spatially gated acoustic data stream within a specified microsecond-level time sliding channel is extracted through a differential filtering algorithm; when the high-frequency energy mutation slope steep bias value exceeds the fixed absolute onset voltage threshold, the transient burst of the initial speech pronunciation is intercepted at the physical level, and a global hardware interrupt trigger signal with an absolute timestamp is emitted.

[0040] Specifically, after stabilizing the output space-selected acoustic data stream in step S2, the system unidirectionally directs it to a dedicated micro digital signal processor (DSP). This DSP is preferably an ARM Cortex-M4 core microcontroller unit operating at 96MHz, equipped with a single-cycle multiply-accumulate instruction set to support efficient digital signal processing. This data stream is fed into the processor via an I²S (Inter-ICSound) serial bus to avoid consuming main CPU resources. Once the DSP receives the data, it continuously initiates low-power cycle transient physical edge scanning calculations. This calculation does not decode the audio content but rather rapidly scans the physical characteristics of the audio waveform at extremely high frequencies.

[0041] During the scanning calculation, the miniature digital signal processor first applies a high-pass digital filter to the input spatially gated acoustic data stream to filter out low-frequency components below 2kHz, thus highlighting high-frequency features such as consonants and plosives in speech. The processor then uses a differential filtering algorithm to extract the steep bias value of the high-frequency energy abrupt change slope within a specified microsecond-level time-sliding channel of the spatially gated acoustic data stream in real time. Specifically, the processor maintains a length of... A sliding window that moves in steps. Move along the data stream. It should be noted that the parameter for specifying the microsecond-level time-sliding channel can be set to the window length. For 128 sampling points, the sliding step size With 16 sampling points and a sampling rate of 48kHz, L=128 corresponds to a time window of approximately 2.67ms. This width not only fully encompasses the transient energy burst at the moment of speech pronunciation, typically lasting 1-5ms, but also avoids the smoothing dilution of energy features caused by an excessively long window. A step size K=16 corresponds to a sliding step of 0.33ms, implying a high overlap between adjacent windows. This high overlap ensures that the system can track the highest point of the energy mutation slope at microsecond-level high frequencies, providing a physical-level guarantee for obtaining the absolute timestamp; at a sampling rate of 48kHz, this corresponds to a time window width of 2.67ms and a time step interval... It is approximately 333 μs.

[0042] At each step position, the processor calculates the sum of the squares of all sample amplitudes within the window, which is taken as the short-time energy at that time point. Then, the difference between the current window energy and the previous window energy is calculated and divided by the window step time interval. Thus, the steep bias value of the high-frequency energy change slope, which characterizes the rate of energy change, is obtained. The calculation formula is as follows: Among them, the aforementioned high-frequency energy mutation slope steep bias value The calculation formula is given, and its unit can be converted to volts squared per second (V² / s). It is the first The short-term energy of a window, It is the first The window, that is, the short-term energy of the previous sliding window; The first one after high-pass filtering and normalization The digital amplitude of each audio sampling point is normalized to ±1.0 at full scale. To facilitate unit uniformity, this paper maps the normalized full-scale digital quantity to the nominal unit volt (V). It is the length of the sliding window, which is the dimensionless number of sample points. is the step size of the window sliding, and is the dimensionless number of sample points. It represents the time interval between each window slide, equal to the step size. Divide by audio sampling rate The unit is seconds (s). Furthermore, in the formula... The current time-series index number of the sliding window, which is a positive integer; This is the sequence number of the audio sample point in the data stream, and it is a positive integer.

[0043] Finally, the system uses the steep bias value of the high-frequency energy mutation slope calculated in real time in a loop judgment. Compared to the fixed absolute trigger voltage threshold stored in the processor firmware A comparison was made. It should be noted that the fixed absolute attack voltage threshold... This is an empirical constant established through statistical analysis of the onset slope of the initial sounds of a large number of standard Mandarin speech samples, such as the commands "open," "close," and "stop." It is stored in the processor's read-only memory to prevent tampering. This threshold is set between 3 and 5 times the maximum transient energy slope of the background steady-state noise. The energy rise curve of environmental wind noise or tire noise is gradually and continuously changing, while the instantaneous impact of the vocal cords and airflow at the onset of human vocalization exhibits a steep step characteristic. The introduction of this strict absolute onset voltage threshold eliminates the interference of steady-state environmental noise on system wake-up.

[0044] Once the high-frequency energy mutation slope steepness bias value is determined to exceed the absolute start-up voltage threshold, in this extremely brief instant, the micro digital signal processor immediately obtains the current high-precision time from the system-level hardware clock source and uses it as an absolute timestamp. This absolute timestamp comes from a 64-bit hardware counter synchronized with the system master clock, with an accuracy of 1 μs. At the same time, a high-level signal is thrown to the system bus controller through a dedicated GPIO (General Purpose Input / Output) pin. This signal is the global hardware interrupt trigger signal, and its priority is set to the highest hardware interrupt request (IRQ) to ensure that the vision processing subsystem can respond immediately.

[0045] For example, following the aforementioned steps, the miniature digital signal processor receives a spatially gated acoustic data stream focused on the direction of the passenger-side window. Assuming the interior is relatively quiet before system time 1578342100 μs, the calculated high-frequency energy abrupt change slope is a steep bias value. All remained below the absolute phonation voltage threshold for curing. For example, it can be set to 50,000 V² / s. Around 1578342100 μs, the user in the passenger seat issues the "open window" command. The initial consonant / k / of the word "open" generates a violent energy surge. The short-term energy is calculated for the m-th time window covering the origin of this surge. It rises sharply to 18.5 V², while the energy of the previous window... It is only 0.5 V².

[0046] The processor calculates V² / s. This value far exceeds the absolute start-up voltage threshold of 50,000 V² / s. Therefore, the micro digital signal processor immediately latches the current system hardware time of 1578342100 μs and uses this timestamp as the absolute timestamp. At the same time, it sends a high-level signal to the main CPU through its IRQ pin, forming a global hardware interrupt trigger signal.

[0047] In one embodiment of the present invention, step S4 includes the following steps: The system captures the global hardware interrupt trigger signal transmitted to the system control bus through an independent channel, and then inverts and elevates the global hardware interrupt trigger signal to the highest-level interrupt register of the visual capture subsystem for response. The system uses the absolute timestamp carried by the global hardware interrupt trigger signal to forcibly freeze the unlimited scrolling overwrite calculation lifecycle of the entire original physical image in the visual capture subsystem's queue buffer storage pool. Using the absolute timestamp as the truncation center axis instruction pointer, the system performs bidirectional hard cropping and discards background frames outside the preset time window without related actions, retaining only a set of image frames immediately before and after the absolute timestamp in a preset number of units, and assembling them to construct a locked interactive frame sequence with exclusive attributes.

[0048] Specifically, after the global hardware interrupt trigger signal is emitted in step S3, it is immediately transmitted to the system control bus inside the central processing unit via a dedicated hardware interrupt line. Upon receiving this high-priority independent channel signal, the system interrupt controller performs interrupt dispatch, elevating the global hardware interrupt trigger signal to the highest-level interrupt register of the vision capture subsystem for mandatory response. The highest-level interrupt register of the vision capture subsystem is a programmable register within the system interrupt controller, with its priority set to 0, representing the highest interrupt level in the system. This operation suspends all currently executing low-priority image processing tasks.

[0049] Once the system's interrupt service routine is activated, it first reads the absolute timestamp generated in step S3 from the data carried by the global hardware interrupt trigger signal or the associated register. Simultaneously, the interrupt service routine issues a pause command to the DMA controller responsible for managing the Time-of-Flight (ToF) sensor data stream, thereby forcibly freezing the lifetime of the unlimited rolling overwrite computation of the full original physical frames in the visual capture subsystem's queue buffer storage pool. This storage pool is a physically circular buffer that continuously stores a sequence of consecutive spatial frames from the most recent period. Specifically, the queue buffer storage pool is a 32 MB circular buffer configured in DDR SDRAM, capable of rolling over approximately 60 consecutive spatial frame sequences generated by the ToF sensor within the last two seconds.

[0050] After the buffer write is frozen, the central processing unit (CPU) uses the absolute timestamp passed from S3 as the truncation center axis instruction pointer to traverse and index the image frames in the buffer. The processor compares the capture timestamp of each frame with the absolute timestamp and locates the frame with the closest timestamp as the center keyframe.

[0051] The processor performs a bidirectional hard crop based on the frame count parameters embedded in the system configuration, discarding background frames outside the preset time window that do not have relevant actions—that is, all other frames that do not meet the time window requirements. After cropping, the system retains only a set of image frames immediately preceding and following a preset number of timestamps; for example, pre-setting to capture 5 frames before and 5 frames after the central keyframe. These retained image frames are reorganized in memory, forcibly combined to construct a data structure with temporal exclusivity, i.e., locking the interactive frame sequence, and awaiting subsequent multimodal fusion processing. After this process is completed, the interrupt service routine sends a restore command to the DMA controller, and the circular buffer resumes normal rolling overwriting.

[0052] The truncation center axis command pointer is not a physical pointer, but a logical time reference used as a matching target in the subsequent search algorithm. Furthermore, based on the preset unit of capturing 5 frames before and after, the final generated lock interaction frame sequence will contain 11 consecutive frames of image data, with a total duration of approximately 367 ms. This duration is sufficient at the system level to cover the entire process of a short gesture.

[0053] For example, continuing the previous steps, after the system captures a global hardware interrupt trigger signal carrying an absolute timestamp of 1578342100 μs, it immediately suspends the write operation to the image circular buffer. At this time, the buffer stores approximately 60 frames of images from timestamps of 1577142000 μs to 1579142000 μs. The CPU searches the buffer using 1578342100 μs as a reference and finds that the capture timestamp of frame 37 is 1578342088 μs, which is closest to the target timestamp. Therefore, frame 37 is identified as the center keyframe.

[0054] Based on the preset rule of capturing 5 frames before and after, the system copies 11 frames of image data, from index 32 (37-5) to 42 (37+5), from the circular buffer to a new contiguous memory block. These 11 frames together constitute the lock interaction frame sequence, which contains the complete subtle pointing or pushing movements of the hand before and after the user utters the "open" sound. Afterward, the write operation to the circular buffer is resumed, and the remaining 49 frames of image data are considered irrelevant information and will be naturally overwritten in subsequent data rollovers.

[0055] In one embodiment of the present invention, step S5 includes the following steps: By penetrating and reading the locked interaction frame sequence, the point with the smallest depth value in the dynamic human point cloud cluster of the locked interaction frame sequence is located as the user's fingertip point. The pixel displacement gradient components of the user's fingertip point during the existence of the locked interaction frame sequence are analyzed to confirm the local push-pull motion vector that triggers the action feedback in the controlled window area of ​​the target. Simultaneously, audio feature island segments connected immediately after the absolute timestamp in the spatially gated acoustic data stream are extracted as short trigger confirmation recordings, and the corresponding audio fundamental frequency fluctuation envelope layer features are extracted. The local push-pull motion vector, which serves as the structural driving verification condition, and the audio fundamental frequency fluctuation envelope layer features, which serve as the acoustic secondary verification condition, are merged and combined with the target window physical displacement identifier obtained by converting the eight-bit binary addressing mask. Through matrix orthogonal filling and splicing, a multi-dimensional target constraint tensor is generated to wake up the underlying electromechanical assembly switch of the device.

[0056] Specifically, after constructing the locked interaction frame sequence in step S4, the central processing unit initiates deep fusion processing of visual and auditory data streams in parallel. First, the system reads the locked interaction frame sequence containing 11 frames of image data and rapidly processes the 3D point cloud data of each frame to analyze the pixel displacement gradient components of the user's fingertip point during the duration of the locked interaction frame sequence. Specifically, in each frame's dynamic human point cloud cluster, the system locates the point with the smallest depth value as the user's fingertip point. This location process can employ a search algorithm based on the minimum depth value; that is, in the depth map returned by the ToF sensor, after excluding the static background, it searches for the pixel point with the smallest z-axis coordinate value. The system records the position coordinates of this point in the vehicle's 3D coordinate system and calculates the 3D displacement vector by comparing the 3D coordinate changes of the fingertip point in the first and last frames of the locked interaction frame sequence. The system retrieves the plane normal vector of the target window and projects this 3D displacement vector onto the normal vector direction to confirm the local push-pull motion vector pointing towards the target controlled window area that triggers the action feedback.

[0057] Simultaneously with visual processing, the system extracts isolated audio feature segments from the spatially gated acoustic data stream, closely following the absolute timestamp, as brief trigger confirmation recordings. The starting point of this segment is the absolute timestamp, and its duration is preset to 400 ms. This duration setting is based on statistical analysis of the duration of commonly used short speech commands controlled by single / disyllabic words in natural language processing, such as "on" and "stop." The typical duration for an adult to pronounce a sudden syllable and complete its final sound is between 200 ms and 350 ms. The preset duration of 400 ms ensures the capture of a complete, untruncated speech fundamental frequency envelope waveform for 40-dimensional feature extraction, while also removing redundant background noise after the command ends, reducing the tensor dimension of subsequent matrix operations, and improving processing efficiency.

[0058] The system then extracts the corresponding audio fundamental frequency fluctuation envelope features from the recording segment. First, the audio data is pre-emphasized and segmented into frames. Second, the short-time energy of each frame is calculated, forming an energy envelope sequence. Third, the fundamental frequency (F0) of each frame is calculated using the autocorrelation function method, forming a fundamental frequency fluctuation sequence. Finally, the energy envelope sequence and the fundamental frequency fluctuation sequence are downsampled and concatenated to form a fixed-dimensional feature vector, which is the audio fundamental frequency fluctuation envelope feature. Specifically, the audio fundamental frequency fluctuation envelope feature is a composite feature vector, where the first 20 elements represent the average energy value every 20 ms interval, and the last 20 elements represent the corresponding average fundamental frequency value, thus forming a 40-dimensional vector.

[0059] The system merges the analysis results from the first two steps to construct a multidimensional target constraint tensor. The system generates the target window identity using a matrix orthogonal filling and splicing method, combining the local push-pull motion vector (serving as the structural driving verification condition), the audio fundamental frequency fluctuation envelope layer features (serving as the acoustic secondary verification condition), and the physical displacement identifier obtained from step S1. It should be understood that the physical displacement identifier can be directly derived from the eight-bit binary addressing mask generated in step S1 and converted into a 5-dimensional one-hot encoding using a lookup table; for example, 00000010 (passenger side window) corresponds to [0, 1, 0, 0, 0].

[0060] The specific operation involves creating a one-dimensional vector of a predetermined length, and then sequentially filling different positions of the one-dimensional vector with the three components of the local push-pull motion vector, all elements of the feature vector of the audio fundamental frequency fluctuation envelope layer, and the one-hot encoding corresponding to the physical displacement identifier. This vector is the final multi-dimensional target constraint tensor used to wake up the underlying electromechanical assembly switch of the device, and is then passed to the next processing step.

[0061] When confirming the action pointing towards the controlled window area, the validity criterion for the local push-pull motion vector is that the projection component of the vector onto the window normal vector must be positive, and its magnitude must be greater than a preset action threshold, such as 3cm. Setting the action threshold to 3cm constitutes a second layer of safety confirmation in the visual spatial dimension. Within a time slice of approximately 367 ms (11 frames), the physical travel of a normal finger tap or push-pull operation is generally between 5cm and -10cm; while the inertial displacement of the human body due to vehicle vibration is usually less than 2cm. Therefore, 3cm is set as a rigid physical boundary condition to distinguish between "active structural morphological changes" and "passive environmental disturbances," preventing false triggering caused solely by static pointing and mispronunciation. Furthermore, since it integrates the motion vector with 3 components, the 40-dimensional acoustic feature vector, and the 5-dimensional physical displacement identifier, the multi-dimensional target constraint tensor in this embodiment is manifested as a one-dimensional tensor with dimensions of 3+40+5=48.

[0062] For example, following the aforementioned steps, the central processing unit begins processing a lock-interaction frame sequence containing 11 frames of images and a spatially gated acoustic data stream starting at 1578342100 μs. The system analyzes the first frame of the lock-interaction frame sequence and locates the coordinates of the user's fingertip. =(-0.75, 0.58, -0.15) m. In the last frame, the coordinates of this point become... =(-0.70, 0.59, -0.25) m. Based on this, the three-dimensional displacement vector is calculated to be (0.05, 0.01, -0.10) m.

[0063] Assuming the normal vector of the passenger-side window approximately points inwards, with a direction vector of (0, 0, -1), the projection magnitude is greater than zero. The system confirms this as a push-to-pull motion towards the window and uses this displacement vector as the local push-pull motion vector. Simultaneously, the system extracts a 400 ms audio segment from 1578342100 μs to 1578742100 μs and extracts its 40-dimensional fundamental frequency fluctuation envelope feature, assuming its value is [E1, E2, ..., E...]. 20 F1, F2…, F 20 Based on the eight-bit binary addressing mask 00000010 from step S1, the system converts the physical displacement identifier to [0, 1, 0, 0, 0]. Finally, the system concatenates these three sets of data to generate a 48-dimensional multidimensional target constraint tensor, whose content is [0.05, 0.01, -0.10, E1, E2…, E…]. 20 F1, F2…, F 20 [0, 1, 0, 0, 0]. This tensor is then sent to the edge gateway node for decoding.

[0064] In one embodiment of the present invention, step S6 includes the following steps: The multidimensional target constraint tensor is fed into the edge gateway node for lookup-level lightweight decoding and matching without semantic computation, so as to obtain a deterministic execution password carrying the unique physical entity identification code of the target window and the target mechanical execution logic sequence. Translate the deterministic execution command to generate a DC pulse duty cycle wide waveform for the drive terminal anti-pinch lifting module and a synchronously bound motor forward and reverse rotation level switching signal; The DC pulse duty cycle wide waveform and the synchronously bound motor forward and reverse rotation level switching signal are sent to the base servo motor that is precisely corresponding to the unique physical entity identification code of the target window. The drive gear produces a constant power physical displacement, thus completing the mechanical action drive process of the window to achieve the expected interactive action result.

[0065] Specifically, after constructing and issuing the multi-dimensional target constraint tensor in step S5, the tensor is transmitted to the vehicle's domain controller or a dedicated edge gateway node for final parsing and execution. It should be noted that the edge gateway node is a hardware module integrating a microcontroller and a CAN bus transceiver.

[0066] After receiving the multidimensional target constraint tensor that has completed spatiotemporal height synchronization verification, the edge gateway node does not perform complex semantic understanding or machine learning inference. Instead, it initiates a lightweight decoding and matching process that avoids semantic computation by looking up a table. The system uses the physical displacement identifier portion of the multidimensional target constraint tensor as an index in the node's local non-volatile memory to look up a preset execution instruction mapping table. The execution instruction mapping table stored in the firmware is a static data structure that maps the 5-dimensional one-hot encoding to the physical addresses and control instruction sets of the motor control units for the four windows and sunroof. For example, [0, 1, 0, 0, 0] is mapped to the passenger-side window lift controller with physical address 0x1A2, along with the corresponding instruction sequence {direction:UP, duration:5000ms}. Through this index, the system obtains a deterministic execution password carrying the unique physical entity identifier code of the target window and the target mechanical execution logic sequence. The physical entity identifier code is the CANID of the target window in the vehicle network, for example, 0x1A2.

[0067] Upon receiving the deterministic execution command, the microcontroller within the edge gateway node translates it. The translation process converts the logical instructions in the deterministic execution command into underlying electrical signal control parameters. For example, instructions containing actions such as "rise" and "fully open" are translated into a series of DC pulse duty cycle wide waveforms driving the anti-pinch lifting module of the driving terminal, along with synchronously bound motor forward / reverse rotation level switching signals. The DC pulse duty cycle wide waveforms determine the speed and force of the motor rotation, while the motor forward / reverse rotation level switching signals control whether the motor rotates forward or backward, thus raising or lowering the window.

[0068] As the final link in the execution, the edge gateway node translates the generated DC pulse duty cycle wide waveform and synchronously binds the motor forward and reverse rotation steering level switching signal, and sends it down to the base servo motor corresponding to the unique physical entity identifier code of the target window specified in the deterministic execution password via the CAN bus of the controller area network or a dedicated line.

[0069] After receiving the electrical signal, the servo motor's internal drive circuit starts working, driving the gear mechanism to produce a constant power physical displacement, thereby controlling the movement of the car window glass. The completion of this series of actions signifies that the expected interactive action result, namely the mechanical action drive process of the car window, has been fully verified in a closed loop.

[0070] In the implementation of the underlying control signals, the DC pulse duty cycle wide waveform is the PWM signal, with a fixed frequency of 1 kHz. This setting is based on engineering optimization for the inductive load characteristics of the window servo motor. This frequency ensures a continuous and smooth current flowing through the motor windings, resulting in a stable and delicate window lifting torque output, avoiding mechanical resonance and jerking caused by low-frequency modulation. Simultaneously, it effectively controls the switching losses of the H-bridge power transistors in the anti-pinch lifting module, preventing overheating of the hardware due to high-frequency switching in the closed-loop control process. The duty cycle is adjustable from 10% to 90% to control the motor speed. Furthermore, the motor's forward and reverse rotation level switching signal specifically controls two logic level signals in the H-bridge drive circuit. For example, a combination of [low level, high level] represents forward rotation (rising), while a combination of [high level, low level] represents reverse rotation (falling).

[0071] For example, following the aforementioned steps, the edge gateway node receives the 48-dimensional multidimensional target constraint tensor generated above. It first parses the last five elements of this tensor [0, 1, 0, 0, 0]. Using these as an index, the system finds the corresponding entry in its local mapping table. This entry specifies the unique physical entity identifier code of the target window as 0x1A2, and the target mechanical execution logic sequence as {action: parameter:FULL}.

[0072] The gateway's microcontroller translates this logic sequence. It generates a level signal to control the motor's direction, assuming [high, low] for descent, and thus generates this signal. Simultaneously, it generates a 1 kHz DC pulse wide-duty-cycle waveform with an 80% duty cycle, which lasts for 5 seconds to ensure the window is fully open. The system packages the data frame with CANID 0x1A2, containing the direction control signal [high, low] and PWM parameters {duty_cycle:80%, duration:5000ms}, and sends it via the CAN bus. Upon receiving this data frame, the passenger-side window's lift motor module immediately drives the motor to descend, stopping after 5 seconds, thus completing the user's gesture and voice-initiated window-opening interaction.

[0073] See appendix Figure 2 The present invention also proposes a non-contact gesture and voice fusion window control interaction system, comprising the following modules: The spatial pointing mask generation module acquires three-dimensional spatial voxel data inside the vehicle and constructs a three-dimensional spatial direction vector to generate an eight-bit binary addressing mask for defining the acoustic receiving area. The acoustic data gating module calls an eight-bit binary addressing mask to cover the memory access channel port of the acoustic controller, performs physical-level region blocking on the preset raw audio sampling stream, and extracts the spatially gating acoustic data stream. The voice initiation monitoring module monitors the temporal energy mutation nodes of the spatially gated acoustic data stream and generates a global hardware interrupt trigger signal for inversely constraining the life boundary of visual computing. The visual frame locking module performs an irreversible temporal truncation and forced cutting operation on the visual image sequence that is continuously stored in the front-end hardware queue based on the absolute timestamp carried by the global hardware interrupt trigger signal, and separates and extracts the locking interaction frame sequence. The multimodal fusion tensor construction module integrates the locked interactive frame sequence and the preset duration voice command segments in the spatially gating acoustic data stream to construct a multidimensional target constraint tensor that includes a dual security confirmation base of physical space coordinate orientation and structural morphological motion change. The control command decoding and execution module parses the multi-dimensional target constraint tensor to generate irreversible transferable terminal physical device linkage control execution commands, completing the mechanical motion drive process of the car window in a closed-loop state.

[0074] Each of the modules can be implemented in whole or in part through software, hardware, or a combination thereof. It supports hardware embedded in or independent of the processor in the computer device, and also supports software stored in the memory of the computer device, so that the processor can call and execute the operations corresponding to each of the above modules.

[0075] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for non-contact gesture and voice fusion window control interaction, characterized in that, Includes the following steps: S1. Acquire three-dimensional spatial voxel data inside the vehicle and construct a three-dimensional spatial direction vector to generate an eight-bit binary addressing mask for limiting the acoustic receiving area. S2. Call an eight-bit binary addressing mask to cover the memory access channel port of the acoustic controller, perform physical-level region blocking on the preset raw audio sampling stream, and extract the spatially gated acoustic data stream; S3. Monitor the temporal energy mutation nodes of the spatially gated acoustic data stream and generate a global hardware interrupt trigger signal for inversely constraining the life boundary of visual computing. S4. Based on the absolute timestamp carried by the global hardware interrupt trigger signal, perform an irreversible temporal truncation and forced cutting operation on the visual image sequence that is rolled into the front-end hardware queue, and separate and extract the locked interaction frame sequence. S5. By integrating the locked interactive frame sequence and the preset duration voice command segments in the spatially gated acoustic data stream, a multi-dimensional target constraint tensor is constructed that includes a dual security confirmation base containing physical space coordinate pointing and structural morphological motion changes. S6. Analyze the multi-dimensional target constraint tensor to generate irreversible transferable terminal physical device linkage control execution instructions, and complete the mechanical motion drive process of the car window in the closed loop state. 2.The non-contact gesture and voice fusion window control interaction method of claim 1, wherein, Generating an 8-bit binary addressing mask to define the acoustic receiving area includes the following steps: Activate the in-vehicle environment surface monitoring sensor to collect real-time continuous spatial frame sequences, and analyze the continuous spatial frame sequences to obtain interactive data of relative displacement density map characterizing the physical occupation of the user's hand; Extract continuous edge extreme points from the relative displacement density map interaction data, calculate the spatial mapping geometric relationship between the continuous edge extreme points and the fixed physical coordinate system of each independent window contained in the target vehicle, and output the real-time iterative three-dimensional spatial direction vector. Based on the preset coordinate space hard-coded forced mapping rules, the three-dimensional spatial direction vector is reduced and discretized to generate an eight-bit binary addressing mask that specifically points to the controlled window area of ​​the target. 3.The non-contact gesture and voice fusion window control interaction method of claim 2, wherein, Outputting the real-time iterative 3D spatial direction vector includes the following steps: Edge topology analysis was performed on the relative displacement density map interaction data, and the moving voxel clusters representing the hand were identified by the connected component labeling algorithm. Calculate the center of mass of the moving voxel cluster that is farthest from the preset user shoulder reference coordinates as the continuous edge extreme point; Extract the three-dimensional coordinates of the geometric center point corresponding to the fixed physical coordinate system of each independent window, calculate the direction vector between the three-dimensional coordinates of the continuous edge extreme points and the geometric center point, and generate a three-dimensional spatial direction vector.

4. The method of claim 1, wherein, Extracting the spatially gated acoustic data stream includes the following steps: The microphone array pickup baseband node located in the controlled space is activated to continuously capture the ambient audio pulse code modulation stream with a preset upper limit frequency as the raw audio sampling stream, and the raw audio sampling stream is cyclically pushed into the acoustic ring buffer of the bypass architecture. The default omnidirectional digital signal reading gate of the clamping acoustic ring buffer is in a physical blocking state, cutting off the overflow path of boundless background wind noise and reverberation noise to the central processing unit. An eight-bit binary addressing mask is used as an addressing and positioning pointer and overwritten into the memory access read address pin which is in a physically blocked state. This forces the beamforming weight matrix dedicated channel of the corresponding target controlled window area to be partially turned on, so that the pickup beam can only capture the ambient sound of the physical location coordinates pointed to by the eight-bit binary addressing mask and output an exclusive spatially selected acoustic data stream.

5. The non-contact gesture and voice integrated window control interaction method of claim 1, wherein, Generating a global hardware interrupt trigger signal for inversely constraining the life boundary of visual computing includes the following steps: The spatially gated acoustic data stream is guided unidirectionally into the micro digital signal processor, continuously initiating transient physical edge scanning calculations in low-power cycles; The steep offset value of the high-frequency energy mutation slope in a specified microsecond-level time-sliding channel is extracted by a differential filtering algorithm; When the steep bias value of the high-frequency energy mutation slope exceeds the fixed absolute onset voltage threshold, the transient surge of the initial speech is intercepted at the physical level, and a global hardware interrupt trigger signal with an absolute timestamp is thrown outward.

6. The non-contact gesture and voice fusion window control interaction method according to claim 5, characterized in that, The high-frequency energy abrupt change slope steep offset value of the spatially gated acoustic data stream within a specified microsecond-level time-sliding channel is extracted using a differential filtering algorithm, including the following steps: High-pass digital filters are applied to the spatially gated acoustic data stream to remove low-frequency components; Calculate the short-time energy of the current window and the short-time energy of the previous window within a specified microsecond-level sliding channel; Calculate the difference between the short-time energy of the current window and the short-time energy of the previous window, and divide the difference by the window stepping time interval to generate a steep bias value for the high-frequency energy mutation slope.

7. The non-contact gesture and voice fusion window control interaction method according to claim 1, characterized in that, Separating and extracting the locked interaction frame sequence includes the following steps: The global hardware interrupt trigger signal transmitted to the system control bus is captured through an independent channel, and the global hardware interrupt trigger signal is inverted and promoted to the highest level interrupt register of the vision capture subsystem for response. The absolute timestamp carried by the global hardware interrupt trigger signal is used to forcibly freeze the lifecycle of the unlimited scrolling overwrite calculation of the full original physical images in the queue buffer storage pool of the visual capture subsystem; Using the absolute timestamp as the truncation center axis instruction pointer, the background frame outside the preset time window without related actions is hard-trimmed in both directions and discarded. Only the set of image frames that are immediately before and after the absolute timestamp are retained and combined to form a locked interactive frame sequence with exclusive properties.

8. The non-contact gesture and voice fusion window control interaction method according to claim 1, characterized in that, Constructing a multidimensional target constraint tensor that includes both physical space coordinate orientation and structural morphological motion changes for safety verification of the base includes the following steps: By penetrating and reading the locked interaction frame sequence, the point with the smallest depth value in the dynamic human point cloud cluster of the locked interaction frame sequence is located as the user's fingertip point. The pixel displacement gradient component of the user's fingertip point during the existence of the locked interaction frame sequence is analyzed to confirm the local push-pull motion vector that triggers action feedback in the controlled window area of ​​the target. Simultaneously extract audio feature island segments that immediately follow the absolute timestamp in the spatially gated acoustic data stream as short trigger confirmation recordings, and extract the corresponding audio fundamental frequency fluctuation envelope layer features; The local push-pull motion vector, which serves as the structural driving verification condition, and the audio fundamental frequency fluctuation envelope layer feature, which serves as the acoustic secondary verification condition, are merged and combined with the target window physical displacement identifier obtained by converting the eight-bit binary addressing mask. Through matrix orthogonal filling and splicing, a multi-dimensional target constraint tensor is generated to wake up the underlying electromechanical assembly switch of the device.

9. A non-contact gesture and voice fusion window control interaction method according to claim 1, characterized in that, The complete closed-loop mechanical movement drive process for the vehicle window includes the following steps: The multidimensional target constraint tensor is fed into the edge gateway node for lookup-level lightweight decoding and matching without semantic computation, so as to obtain a deterministic execution password carrying the unique physical entity identification code of the target window and the target mechanical execution logic sequence. Translate the deterministic execution command to generate a DC pulse duty cycle wide waveform for the drive terminal anti-pinch lifting module and a synchronously bound motor forward and reverse rotation level switching signal; The DC pulse duty cycle wide waveform and the synchronously bound motor forward and reverse rotation level switching signal are sent to the base servo motor that is precisely corresponding to the unique physical entity identification code of the target window. The drive gear produces a constant power physical displacement, thus completing the mechanical action drive process of the window to achieve the expected interactive action result.

10. A non-contact gesture and voice integrated window control interaction system, characterized in that, Includes the following modules: The spatial pointing mask generation module acquires three-dimensional spatial voxel data inside the vehicle and constructs a three-dimensional spatial direction vector to generate an eight-bit binary addressing mask for defining the acoustic receiving area. The acoustic data gating module calls an eight-bit binary addressing mask to cover the memory access channel port of the acoustic controller, performs physical-level region blocking on the preset raw audio sampling stream, and extracts the spatially gating acoustic data stream. The voice initiation monitoring module monitors the temporal energy mutation nodes of the spatially gated acoustic data stream and generates a global hardware interrupt trigger signal for inversely constraining the life boundary of visual computing. The visual frame locking module performs an irreversible temporal truncation and forced cutting operation on the visual image sequence that is continuously stored in the front-end hardware queue based on the absolute timestamp carried by the global hardware interrupt trigger signal, and separates and extracts the locking interaction frame sequence. The multimodal fusion tensor construction module integrates the locked interactive frame sequence and the preset duration voice command segments in the spatially gating acoustic data stream to construct a multidimensional target constraint tensor that includes a dual security confirmation base of physical space coordinate orientation and structural morphological motion change. The control command decoding and execution module parses the multi-dimensional target constraint tensor to generate irreversible transferable terminal physical device linkage control execution commands, completing the mechanical motion drive process of the car window in a closed-loop state.