Juicer selling system and method based on AI vision and voice interaction

CN122676594APending Publication Date: 2026-09-01HUIDEBAO INTELLIGENT MANUFACTURING (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610849809.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-12
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0005]本发明解决的技术问题在于,现有的自助售卖设备在嘈杂环境下,仅依赖音频拾音技术难以将目标用户的语音与背景噪声或非交互人员的语音有效分离,易引发设备误触发或执行错误指令

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122676594A_ABST
    Figure CN122676594A_ABST
Patent Text Reader

Abstract

This invention discloses a juicer vending system and method based on visual and voice interaction, belonging to the field of human-computer interaction technology. Addressing the issues of voice interaction in public settings being susceptible to environmental noise and interference from others leading to false triggering, and the lack of proactive customer acquisition and network outage tolerance in existing devices, this invention constructs a time-window synchronized audio and video stream, extracts the audio energy envelope sequence, and simultaneously utilizes the target's neck and nose tip nodes to construct an anatomical principal axis. This axis is then subjected to dot product projection and non-negative truncation operations with the two-dimensional dense optical flow vector of the target's mandibular region to generate a visual motion purification envelope sequence. Subsequently, a dynamic time warping algorithm is used to calculate the coherence score between the audio and visual envelope sequences. Based on this score, the actual vocal target is located, and the microphone array is guided to perform beamforming noise reduction. The subsequent command issuance system further integrates distance perception to trigger proactive greetings and dynamically generates marketing and sales scripts based on environmental parameters. Finally, a complete business loop including order placement, payment, beverage preparation, and cup retrieval reminders is executed. An edge network outage fallback module is also built in to ensure offline availability. This invention achieves cross-modal binding between facial vocalization actions and sound field energy fluctuations, effectively suppressing intent recognition interference in complex public acoustic environments and improving the order conversion rate and operational reliability of unmanned equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, specifically to a juice vending system and method based on AI vision and voice interaction. Background Technology

[0002] With the increasing prevalence of unmanned retail terminals, self-service juice machines and other similar devices are gradually being deployed in various commercial venues. To provide more convenient contactless services, voice recognition-based interaction technology is widely used in the product selection and control processes of these self-service devices. In relatively quiet scenarios with only one user, existing voice interaction systems can effectively parse user commands and drive the devices.

[0003] However, vending machines are typically deployed in public open environments such as shopping malls, pedestrian streets, or transportation hubs, where there is high ambient noise and random conversations among non-interacting personnel. Traditional device sound pickup solutions mainly rely on a single microphone or conventional sound field noise reduction algorithms, which struggle to accurately distinguish the voice of the target user standing in front of the device from surrounding noise when faced with complex external acoustic environments. When the conversations of bystanders or environmental broadcasts contain words similar to preset instructions, the vending system is prone to generating incorrect recognition responses, thereby triggering unexpected mechanical dispensing or payment processes.

[0004] To address the limitations of pure audio pickup, some existing technologies attempt to incorporate visual images to assist voice interaction. However, current methods primarily rely on a rough judgment of the user's facial orientation and a simple logical superposition of audio endpoint detection. They lack the extraction and calculation of the temporal causal relationship between the physical muscle movements of the target person and the energy fluctuations of the acoustic signal, failing to accurately pinpoint the physical sound source of the actual voice command. Furthermore, most current self-service vending systems are in a passive, waiting-for-interaction state, lacking proactive greetings based on environmental awareness and multi-round voice marketing guidance capabilities, resulting in difficulty in improving order conversion and repurchase rates. Simultaneously, existing devices heavily rely on cloud networks for voice recognition and transaction settlement. In extreme situations such as network congestion or outages in shopping malls, these devices often become completely paralyzed, lacking local edge computing fallback strategies, severely impacting actual business operations and user experience. Summary of the Invention

[0005] The technical problem solved by this invention is that existing self-service vending machines, in noisy environments, cannot effectively separate the voice of the target user from background noise or the voice of non-interactive personnel by relying solely on audio pickup technology, which can easily lead to the device being triggered erroneously or executing incorrect commands.

[0006] To address the above problems, the present invention provides the following technical solution:

[0007] The first aspect of the present invention provides a juicer vending system based on vision and voice interaction, comprising: a data synchronization acquisition module, an audio envelope construction module, a pose and main axis construction module, an optical flow filtering and gating module, a coherence calculation module, a dynamic shopping guide and marketing module, a scheduling and execution module, and an edge network outage backup module;

[0008] This system achieves temporal synchronization of audio and video by constructing a sliding time window, and extracts the longitudinal optical flow component of the target mandibular region using the anatomical principal axis vector as a spatial reference.

[0009] By combining the dynamic time warping algorithm to calculate the alignment distance between the audio energy envelope and the visual motion envelope, it is possible to spatially bind the real sound source with the visual target.

[0010] Building on this foundation, the system further integrates visual ranging-based proactive greeting, multi-round voice shopping guidance, full-process order closed-loop control, and edge network outage backup technology.

[0011] The audio envelope construction module is configured to perform speech activity detection and short-time energy calculation on audio data within a sliding time window to generate a one-dimensional audio energy envelope discrete time series.

[0012] The pose and principal axis construction module is configured to extract the facial node coordinates and bounding box size of candidate targets in video frames, generate an interaction priority queue, and calculate the two-dimensional normalized anatomical principal axis vector of the targets in the interaction priority queue.

[0013] The optical flow filtering and gating module is configured to calculate the two-dimensional dense optical flow vector within the local region of the target bounding box, perform dot product projection and spatial integration on the two-dimensional dense optical flow vector and the two-dimensional normalized anatomical principal axis vector to generate a one-dimensional visual motion purification envelope sequence; the coherence calculation module is configured to use a dynamic time warping algorithm to perform path mapping alignment on the one-dimensional audio energy envelope discrete time sequence and the one-dimensional visual motion purification envelope sequence of the target, and extract a coherence score based on the warping distance;

[0014] The scheduling and execution module is configured to trigger the corresponding feature intent binding and control instruction issuance process based on the comparison result between the coherence score and the preset coherence threshold.

[0015] This system achieves temporal synchronization of audio and video by constructing a sliding time window, and extracts the longitudinal optical flow component of the target mandibular region using the anatomical principal axis vector as a spatial reference.

[0016] By combining dynamic time warping algorithms to calculate the alignment distance between the audio energy envelope and the visual motion envelope, nonlinear time delays between audio and video data can be tolerated, and the coherence between facial vocalization actions and actual sound field energy fluctuations can be quantified.

[0017] This solution can spatially bind real sound sources to visual targets, suppressing acoustic interference caused by non-interactive personnel and ambient background noise.

[0018] Preferably, the audio envelope construction module is specifically configured as follows:

[0019] If it is determined that there is speech activity in the audio data, the audio sampling points are segmented according to the video frame rate within the sliding time window to construct an audio frame sequence.

[0020] Apply a Hamming window to the audio frame sequence and calculate the short-time energy envelope value; perform normalization on the short-time energy envelope value and output a one-dimensional discrete-time audio energy envelope sequence.

[0021] The above configuration smooths out frequency domain leakage by performing frame segmentation and windowing on the audio signal, converting high-frequency sound waves into an energy envelope sequence corresponding to the video frame rate dimension, and providing a scale-uniform data foundation for the alignment calculation of cross-modal data.

[0022] Preferably, the pose and main axis construction module is specifically configured to: obtain the coordinates of the nose tip node and the neck node in the facial node coordinates;

[0023] Construct an initial two-dimensional direction vector pointing from the coordinates of the neck node to the coordinates of the nose tip node;

[0024] Perform a magnitude normalization operation on the initial two-dimensional direction vector to generate a two-dimensional normalized anatomical principal axis vector.

[0025] Since vocalization mainly causes longitudinal displacement of the jaw and neck muscles, the aforementioned anatomical principal axis vector is constructed to establish a local spatial reference axis that is not affected by the camera's perspective for subsequent extraction of visual vocalization features.

[0026] Preferably, when the optical flow filtering and gating module performs dot product projection and spatial integration on the two-dimensional dense optical flow vector and the two-dimensional normalized anatomical principal axis vector to generate a one-dimensional visual motion purification envelope sequence, the specific configuration is as follows:

[0027] The lower half of the target bounding box is selected as the analysis region, and the pixel-level two-dimensional dense optical flow vector between adjacent video frames within the analysis region is calculated frame by frame.

[0028] Perform a dot product operation between the pixel-level two-dimensional dense optical flow vector and the two-dimensional normalized anatomical principal axis vector to output a projection scalar; perform a non-negative truncation operation on the projection scalar, and set the projection scalar with a value less than zero to zero;

[0029] A one-dimensional visual motion purification envelope sequence is constructed by performing spatial integration and summation on the projected scalar after non-negative truncation in pixel space.

[0030] The aforementioned non-negative truncation and integration operations filter out optical flow interference components caused by lateral head swaying or unrelated displacement, and directionally extract longitudinal motion features related to vocalization, thereby improving the correlation between the generated visual motion envelope and the real speech energy fluctuations.

[0031] Preferably, the optical flow filtering and gating module is further configured to: calculate the statistical variance of the one-dimensional visual motion purification envelope sequence;

[0032] If the statistical variance is greater than the preset noise floor variance threshold, the visual envelope validity status flag of the corresponding target is set to valid.

[0033] By filtering the visual purification envelope using a variance threshold, background targets without significant facial movements can be eliminated in advance, avoiding the introduction of invalid targets into the matching operation and reducing the system's computational resource consumption.

[0034] Preferably, the coherence calculation module is specifically configured as follows: performing standardization processing on the one-dimensional audio energy envelope discrete time sequence and the one-dimensional visual motion purification envelope sequence respectively to generate a normalized audio sequence and a normalized visual sequence;

[0035] Establish a cumulative distance matrix and set the global constraint window width during the state transition calculation; for matrix elements whose absolute difference in coordinate indices is greater than the global constraint window width, set their state transition cost to infinity;

[0036] Extract the endpoint coordinates from the cumulative distance matrix as the cumulative minimum path distance;

[0037] An exponential decay function is used to perform a mapping operation on the cumulative minimum path distance, and a coherence score is output.

[0038] A global constraint window is introduced in the cumulative distance calculation to limit the maximum offset range of time series alignment, prevent mismatch caused by over-regulation, and make the mapping result conform to the reasonable time delay between physical sound generation and visual capture.

[0039] Preferably, the scheduling and execution module is specifically configured as follows: when the coherence score is greater than the preset coherence threshold, the corresponding target is set as the sound source target;

[0040] The pixel coordinates of the sound source target are back-projected into a three-dimensional spatial direction ray, and the azimuth and pitch parameters are calculated to construct a spatial steering vector.

[0041] Adaptive beamforming filtering is performed on the acquired multi-channel frequency domain signal of the microphone array based on the spatial steering vector, and a single-channel target frequency domain signal is output.

[0042] The single-channel target frequency domain signal is converted into a time domain audio sequence for speech recognition and text parsing to generate control commands.

[0043] The system guides the microphone array to perform beamforming based on the spatial coordinate parameters of the bound sound source target, forming a pickup lobe at the target location and suppressing signal interference from non-sound source directions, thereby improving the signal-to-noise ratio and the accuracy of command parsing input to the speech recognition module.

[0044] In a preferred embodiment of the present invention, the system further includes an edge intelligent terminal, a microphone pickup array, a visual camera peripheral, and a lower-level machine execution mechanism;

[0045] The microphone pickup array and visual camera peripherals are respectively connected to the edge intelligent terminal to provide audio streams and video frames to the system;

[0046] The data synchronization acquisition module, audio envelope construction module, pose and main axis construction module, optical flow filtering and gating module, coherence calculation module, and scheduling and execution module are all integrated and deployed in the edge intelligent terminal.

[0047] By utilizing edge intelligent terminals for localized multimodal data processing, the impact of network communication latency on audio and video timestamp synchronization and alignment is reduced.

[0048] Preferably, the lower-level actuator includes a programmable logic controller and an electromagnetic dispensing valve; the edge intelligent terminal is configured to receive control commands generated by the scheduling and execution module and send the control commands to the programmable logic controller to drive the electromagnetic dispensing valve to perform the juice filling and dispensing operation.

[0049] This establishes a complete hardware control chain from sensor acquisition and multimodal intent recognition to physical actuator response.

[0050] The second aspect of the present invention provides a method for selling juice machines based on visual and voice interaction, comprising: performing time-series alignment of the acquired audio stream and video frames based on hardware timestamps, and constructing a sliding time window containing multiple frames of data;

[0051] Speech activity detection and short-time energy calculation are performed on the audio data within the sliding time window to generate a one-dimensional audio energy envelope discrete time series.

[0052] Extract the facial node coordinates and bounding box size of candidate targets in video frames, generate an interaction priority queue, and calculate the two-dimensional normalized anatomical principal axis vector of the targets in the interaction priority queue.

[0053] Calculate the two-dimensional dense optical flow vector within the local region of the target bounding box, perform dot product projection and spatial integration on the two-dimensional dense optical flow vector and the two-dimensional normalized anatomical principal axis vector to generate a one-dimensional visual motion purification envelope sequence;

[0054] A dynamic time warping algorithm is used to perform path mapping alignment between the discrete time series of the one-dimensional audio energy envelope of the target and the one-dimensional visual motion purification envelope series, and to extract a coherence score based on the warping distance.

[0055] Based on the comparison results between the coherence score and the preset coherence threshold, the corresponding feature intent binding and control command issuance process is triggered.

[0056] The method provided by the second aspect of the present invention has the same specific principle and beneficial effects as the system provided by the first aspect, and will not be repeated here.

[0057] This invention provides a juice vending system and method based on AI vision and voice interaction. It has the following beneficial effects:

[0058] 1. This invention constructs a two-dimensional normalized anatomical principal axis vector pointing from the neck node to the nose tip node, and performs dot product projection and non-negative truncation operations on the two-dimensional dense optical flow vector in the local area of ​​the target bounding box with it. This enables the targeted extraction of the longitudinal motion component of the mandible related to the actual vocalization action of the human body, filtering out visual motion interference caused by the lateral shaking of the target head or irrelevant body displacement, and improving the targeting and accuracy of extracting visual vocalization features.

[0059] 2. This invention employs a dynamic time warping algorithm with a globally constrained window width to perform path mapping and alignment between the discrete time series of audio energy envelopes and the purified visual motion envelope series. This algorithm can accommodate the nonlinear time delays between cross-modal signals during physical transmission and between human muscle responses. By quantifying the physical coherence between facial opening and closing movements and acoustic energy fluctuations, it achieves accurate spatial binding of real sound targets and eliminates interference from pseudo-instructions caused by conversations between non-interactive personnel.

[0060] 3. This invention performs three-dimensional spatial back projection based on the pixel coordinates of the sound source target locked by coherence score, constructs a spatial steering vector, and uses this vector to guide the microphone array to perform adaptive beamforming filtering on the multi-channel frequency domain signal, forming a pickup main lobe at the actual location of the target, suppressing environmental background noise in non-target directions, improving the signal-to-noise ratio of the input to the subsequent speech recognition stage, and ensuring the accuracy of juicer control command parsing and mechanical execution.

[0061] 4. This invention breaks through the limitations of passive waiting in traditional vending machines. By introducing an active voice greeting triggered at a preset distance of 1.5 to 3 meters, and a dynamic multi-round voice guidance and marketing promotion mechanism based on the number of people and time periods, it realizes a business upgrade from passive transactions to active customer acquisition, significantly improves the order conversion rate and user repurchase rate of unmanned equipment, and greatly reduces the operating costs of manual sales guides.

[0062] 5. This invention completes the business loop of automated vending and innovatively introduces an edge network failure fallback module. In the event of a network failure, the system can seamlessly switch to a local pre-built script library and a lightweight offline acoustic model, ensuring basic voice interaction and local order processing in the network failure environment, and greatly improving the operational stability and risk resistance of the equipment in various business environments. Attached Figure Description

[0063] Figure 1 This is a diagram of the interactive system architecture for selling juice machines according to an embodiment of the present invention.

[0064] Figure 2 This is a flowchart of the juicer sales interaction method according to an embodiment of the present invention;

[0065] Figure 3 This is a timing diagram illustrating the multimodal data synchronization and sliding time window construction principle of an embodiment of the present invention.

[0066] Figure 4 This is a schematic diagram illustrating the principle of audio low-level feature parsing and envelope construction in an embodiment of the present invention.

[0067] Figure 5 This is a schematic diagram illustrating the pose pre-screening and spindle construction principle of an embodiment of the present invention.

[0068] Figure 6 This is a schematic diagram of the directional optical flow projection and variance gating principle in an embodiment of the present invention;

[0069] Figure 7 This is a schematic diagram illustrating the principle of cross-modal dynamic time warping and coherence evaluation in an embodiment of the present invention.

[0070] Figure 8 This is a schematic diagram illustrating the principle of directional beamforming and service command generation in an embodiment of the present invention.

[0071] Figure 9 This is a dynamic time warping alignment and coherence mapping diagram of a dual-modal sequence according to an embodiment of the present invention. Detailed Implementation

[0072] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0073] Example:

[0074] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0075] See attached document Figure 1 , Figure 1 This is an architecture diagram of a juicer vending interactive system according to an embodiment of the present invention.

[0076] This invention provides a juice machine vending interactive system, which may include:

[0077] The system includes a data synchronization acquisition module, an audio envelope construction module, a pose and spindle construction module, an optical flow filtering and gating module, a coherence calculation module, and a scheduling and execution module.

[0078] The hardware topology of the juice machine vending interaction system includes edge computing nodes, audio acquisition peripherals, visual acquisition peripherals, and lower-level execution mechanisms. The edge computing node serves as the system's main control computing center, integrating a central processing unit, a graphics processing unit, and a neural network processing unit.

[0079] The audio acquisition peripheral uses a single-channel electret microphone.

[0080] The audio acquisition peripheral establishes a data transmission path with the edge computing node through the system's internal audio bus to acquire single-channel audio stream data in the device's physical environment.

[0081] The visual acquisition peripherals use RGB cameras with global exposure or rolling shutter exposure specifications.

[0082] The visual acquisition peripheral is connected to the edge computing node via a video bus to acquire video stream data in front of the juicer.

[0083] The lower-level actuator includes a programmable logic controller, a cup-dropping module, a solenoid valve, and a hydraulic pump.

[0084] Edge computing nodes connect to programmable logic controllers via standard serial communication protocols and issue hardware driver commands to control the cup-dispensing module, solenoid valve, and hydraulic pump to perform cup dispensing and filling actions.

[0085] The data synchronization acquisition module is configured to receive raw data streams from audio acquisition peripherals and visual acquisition peripherals.

[0086] The data synchronization acquisition module uses a unified hardware timestamp at the underlying level to perform physical time-domain timing alignment between the acquired continuous audio stream and discrete video frames.

[0087] The data synchronization acquisition module divides the data according to the set period and constructs a sliding time window containing multiple consecutive frames of data.

[0088] The audio envelope construction module is configured to receive sliding time window data output by the data synchronization acquisition module. The audio envelope construction module performs speech activity detection on the audio data within the time window.

[0089] When a valid speech signal is detected, the audio envelope construction module performs a windowing operation on each discrete audio frame to extract short-time energy in the time domain.

[0090] The audio envelope construction module combines the extracted short-time energies from each frame to generate a one-dimensional audio energy envelope discrete-time series.

[0091] The pose and axis building module is configured to process video data within a time window.

[0092] The pose and main axis construction module calls the object detection and lightweight pose estimation algorithm to obtain the bounding box size of the candidate object and the pixel coordinates of the key facial nodes.

[0093] The pose and main axis construction module filters candidate targets based on the set interactive view frustum angle threshold and distance area threshold.

[0094] The pose and principal axis construction module sorts targets that meet the threshold conditions in descending order based on the bounding box area, extracts targets limited by the maximum number of concurrent processing targets of edge computing nodes, and generates an effective interaction priority queue.

[0095] For each target in the queue, the pose and principal axis construction module calculates the two-dimensional normalized anatomical principal axis vector from the neck to the nose tip based on the obtained nose tip node coordinates and neck node coordinates.

[0096] The optical flow filtering and gating module is configured to receive the effective interaction priority queue and the anatomical principal axis vectors of each target from the pose and principal axis construction module. The optical flow filtering and gating module extracts the lower half of the target bounding box as the analysis region.

[0097] The optical flow filtering and gating module calculates and analyzes pixel-level two-dimensional dense optical flow vectors between adjacent video frames within the region.

[0098] The optical flow filtering and gating module performs a dot product operation between the two-dimensional dense optical flow vector and the corresponding two-dimensional normalized anatomical principal axis vector.

[0099] The optical flow filtering and gating module performs spatial integration and summation on the non-negative projection scalars generated by the dot product operation to generate a one-dimensional visual motion purification envelope sequence for each target.

[0100] The optical flow filtering and gating module calculates the statistical variance of the one-dimensional visual motion purification envelope sequence.

[0101] Based on the comparison results between the statistical variance and the preset noise floor variance threshold, the optical flow filtering and gating module sets the visual envelope validity status flag of the target.

[0102] The coherence calculation module is configured to receive timing data output from the audio envelope construction module and the optical flow filtering and gating module.

[0103] The coherence calculation module extracts the intent command text corresponding to the audio stream within the time window.

[0104] For targets where the visual envelope validity status flag is in a valid state, the coherence calculation module uses a dynamic time warping algorithm to perform path mapping alignment between the one-dimensional audio energy envelope discrete time series and the one-dimensional visual motion purification envelope series.

[0105] The coherence calculation module calculates the coherence score of the two time series after mapping based on the regular distance.

[0106] The scheduling and execution module is configured to trigger the corresponding control branch based on the coherence score output by the coherence calculation module.

[0107] When the relevance score of a specific target is greater than the preset relevance threshold, the scheduling and execution module establishes a binding association between the visual target features and the intent command text, and calls the external payment interface.

[0108] After verification, the scheduling and execution module drives the lower-level machine execution mechanism to complete the filling and shipping process. When the coherence score of all targets is lower than the coherence threshold, and the target with the largest area in the effective interaction priority queue is in a visual envelope failure state, the scheduling and execution module suspends the coherence binding stream.

[0109] The scheduling and execution module triggers the system's UI to enter touch confirmation mode, and determines the subsequent control flow based on the physical touch feedback signal.

[0110] See attached document Figure 2 , Figure 2 This is a flowchart of a juicer sales interaction method according to an embodiment of the present invention. The present invention provides a juicer sales interaction method, including the following steps:

[0111] In step S100, the pose perception and active welcoming module perceives the environmental passenger flow and triggers the welcoming.

[0112] The system captures candidate targets in front of the device in real time through a visual acquisition peripheral.

[0113] The physical depth of the target is calculated based on the principle of binocular vision parallax or the perspective relationship of bounding box pixels.

[0114] When a candidate target's depth value is detected to be approaching within a preset welcoming distance range of 1.5 to 3 meters, and its facial yaw angle is continuously pointing towards the device, it is determined that there is a purchase intention. The system immediately plays welcoming voice messages such as "Hello, welcome to Fresh Juice Bureau" through an external speaker to achieve proactive interception.

[0115] Meanwhile, the system extracts the neck and nose tip nodes of the target and constructs a two-dimensional normalized anatomical principal axis vector for subsequent noise reduction.

[0116] In step S200, the dynamic shopping guide and marketing module initiates multiple rounds of interaction.

[0117] After proactively greeting the guest, the system reads the current clock and counts the number of people in the same frame. If the time period is in the morning and there is only one person, an energizing message is sent: "Have a glass of vitamin C-rich orange juice to start your day with energy."

[0118] If the system detects that two people are traveling together, it will send a promotional message: "We currently have a 'buy one get one half price' promotion."

[0119] Through the above logic, the system transforms the traditional passive waiting into proactive sales guidance, encouraging customers to speak.

[0120] S300: Extract the facial node coordinates and bounding box size of candidate targets in video frames, generate an effective interaction priority queue based on threshold conditions, and calculate the two-dimensional normalized anatomical principal axis vector of the targets in the queue.

[0121] Step S400: Coherence calculation and sound source locking.

[0122] The system uses a dynamic time warping algorithm to perform path mapping and alignment of audio and visual envelope sequences and calculates coherence scores.

[0123] When the score exceeds the preset threshold, the target in the visual image is absolutely bound to the sound captured by the microphone array to confirm that the instruction was indeed issued by the customer standing in front of the machine, and not by a passerby.

[0124] Step S500: Closed loop execution of business scheduling.

[0125] The directional voice pickup and command parsing module performs text recognition on the locked, clean voice (or the system directly receives the customer's touch selection on the screen). After clarifying the purchase intention, the scheduling and execution module initiates the entire sales process:

[0126] (1) Orders and Payments: The system generates an order for the corresponding beverage, renders the payment QR code on the multimedia display screen, and announces "Your orange juice order has been placed, please scan the code to pay";

[0127] (2) Status verification: Poll the payment gateway, and switch to the creation interface after receiving a payment success callback;

[0128] (3) Cup making execution: The instruction is sent to the programmable logic controller (PLC) to drive the cup dropper and the electromagnetic liquid dispensing valve to perform juice filling;

[0129] (4) Cup Retrieval Reminder: When the PLC feedback indicates that the safety door has been opened and filling is complete, the system plays "Your juice is ready. Please take your cup. Welcome to visit us again." This completes the business loop.

[0130] Step S600: Global edge network outage fallback monitoring.

[0131] Based on the comparison results between the coherence score and the preset coherence threshold, and combined with the status flag of the main target in the effective interaction priority queue, the corresponding feature intent is triggered to bind to the shipment control flow or the ambiguity resolution touch control flow.

[0132] The following section will provide a detailed explanation of the specific execution steps and underlying algorithm logic of the method provided in the embodiments of the present invention, in conjunction with specific application scenarios and mathematical derivation processes.

[0133] See attached document Figure 3 , Figure 3 This is a timing diagram illustrating the multimodal data synchronization and sliding time window construction principle according to an embodiment of the present invention.

[0134] The system based on temporal coherence and directional projection filtering provided by this invention ensures the physical consistency of audio and video signals in subsequent calculations by executing a low-level data synchronization mechanism.

[0135] The execution process of step S100 can be further subdivided into three specific processing steps: hardware timestamp allocation, cross-modal timing intercept alignment, and sliding time window memory organization.

[0136] In step S101, the data synchronization acquisition module obtains the sensor's underlying data stream at the driver layer and assigns a hardware timestamp.

[0137] In a multimodal streaming data acquisition environment, there is an inherent difference between the audio sampling frequency of a single-channel microphone and the video output frame rate of an RGB camera.

[0138] The data synchronization acquisition module intercepts the low-level interrupt signals transmitted to the system bus from the audio acquisition peripheral and the visual acquisition peripheral, and calls the system hardware clock inside the edge computing node to uniformly allocate nanosecond-level hardware timestamps for the data stream entering the memory buffer.

[0139] For continuously sampled audio data streams, the data synchronization acquisition module timestamps the buffer data blocks composed of a fixed number of sampling points.

[0140] The number of fixed sampling points is determined by multiplying the analog-to-digital conversion sampling rate of the audio acquisition peripheral with the underlying interrupt polling cycle set by the system, in order to balance data throughput efficiency and real-time requirements.

[0141] For discrete output video streams, the data synchronization acquisition module records the corresponding timestamp when the image sensor completes the exposure of a single frame image and outputs it to the system's video memory.

[0142] In step S102, the data synchronization acquisition module performs a cross-modal timing alignment operation based on a low-frequency clock.

[0143] Since the acquisition frame rate of the visual peripheral is lower than that of the audio peripheral, the data synchronization acquisition module adopts a data alignment strategy based on the timestamps of discrete video frames.

[0144] The data synchronization acquisition module reads the exposure start hardware timestamp of the current specific video frame and the exposure start hardware timestamp of the next adjacent frame, and uses the difference between the two to form the physical time interval of a single frame.

[0145] Based on the determined single-frame physical time interval, the data synchronization acquisition module matches the interval segment corresponding to the timestamp in the system audio memory buffer and extracts a continuous audio sampling point sequence.

[0146] During this matching process, the data synchronization acquisition module performs a difference check on the physical time interval of a single frame. If the difference between two adjacent video frames exceeds the frame drop tolerance threshold set by the system, the data synchronization acquisition module will determine that a visual frame drop anomaly has occurred and perform zero-padding at the end of the truncated audio sampling point sequence to maintain data length consistency.

[0147] The extracted audio sampling point sequence is aligned with the specific video frame in the physical time domain to form a pair of synchronized multimodal data frames.

[0148] The specific implementation process of microphone audio analog-to-digital conversion acquisition and camera image signal processing exposure output can be completed by conventional audio codec chips and image signal processing devices in conjunction with corresponding underlying drivers. The hardware interrupt triggering and digital signal conversion mechanism are well-known technologies in this field and will not be elaborated here.

[0149] Step S103: The data synchronization acquisition module constructs and maintains a sliding time window for time series feature analysis.

[0150] Since a single frame of data cannot capture the continuous changes in the target object's vocalization and facial muscle movements, the data synchronization acquisition module allocates a circular buffer in the system's shared memory area to store and update synchronized multimodal data frames. The data synchronization acquisition module sets the capacity of this circular buffer to a length of [missing information]. The analysis period consisting of a series of consecutive synchronized frames is defined as the sliding time window W.

[0151] Among them, capacity length The value is determined based on the expected interaction response latency and the available video memory of the edge computing node.

[0152] In one specific implementation, to cover the complete pronunciation cycle of a short speech command, the video frame rate is set to 15 frames per second. The value ranges from 15 to 45, corresponding to a physical analysis time of 1 to 3 seconds.

[0153] Within the time window W, the system introduces a discrete-time index. ,and .

[0154] Using time index The set of video frames within a time window is defined as a discrete video frame sequence. And define the corresponding aligned audio stream data as a synchronized audio frame sequence. .

[0155] The data synchronization acquisition module updates the data sequence within the sliding time window W according to the first-in-first-out cache eviction mechanism.

[0156] When a new pair of synchronized multimodal data frames is written to the circular buffer, the data synchronization acquisition module moves the earliest time-indexed pair of synchronized frames out of the memory area. This sliding time window W provides a fixed-length, equal-dimensional data source for subsequent modules to calculate the one-dimensional motion envelope sequence and short-time energy sequence.

[0157] See attached document Figure 4 , Figure 4 This is a schematic diagram illustrating the principle of audio low-level feature parsing and envelope construction according to an embodiment of the present invention.

[0158] The solution provided by this invention extracts temporal features from the audio stream and filters out high-frequency irrelevant phase information, providing basic one-dimensional data reflecting the sound intensity for subsequent cross-modal temporal mapping.

[0159] The execution process of step S200 can be further subdivided into three specific processing steps: speech activity state discrimination, discrete audio frame windowing calculation, and time sequence organization.

[0160] In step S201, the audio envelope construction module receives sliding time window data containing a sequence of synchronized audio frames and performs voice activity detection.

[0161] In open juicer vending scenarios, the devices are in standby mode most of the time. To avoid the system performing ineffective high-concurrency visual calculations on ambient noise, the audio envelope construction module calculates the short-time zero-crossing rate and logarithmic energy features of each audio frame and compares the extracted features with the system's preset static noise threshold.

[0162] The static noise floor threshold is established by collecting ambient noise of a fixed duration without interaction during the device startup initialization phase and calculating its average energy level.

[0163] If more than a set proportion of audio frames within the time window meet the criteria for valid vocal segments, the audio envelope construction module triggers the subsequent short-time energy extraction process.

[0164] The set percentage is usually set between 60% and 80% to exclude occasional sudden environmental noise interference.

[0165] If the set ratio is not reached, the system determines that the current time window is in a low ambient noise or silent state. Then, it bypasses the array multi-channel audio data in the current time window and outputs it to the environmental noise evaluation buffer for subsequent updates of the spatial covariance matrix. After that, it releases the cached data of the main analysis process and waits for the data to be written in the next cycle to reduce the computing power overhead of the computing nodes.

[0166] For the specific algorithm implementation of speech activity detection, those skilled in the art can use conventional methods such as Gaussian mixture models or gated loop units to complete the endpoint detection and silence removal of audio segments. The underlying statistical discrimination mechanism is a well-known technology in this field and will not be elaborated here.

[0167] In step S202, the audio envelope construction module performs weighted summation and short-time energy calculation on discrete audio frames within the time window of the detected valid speech signal.

[0168] The sliding time window contains Indexed by time Arranged synchronized audio frames .

[0169] Assuming a single frame of audio Include The nth discrete sampling point is defined as the nth... The amplitude value of each sampling point is The time index Sampling point index

[0170] Since discrete audio frames are forcibly segmented in physical timing, directly calculating energy will cause spectral leakage and data abrupt changes at the frame edges.

[0171] The audio envelope construction module introduces a smoothing window function to the audio frame. Perform weighted processing.

[0172] From the perspective of signal processing principles, the function of the smoothing window function is to make the amplitude of the sampling points at both ends of the audio frame gradually transition to zero, thereby eliminating the abrupt changes at the truncation edge.

[0173] The audio envelope construction module squares the amplitude values ​​of the windowed discrete sampling points and accumulates them within the frame to obtain the short-time energy of the audio frame. The calculation formula is as follows:

[0174] ;

[0175] Among them, parameters The value of is limited by the analog-to-digital conversion sampling rate of the audio acquisition peripheral and the physical frame spacing of the video during cross-modal alignment.

[0176] When the system's hardware audio sampling rate is set to 16000 Hz and the synchronous video frame rate is 15 frames per second, the number of sampling points contained in a single frame of discrete audio is... The value is usually 1066 or 1067, and this parameter is directly calculated based on the underlying hardware configuration parameters of the system.

[0177] Smooth window function In practical engineering implementation, Hamming windows can be used, and their expression is defined as follows:

[0178] In step S203, the audio envelope construction module integrates the calculation results of each frame to generate a one-dimensional discrete time series.

[0179] The audio envelope construction module is based on time indexing. The increasing order will be used to select all [times] within the sliding time window. Short-time energy scalar corresponding to each synchronized audio frame The data is sequentially stored in a contiguous memory array of the master node. Through temporal data combination, the audio envelope construction module constructs a one-dimensional audio energy envelope discrete-time sequence E, with a length equal to the number of video frames.

[0180] ;

[0181] The one-dimensional audio energy envelope discrete time series E discards the high-frequency phase information in the original audio and extracts the temporal variation trend reflecting the sound intensity.

[0182] The one-dimensional audio energy envelope discrete time series E serves as the reference data source for subsequent cross-modal dynamic time warping operations.

[0183] See attached document Figure 5 , Figure 5 This is a schematic diagram of the pose pre-screening and spindle construction principle according to an embodiment of the present invention.

[0184] The solution provided by this invention filters the spatial pose of the target within the field of view and establishes a reference principal axis representing the direction of the human vocal organs on the two-dimensional image plane, providing a benchmark for subsequent optical flow calculation.

[0185] The execution process of step S300 can be further subdivided into three specific processing steps: feature parameter extraction, effective interaction priority queue generation, and normalized principal axis vector construction.

[0186] Step S301: The pose and principal axis construction module extracts the surface pose parameters of the candidate targets within the current video frame.

[0187] After completing the candidate target detection, the pose and principal axis construction module calculates the physical depth value of the target from the device based on the principle of binocular visual parallax or the geometric perspective relationship between the pixel height of the target bounding box in the image and the average height of the real human body.

[0188] When a candidate target is detected to be approaching from a distance of 1.5 to 3 meters from a distance, and its facial yaw angle is continuously pointing towards the device screen, the system determines that the target has potential purchasing intentions and immediately triggers the system's proactive welcoming service flow. The system plays a welcoming voice message such as "Hello, welcome to the Fresh Juice Machine, what would you like to drink today?" through an external speaker, thus achieving proactive interception and customer flow guidance.

[0189] Step S302: The pose and spindle construction module generates an effective interaction priority queue in combination with hardware computing power constraints.

[0190] In an open environment where juice machines are deployed, there may be people who do not intend to interact, such as those passing by, those with their backs to the machine, or those at a distance.

[0191] The pose and main axis construction module sets the interactive view frustum angle threshold and distance area threshold to make a preliminary judgment on candidate targets.

[0192] The pose and spindle construction module determine the target's yaw angle. With pitch angle Whether all values ​​are within the preset frustum angle threshold range. This frustum angle threshold range is set as follows: It is used to filter out targets with their faces turned to the side or back to the device.

[0193] The pose and main axis construction module calculates the pixel area of ​​the target based on the bounding box coordinates and determines whether its pixel area is greater than a preset distance area threshold.

[0194] The distance area threshold is set to 5% of the total pixel area of ​​the entire image, which is used to filter out targets that are too far away and whose facial feature resolution is insufficient.

[0195] After the threshold determination is completed, the pose and principal axis construction module sorts the valid targets that meet the conditions in descending order based on the bounding box pixel area.

[0196] To prevent high-concurrency operations from overloading edge computing node resources, the pose and spindle construction module extracts and sorts the results based on a set upper limit for concurrency, generating an effective interaction priority queue. .

[0197] The maximum number of concurrent operations is determined based on the evaluation of the parallel computing power of the neural network processing unit within the master node. In this embodiment, it is set to 3, meaning that the system can process the feature operations of a maximum of 3 main targets within a single time window.

[0198] Step S303: The pose and principal axis construction module constructs a two-dimensional normalized anatomical principal axis vector for the targets in the priority queue.

[0199] When the human body speaks, the physical movement of the mandible and the vocal muscles in the neck is parallel to the longitudinal anatomical axis from the neck to the center of the face. Extracting the direction of this axis can provide a reference for subsequent filtering of lateral interference movements.

[0200] Priority queue for valid interactions For each target in the model, the pose and main axis construction module reads the corresponding neck node coordinates. coordinates of the nose tip node Calculate the initial two-dimensional direction vector from the neck to the tip of the nose. The calculation formula is as follows:

[0201] ;

[0202] Before performing subsequent calculations, the pose and main axis construction module calculates the Euclidean distance between two key nodes. If the distance between the nodes is less than the minimum tolerance value set by the system (set to 2 pixels in this embodiment), it indicates that the node coordinate predictions coincide or severe occlusion occurs. The system will discard the pose data of the target in the current frame to prevent division by zero anomalies.

[0203] For a target that passes the distance verification, the pose and principal axis construction module initializes the two-dimensional direction vector. Perform modulus normalization to generate two-dimensional normalized anatomical principal axis vectors that eliminate scale differences. The calculation formula is as follows:

[0204] ;

[0205] Calculated generated two-dimensional normalized anatomical principal axis vectors It characterizes the spatial orientation of the target vocal organ in the two-dimensional image coordinate system and outputs it as an independent directional feature variable to subsequent modules.

[0206] Step S304: After establishing an effective interaction priority queue and locking the target, the dynamic shopping guide and marketing module is connected to the interaction process.

[0207] The dynamic shopping guide and marketing module reads the current system's physical clock and receives the target number of people in the same frame from the visual acquisition peripheral.

[0208] The system has a built-in multi-dimensional marketing message decision tree: for example, if the time period is in the morning and the number of people is 1, the system will push an energizing message such as "Good morning, have a glass of orange juice rich in vitamin C to start your day with energy";

[0209] If the visual system identifies two people or a family, it dynamically generates a promotional message such as "Hello, we currently have a buy-one-get-one-half-price promotion. We recommend you try our mixed juice."

[0210] Upon receiving an initial vague voice response from a user (such as "What recommendations do you have?"), the system maintains the coherence feature binding, initiates multiple rounds of dialogue interaction, and guides the user to complete membership registration or place an order for high-profit beverages, transforming traditional command execution into a sales-oriented intelligent shopping guide.

[0211] See attached document Figure 6 , Figure 6 This is a schematic diagram of directional optical flow projection and variance gating according to an embodiment of the present invention.

[0212] The solution provided by this invention performs spatial orientation filtering on motion vectors in a specific local area to remove environmental interference and irrelevant limb movements, generates a purified one-dimensional visual motion envelope sequence, and establishes a gating mechanism based on statistical features.

[0213] The execution process of step S400 can be further subdivided into four specific processing steps: local region truncation and optical flow calculation, directional projection and non-negative truncation, spatial integral sequence construction, and variance gating flag setting.

[0214] Step S401: The optical flow filtering and gating module extracts a specific local region of the target and calculates pixel-level two-dimensional dense optical flow. Facial muscle displacement is mainly concentrated in the jaw and neck regions.

[0215] The optical flow filtering and gating module obtains the target bounding box coordinates output from the previous step and divides the bounding box into upper and lower parts based on the height parameter.

[0216] Before segmentation, the optical flow filtering and gating module verifies the vertical pixel height of the bounding box. If the height is lower than the system's preset effective truncation lower limit, it is proportionally filled and enlarged according to the lower limit value to avoid calculation failure caused by the truncation area being too small.

[0217] The optical flow filtering and gating module extracts the lower half of the bounding box as the region of interest to remove motion interference caused by eye blinking or head background.

[0218] Within the sliding time window, for the region of interest, the optical flow filtering and gating module calculates the pixel-level two-dimensional dense optical flow field between adjacent video frames frame by frame.

[0219] For a pixel at a specific discrete time index The calculated dense optical flow vector is defined as ,in and These represent the physical displacement of the pixel in the horizontal and vertical directions, respectively.

[0220] For the specific calculation process of the two-dimensional dense optical flow field, those skilled in the art can use a polynomial expansion algorithm and use an image pyramid to perform multi-scale layer-by-layer estimation. The pixel matching and displacement field construction mechanism are well-known technologies in this field and will not be elaborated here.

[0221] In step S402, the optical flow filtering and gating module performs dot product projection and non-negative truncation on the two-dimensional dense optical flow vector and the two-dimensional normalized anatomical principal axis vector.

[0222] Human movement in real-world environments often involves lateral displacements such as head turning and limb swaying.

[0223] The optical flow filtering and gating module calls the two-dimensional normalized anatomical principal axis vector generated in the previous step. The two-dimensional dense optical flow vector of each pixel in the region of interest. With principal axis vector Perform a mathematical dot product operation to generate a projected scalar representing the intensity of the longitudinal motion of the sound emission. The calculation formula is as follows:

[0224] ;

[0225] Optical flow filtering and gating module for projected scalar Perform a non-negative truncation operation. This operation clears the negative projection to zero, preserves the positive motion component, suppresses reverse motion noise, and extracts the unidirectional force characteristics of the vocal muscles.

[0226] Truncated nonnegative projective scalar The calculation formula is as follows:

[0227] In step S403, the optical flow filtering and gating module performs spatial integration on the non-negative projection scalar to generate a one-dimensional visual motion purification envelope sequence.

[0228] The optical flow filtering and gating module performs optical flow filtering and gating on discrete coordinates within the full pixel space of the region of interest. The corresponding nonnegative projective scalar Perform spatial integration and summation.

[0229] This operation converts the local displacement within the two-dimensional image region into a single-frame macroscopic scalar describing the overall sound intensity of the target. The calculation formula is as follows:

[0230] ;

[0231] in, This represents the set of all pixels within the region of interest, and the total number of pixels it contains determines the computational scale of the integral summation for a single frame.

[0232] Since optical flow calculation is based on the difference relationship between adjacent video frames, the capacity is... Video frame sequence generation An effective single-frame macro scalar.

[0233] The optical flow filtering and gating module will filter all the optical flow within the sliding time window. A single-frame macroscopic scalar is combined according to physical time sequence to generate a one-dimensional visual motion purification envelope sequence.

[0234] ;

[0235] In step S404, the optical flow filtering and gating module calculates the statistical variance of the one-dimensional visual motion purification envelope sequence and sets the gating flag.

[0236] In low-light environments or when the target is physically obscured, the values ​​captured by the optical flow algorithm are mostly dark current noise or background disturbances from the image sensor.

[0237] Optical flow filtering and gating modules calculate one-dimensional visual motion purification envelope sequences. mean and statistical variance The calculation formula is as follows:

[0238] ;

[0239] ;

[0240] The optical flow filtering and gating module will statistically analyze the variance. Compare with the preset noise floor variance threshold.

[0241] The noise floor variance threshold is determined by capturing empty background images of equal-length sliding time windows during the device initialization phase or the silent period when no one is interacting with the system, calculating their optical flow variance, and then multiplying it by a preset width density coefficient, wherein the width density coefficient ranges from 1.2 to 1.5.

[0242] If statistical variance If the variance is greater than the noise floor threshold, it indicates that the sequence contains effective structured motion information. The optical flow filtering and gating module sets the effective state flag of the visual envelope of the target to an effective state.

[0243] If statistical variance If the value is not greater than the noise floor variance threshold, it indicates that the extracted motion data does not have physical representation significance, and the optical flow filtering and gating module sets the status flag to a failure state.

[0244] This flag provides a basis for blocking subsequent processes, preventing the system from performing timing alignment operations on invalid noise.

[0245] See attached document Figure 7 , Figure 7 This is a schematic diagram of the principle of cross-modal dynamic time warping and coherence evaluation according to an embodiment of the present invention.

[0246] The solution provided by this invention calculates the non-rigid temporal alignment distance between the audio energy envelope and the visual motion envelope to adapt to the physiological delay difference between human muscle movements and vocal cord vocalization, and then selects the corresponding target interactive sound source through feature coherence evaluation.

[0247] The execution process of step S500 can be further subdivided into three specific processing steps: pre-gating verification and sequence preprocessing, windowed dynamic time warping calculation, and coherence score mapping and target decision.

[0248] In step S501, the dynamic time warping and coherence assessment module receives the one-dimensional data sequence input from the previous step and performs gating state verification and data preprocessing.

[0249] In multi-target concurrent scenarios, the dynamic time warping and coherence assessment module checks the visual envelope validity status flags of each target in the effective interaction priority queue.

[0250] If the flags of all candidate targets in the queue are invalid, the dynamic time warping and coherence evaluation module will directly determine that there is no effective facial vocalization movement in the current time window and terminate the subsequent time warping calculation to prevent the algorithm's computing resources from being excessively occupied to deal with meaningless background interference.

[0251] For targets with valid status flags, the dynamic time warping and coherence assessment module reads their corresponding one-dimensional visual motion purification envelope sequence. .

[0252] Since the length of the visual sequence is The length of the synchronously extracted one-dimensional audio energy envelope discrete-time sequence E is The dynamic time warping and coherence evaluation module discards the first element of the audio sequence E and truncates it to a length of the same value. Aligned audio sequences .

[0253] Furthermore, since the audio energy value and the pixel displacement integral value are on different physical dimensions, directly calculating the numerical distance will cause the large-scale signal to mask the characteristics of the small-scale signal.

[0254] The dynamic time warping and coherence evaluation module aligns the audio sequence E' and the visual sequence. Perform Z-score standardization on each.

[0255] With visual sequences For example, its standardized elements The calculation formula is as follows:

[0256] ;

[0257]

[0258] Among them, discrete index mean with standard deviation These are the basic statistical parameters for this sequence.

[0259] Similarly, the dynamic time warping and coherence evaluation module calculates the standardized audio sequence elements. Generate a normalized audio sequence that eliminates dimensional differences. With normalized visual sequences .

[0260] In step S502, the dynamic time warping and coherence evaluation module uses the dynamic time warping algorithm to calculate the cumulative minimum path distance of the bimodal sequence.

[0261] Because there is a non-linear physical offset between the facial muscle movements reflected in the visual image and the airborne sound captured by the microphone on the timeline, the linear correlation coefficient is difficult to accurately measure this synchronization characteristic accompanied by delay. The dynamic time warping algorithm can find the optimal alignment path between bimodal sequences by non-rigidly stretching and scaling the time axis.

[0262] The size of the dynamic time warping and coherence assessment module is [size missing]. Cumulative distance matrix Its row index With column index The value range is 0 to .

[0263] To prevent the algorithm from going out of bounds during pathfinding, the dynamic time warping and coherence evaluation module initializes the matrix boundary conditions, setting the starting condition as follows: and to satisfy and boundary elements and All are initialized to infinity.

[0264] After completing the boundary initialization, for elements in the matrix with coordinate indices greater than 0... The state transition calculation formula is as follows:

[0265] ;

[0266] To prevent the algorithm from over-warping and forcing unrelated physiological actions to be matched to audio signals, the dynamic time warping and coherence evaluation module introduces a Sakoe-Chiba global constraint window during the solution of the cumulative distance matrix D.

[0267] The dynamic time warping and coherence assessment module sets the constraint window width parameter to be... And during state transition calculations, the coordinate indices are required to satisfy... .

[0268] The constraint window width parameter The value is determined based on the maximum physiological audiovisual delay difference allowed by the system. When the system sampling frame rate is 15 frames per second and the preset maximum tolerance delay is 200 milliseconds, The value is set to 3.

[0269] If the coordinate index exceeds the set window width, the corresponding transfer cost will be set to infinity.

[0270] After the matrix is ​​solved, the endpoint element It is the shortest regular distance between two sequences under the constraints.

[0271] For the specific implementation process of backtracking path finding in the dynamic time warping algorithm, those skilled in the art can use the standard dynamic programming solution framework to write the algorithm logic. Its underlying mathematical mechanism is a well-known technology in this field and will not be elaborated here.

[0272] In step S503, the dynamic time warping and coherence assessment module maps the warping distance to a coherence score and establishes the final sound source target based on the score.

[0273] Shorter regular distances indicate a physical correspondence between visual facial motion and audio energy fluctuations over time.

[0274] The dynamic time warping and coherence assessment module uses an exponential decay function to calculate the shortest warping distance. Convert to coherence score within the normalized interval [0,1] The calculation formula is as follows:

[0275] ;

[0276] Among them, parameters The mapping attenuation coefficient is a preset parameter for the system. This parameter is calibrated based on the distribution experience of multiple sets of silent and interactive samples, and its value range is set to 0.1 to 0.5.

[0277] The dynamic time warping and coherence assessment module iterates through the coherence scores of all valid targets in the queue and extracts the maximum score.

[0278] The dynamic time warping and coherence assessment module compares the maximum score with a preset judgment threshold, which is used to isolate accidental random synchronous actions.

[0279] The threshold for this determination is determined by collecting a specific number of valid interactive audio-visual clips and environmental interference audio-visual clips offline, calculating their coherence scores and plotting data distribution histograms, and taking the value of the intersection point of the two types of sample distribution curves. The value range is set to 0.6 to 0.8.

[0280] If the maximum score is greater than the judgment threshold, the dynamic time warping and coherence assessment module determines the candidate target with the maximum score as the sound source within the current sliding time window, and outputs the physical coordinates and associated audio data of the target to the downstream business system.

[0281] If the maximum score is not greater than the judgment threshold, the dynamic time warping and coherence evaluation module determines that there is no effective interactive sound source in the field of view and sends an idle standby command to the system bus.

[0282] See attached document Figure 8 , Figure 8 This is a schematic diagram of directional beamforming and service instruction generation according to an embodiment of the present invention.

[0283] The solution provided by this invention performs spatial directional beamforming on the locked interactive target to suppress environmental noise and non-target human voices, extracts the noise-reduced and enhanced voice signal and parses it into device control commands, thus completing the physical closed loop of human-computer interaction.

[0284] The execution process of step S600 can be further subdivided into three specific processing stages: spatial sound source localization, adaptive beamforming filtering, and service command parsing.

[0285] In step S601, the directional sound pickup and command parsing module receives the target physical coordinates output from the previous step and calculates the guidance vector of the spatial sound source.

[0286] The physical coordinates output by the pre-processing step are based on the two-dimensional pixel plane of the vision camera.

[0287] The directional sound pickup and command parsing module calls the camera intrinsic and extrinsic translation and rotation matrices obtained during the system calibration phase, back-projects the target's pixel coordinates into a three-dimensional spatial direction ray with the microphone array center as the origin, and calculates the azimuth angle of the sound source relative to the array reference center based on this. With pitch angle .

[0288] Assuming the device is configured with a microphone array containing Given a set of microphone nodes arranged in a known geometric configuration, with the array reference center set as the origin, the [missing information - likely a specific location or point]. The spatial position vectors of the microphone nodes are: , where index .

[0289] The directional pickup and command parsing module calculates the relative phase shift of each channel in the frequency domain and constructs a spatial steering vector. .

[0290] For angular frequency is signal components, steering vector The calculation formula is as follows:

[0291] ;

[0292] in, This indicates that the sound wave reaches the target direction from the first... The physical transmission delay difference between each microphone node and the reference center, which is determined by the spatial position vector. It is calculated by dividing the geometric projection distance in the direction of the sound source by the ambient sound speed. Imaginary unit, superscript This indicates the matrix transpose.

[0293] This spatial steering vector defines the acoustic propagation path characteristics of the target sound source in the frequency domain.

[0294] In step S602, the directional sound pickup and command parsing module uses the calculated spatial steering vector to perform adaptive beamforming filtering on the multi-channel audio stream.

[0295] To enhance the target speech and suppress mechanical noise from surrounding equipment, the directional sound pickup and command parsing module employs a minimum variance distortion-free response algorithm. By aligning and interfering the phase differences of sound waves in different spatial directions, it achieves spatial filtering of sound sources in specific directions.

[0296] The directional sound pickup and command parsing module extracts array audio data within the time window determined to be silent or in a faulty state in the previous steps, and calculates the spatial covariance matrix of the ambient noise signal. .

[0297] To prevent algorithmic dead zones caused by insufficient environmental noise sampling leading to non-full rank covariance matrix, the directional sound pickup and instruction parsing module introduces a diagonal loading operation on the spatial covariance matrix, that is, superimposing a constant regularization parameter on the main diagonal of the matrix.

[0298] The regularization parameter is set to a value of 1% to 5% of the estimated energy of environmental noise, in order to maintain the original spatial characteristics of the matrix while ensuring numerical stability.

[0299] After completing the matrix inversion, the directional sound pickup and command parsing module calculates the optimal weight vector in the frequency domain. The calculation formula is as follows:

[0300] ;

[0301] The inverse of the acoustic space covariance matrix, with superscript This indicates the complex conjugate transpose.

[0302] The directional sound pickup and command parsing module will calculate the obtained optimal weight vector. With the microphone array multi-channel frequency domain signal vector within the current time window Perform inner product operations to generate the enhanced single-channel target frequency domain signal. The calculation formula is as follows:

[0303] ;

[0304] Subsequently, the directional sound pickup and command parsing module uses inverse fast Fourier transform to convert the single-channel target frequency domain signal. The denoised one-dimensional speech sequence is restored to the time domain to suppress acoustic crosstalk caused by conversations of people nearby.

[0305] In step S603, the directional sound pickup and instruction parsing module performs speech recognition and business instruction mapping on the denoised one-dimensional speech sequence, and drives the sales business closed loop.

[0306] Before recognition is performed, the directional sound pickup and instruction parsing module performs frame segmentation, Hamming windowing, and short-time Fourier transform operations on the denoised one-dimensional speech sequence to extract a continuous 80-dimensional Mel filter bank acoustic feature sequence.

[0307] The directional sound pickup and command parsing module inputs the acoustic feature sequence into the system's built-in end-to-end acoustic model for decoding and outputs the corresponding text string.

[0308] For the aforementioned end-to-end acoustic model, this embodiment employs a neural network based on a convolutional enhanced Transformer architecture.

[0309] The model contains a convolutional front-end network that processes local audio features, a self-attention encoder layer that captures long sequence dependencies, and a fully connected decoder head that outputs the probability distribution of characters.

[0310] After the input features are forward propagated through the model, the output is a sequence of text characters representing the target user's purchase intent.

[0311] For the sample acquisition and training computation of this end-to-end acoustic model, those skilled in the art can use publicly available Mandarin Chinese speech datasets as training samples and optimize the network weight parameters using connectionist temporal classification loss function. The backpropagation and gradient descent optimization methods are well-known techniques in this field and will not be elaborated here.

[0312] After obtaining the text string, the directional sound pickup and instruction parsing module calls the local natural language matching rule library to extract key entity words from the text.

[0313] After clarifying the user's purchase intent (e.g., "I want a glass of orange juice"), the scheduling and execution module officially initiates the closed-loop process of the sales order:

[0314] (1) Order generation and display: The system backend generates the corresponding product order, renders and displays the payment interface containing product information and payment QR code through a multimedia display screen, and announces "Your orange juice order has been placed, please scan the code to pay";

[0315] (2) Payment verification: The system polls the cloud payment gateway interface. Once it receives an asynchronous callback signal indicating successful payment, it immediately switches the interface to the progress bar creation screen.

[0316] (3) Mechanical execution: The scheduling and execution module sends control commands to the programmable logic controller (PLC) to drive the cup-dropping module and the electromagnetic liquid outlet valve to perform the juice filling operation;

[0317] (4) Cup Retrieval Reminder: After the PLC reports that filling is complete and the safety door is opened, the system plays a cup retrieval reminder voice message: "Your juice is ready. Please take it carefully. Welcome to visit us again."

[0318] In the juicer sales scenario, key entity words typically include beverage type and quantity identifiers. The directional audio pickup and command parsing module encapsulates the extracted entity words into standard device control command packages and sends them to the juicer's main control unit. This drives the corresponding internal valves and stirring motor to perform the beverage making action, completing the execution flow from interactive commands to hardware actions.

[0319] If the natural language matching rule base fails to extract complete key entity words from the text, the directional sound pickup and instruction parsing module determines that the current instruction has ambiguity or recognition deficiencies, and sends a voice retry prompt instruction to the main control unit to trigger the device to play interactive voice prompts to guide the user to speak again.

[0320] In step S700, the system executes the edge network outage fault tolerance and local fallback takeover mechanism. The edge network outage fallback module sends network heartbeat packets to the cloud server in the background at a set frequency (e.g., once every 3 seconds).

[0321] If the system fails to receive a heartbeat confirmation from the cloud three times in a row or detects that the network latency exceeds the preset timeout threshold, it determines that a network outage has occurred.

[0322] At this time, the edge network disconnection backup module takes over the main control: (1) Pause the end-to-end acoustic model call in the cloud and seamlessly switch to the lightweight offline acoustic model (such as a pre-tailed lightweight hidden Markov model or a small parameter Transformer) deployed in the edge computing node (such as NPU), which only retains the recognition ability of high-frequency juice words;

[0323] (2) Suspend the cloud-based dynamic marketing module and use the pre-built basic script library in the local storage space instead (e.g., "The current network is poor, only basic ordering is supported").

[0324] (3) During the payment process, switch to the local area network offline payment mode or the temporary order mode, and perform asynchronous batch settlement after the network is restored.

[0325] This fallback strategy ensures that the juicer can maintain its core ordering and juice dispensing processes even in a network-free environment, avoiding business losses caused by equipment downtime.

[0326] Figure 9 This is a dynamic time warping alignment and coherence mapping diagram of the bimodal sequence in this application embodiment.

[0327] Scene setting:

[0328] This juicer was placed in a noisy shopping mall (ambient background noise level approximately 75dB).

[0329] When target user A passes by the device at a distance of 2.5 meters, the system automatically triggers a welcoming voice message: "Hello, would you like a glass of freshly squeezed juice?"

[0330] User A is attracted and stays directly in front of the device (azimuth angle in coordinate system). Pitch angle As you prepare to buy a drink, a passerby B (from the azimuth angle) appears beside you. (He / She) is making a loud phone call.

[0331] The system is set to a video frame rate of 15 FPS and an audio sampling rate of 16000 Hz.

[0332] Slide the time window size A frame (i.e., 1 second of physical time).

[0333] Steps S100-S200: Audio Feature Extraction

[0334] 1. Data Extraction: The number of audio sampling points corresponding to each frame of video is... .

[0335] 2. Envelope Construction: After detecting speech activity (a mixture of the voices of passerby B and user A), the envelope is constructed... Each audio frame is subjected to a Hamming window and its short-time energy is calculated. Assuming that after normalization, a one-dimensional audio energy envelope discrete-time sequence is generated. (Length is 15) as follows (showing a trend of being higher in the middle and lower at both ends):

[0336] [0.10,0.15,0.40,0.85,0.90,0.70,0.45,0.20,0.12,0.10,0.10,0.08,0.05,0.05,0.05].

[0337] Step S300: Pose and Spindle Construction

[0338] 1. The visual model captures the facial nodes of user A and extracts the coordinates of the nose tip (320, 200) and the neck (320, 350).

[0339] 2. Calculate the initial direction vector .

[0340] 3. Normalize to generate anatomical principal axis vectors (This represents the reference axis of the vocal muscles moving vertically upwards from the neck).

[0341] Step S400: Optical Flow Directional Projection and Filtering

[0342] 1. Extract the lower half of user A's bounding box (jaw region) and calculate the optical flow field of adjacent frames (14 pairs in total).

[0343] 2. Interchange the optical flow vector of each pixel with... Dot product and non-negative truncation are performed to filter out lateral interference from user A's left and right head movements, extract the longitudinal vocalization action, and generate a length of [length missing] after spatial integration. One-dimensional visual motion purification envelope sequence;

[0344] .

[0345] 3. Calculate the sequence Statistical variance .

[0346] Assuming the noise floor variance threshold is 0.015, since 0.068 > 0.015, the visual envelope validity status flag of user A is set to valid.

[0347] (Note: Passerby B's variance is below the threshold due to his profile or being too far away, so he is directly blocked by the gating module and does not enter the subsequent calculation.)

[0348] Step S500: DTW Dynamic Time Warping and Coherence Calculation (corresponding to...) Figure 9 )

[0349] 1. Truncate the audio sequence The first frame yields a length of 14. .

[0350] 2. Regarding and Perform Z-score standardization separately.

[0351] 3. Apply a window with Sakohe-Chiba constraints ( The DTW algorithm is used to find the optimal homogeneous path.

[0352] Due to user A's mouth movements (sequence) ) and the actual emitted sound energy (sequence) The high degree of matching results in a very small cumulative minimum path distance. (Assuming...) .

[0353] 4. Mapping coherence score: Set attenuation coefficient ,but .

[0354] 5. Assuming the judgment threshold is 0.6, since The system successfully identified user A as the real source of the interactive sound.

[0355] Step S600: Beamforming and Command Parsing

[0356] 1. Utilize user A's visual coordinates to transform the spatial guidance vector, targeting the azimuth angle. Pitch angle This forms the main pickup lobe.

[0357] 2. Based on the environmental covariance matrix extracted during the silent period, perform MVDR adaptive beamforming in the area where pedestrian B is located. The direction forms a spatial null.

[0358] 3. The noise-reduced audio input acoustic model identifies the text "A glass of orange juice".

[0359] The instruction parsing module extracts the number of entity words: "one cup" (category: orange juice), and then the system screen generates an order and renders a payment QR code.

[0360] After the user successfully pays by scanning the code, the edge node sends a PLC to drive the solenoid valve to dispense the juice. Once completed, a voice announcement is made saying "The orange juice is ready, please take your cup," thus completing the full transaction loop.

[0361] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A juice vending system based on AI vision and voice interaction, characterized in that, include: The pose perception and active greeting module is configured to extract the facial node coordinates and bounding box size of candidate targets in video frames, generate an interaction priority queue, calculate physical distance based on visual features, and when it is determined that the candidate target enters the preset greeting distance range of 1.5 meters to 3 meters, the active greeting voice broadcast is triggered first, and the two-dimensional normalized anatomical principal axis vector of the target in the interaction priority queue is calculated at the same time. The dynamic shopping guide and marketing module is configured to dynamically generate multiple rounds of voice shopping guide scripts and member promotion strategies based on the current system time period, the number of target people captured in the same frame by visual capture, and the preset promotional activities issued by the cloud after triggering the active greeting, so as to guide users to initiate voice interaction; The data synchronization acquisition module is configured to perform time-series alignment of the acquired audio stream and video frames based on the hardware timestamp when the user initiates a voice response to the sales guide script, and construct a sliding time window containing multiple frames of data. The audio envelope construction module is configured to perform speech activity detection and short-time energy calculation on the audio data within the sliding time window to generate a one-dimensional audio energy envelope discrete time series. The optical flow filtering and gating module is configured to calculate the two-dimensional dense optical flow vector within the local region of the target bounding box, and perform dot product projection and spatial integration on the two-dimensional dense optical flow vector and the two-dimensional normalized anatomical principal axis vector to generate a one-dimensional visual motion purification envelope sequence. The coherence calculation module is configured to use a dynamic time warping algorithm to perform path mapping alignment between the discrete time sequence of the one-dimensional audio energy envelope of the target and the one-dimensional visual motion purification envelope sequence, and extract a coherence score based on the warping distance to lock the real sound source target. The scheduling and execution module is configured to trigger the corresponding feature intent binding based on the comparison result of the coherence score and the preset coherence threshold, parse the voice sequence of the sound source into a purchase instruction or receive the user's touch screen order operation, and then sequentially execute the order generation, payment QR code rendering and display, payment status verification, issuing control instructions for the juicer's electromagnetic dispensing valve, and the cup removal voice reminder process after the cup is placed. The edge network disconnection fallback module is configured to globally monitor the network heartbeat connection between the edge smart terminal and the cloud server. When a network disconnection is detected, it takes over the main control, calls the locally pre-built dialogue library and lightweight offline acoustic model, and maintains basic offline voice interaction, local cached order placement and local area network payment processes.

2. The juice vending system based on AI vision and voice interaction according to claim 1, characterized in that, The audio envelope construction module is specifically configured as follows: under the condition that the audio data contains speech activity, the audio sampling points are segmented and truncated according to the video frame rate within the sliding time window to construct an audio frame sequence. A Hamming window is applied to the audio frame sequence and the short-time energy envelope value is calculated; Normalization is performed on the short-time energy envelope values ​​to output the one-dimensional audio energy envelope discrete time series.

3. The juice vending system based on AI vision and voice interaction according to claim 1, characterized in that, The pose and spindle construction module is specifically configured as follows: Obtain the coordinates of the nose tip node and the neck node from the facial node coordinates; Construct an initial two-dimensional direction vector pointing from the coordinates of the neck node to the coordinates of the nose tip node; The initial two-dimensional direction vector is normalized to generate the two-dimensional normalized anatomical principal axis vector.

4. The juice vending system based on AI vision and voice interaction according to claim 1, characterized in that, The optical flow filtering and gating module is specifically configured as follows when performing dot product projection and spatial integration on the two-dimensional dense optical flow vector and the two-dimensional normalized anatomical principal axis vector to generate a one-dimensional visual motion purification envelope sequence: The lower half of the target bounding box is selected as the analysis region, and the pixel-level two-dimensional dense optical flow vector between adjacent video frames within the analysis region is calculated frame by frame. Perform a dot product operation between the pixel-level two-dimensional dense optical flow vector and the two-dimensional normalized anatomical principal axis vector to output a projection scalar. Perform a non-negative truncation operation on the projected scalar, setting projected scalars with values ​​less than zero to zero; The one-dimensional visual motion purification envelope sequence is constructed by performing spatial integration and summation on the projection scalar after non-negative truncation in pixel space.

5. A juice vending system based on AI vision and voice interaction according to claim 4, characterized in that, The optical flow filtering and gating module is also configured to: Calculate the statistical variance of the one-dimensional visual motion purification envelope sequence; When the statistical variance is greater than the preset noise floor variance threshold, the visual envelope validity status flag of the corresponding target is set to valid.

6. A juice vending system based on AI vision and voice interaction according to claim 1, characterized in that, The coherence calculation module is specifically configured as follows: The one-dimensional audio energy envelope discrete time sequence and the one-dimensional visual motion purification envelope sequence are respectively subjected to standardization processing to generate a normalized audio sequence and a normalized visual sequence. Establish a cumulative distance matrix and set a global constraint window width during the state transition calculation; for matrix elements whose absolute difference in coordinate indices is greater than the global constraint window width, set their state transition cost to infinity; Extract the endpoint coordinate element values ​​from the cumulative distance matrix as the cumulative minimum path distance; The cumulative minimum path distance is mapped using an exponential decay function, and the coherence score is output.

7. A juice vending system based on AI vision and voice interaction according to claim 1, characterized in that, The scheduling and execution module is specifically configured as follows: If the coherence score is greater than a preset coherence threshold, the corresponding target is set as the sound source target. The pixel coordinates of the sound source target are back-projected into a three-dimensional spatial direction ray, and the azimuth and pitch parameters are calculated to construct a spatial guidance vector. Adaptive beamforming filtering is performed on the acquired multi-channel frequency domain signal of the microphone array based on the spatial steering vector to output a single-channel target frequency domain signal; The single-channel target frequency domain signal is converted into a time domain audio sequence for speech recognition and text parsing to generate the control command.

8. A juice vending system based on AI vision and voice interaction according to claim 1, characterized in that, The system also includes an edge intelligent terminal, a microphone pickup array, a visual camera peripheral, a lower-level machine execution mechanism, and a multimedia display screen; The microphone array and the visual camera peripheral are respectively communicatively connected to the edge intelligent terminal, and are used to provide the audio stream and the video frame to the system; The data synchronization acquisition module, the audio envelope construction module, the pose and main axis construction module, the optical flow filtering and gating module, the coherence calculation module, the dynamic shopping guide and marketing module, the scheduling and execution module, and the edge network outage backup module are all integrated and deployed in the edge intelligent terminal.

9. A juice vending system based on AI vision and voice interaction according to claim 8, characterized in that, The lower-level machine actuator includes a programmable logic controller and an electromagnetic liquid outlet valve; The edge intelligent terminal is configured to receive control commands generated by the scheduling and execution module and send the control commands to the programmable logic controller to drive the electromagnetic dispensing valve to perform juice filling and dispensing operations.

10. A method for selling juice machines based on AI vision and voice interaction, comprising a juice machine selling system based on AI vision and voice interaction as described in any one of claims 1-9, characterized in that... include: Extract the facial node coordinates and dimensions of candidate targets from the video source, calculate the physical distance, trigger an active welcoming broadcast when the candidate target is determined to be within the distance range of 1.5 meters to 3 meters, and calculate the two-dimensional normalized anatomical principal axis vector of the target; Based on the current time period, the target number of people identified visually, and the preset promotional activities, the dynamic shopping guide and marketing module is invoked to generate multiple rounds of voice shopping guide scripts to guide the front-end customer flow; After receiving the user's voice response, the acquired audio stream and the corresponding video frame are time-aligned based on the hardware timestamp to construct a sliding time window containing multiple frames of data; voice activity detection and short-time energy calculation are performed on the audio data within the sliding time window to generate a one-dimensional audio energy envelope discrete time series. Calculate the two-dimensional dense optical flow vector within the local region of the target bounding box, and perform dot product projection and spatial integration on the two-dimensional dense optical flow vector and the two-dimensional normalized anatomical principal axis vector to generate a one-dimensional visual motion purification envelope sequence. A dynamic time warping algorithm is used to perform path mapping alignment between the discrete time sequence of the one-dimensional audio energy envelope of the target and the one-dimensional visual motion purification envelope sequence, and to extract the coherence score based on the warping distance to lock the sound source target. Based on the coherence score locking result, extract pure voice and parse the purchase instruction, or receive screen touch instruction, and sequentially execute the order generation, payment QR code rendering and display, payment status verification, issue juicer control instruction to drive cup making, and the process of taking the cup after making is completed and voice reminder. The system monitors network connectivity in real time in the background. If a network outage occurs, it triggers the edge network outage fallback module, which calls the locally pre-set dialogue scripts and offline acoustic models to take over the basic interaction and offline transaction process.