Intelligent scheduling method and system for vending machine based on artificial intelligence

By collecting and analyzing non-semantic physiological and behavioral signals of vending machine users, and combining them with time series analysis models, advertising and promotional strategies can be dynamically adjusted. This solves the problem of difficulty in capturing changes in user decision-making in existing technologies, and achieves a higher user experience and operational efficiency.

CN121094468BActive Publication Date: 2026-07-21DONGGUAN JIAFENG MECHANICAL EQUIP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DONGGUAN JIAFENG MECHANICAL EQUIP
Filing Date
2025-09-12
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

The existing intelligent scheduling technology of vending machines is unable to reflect the dynamic changes in user decision-making in real time. The matching degree between the scheduling strategy and the user's current attention distribution and immediate decision-making state is insufficient, resulting in the failure to fully realize the user shopping experience and operational efficiency.

Method used

Non-semantic physiological and behavioral signals of users interacting with vending machines are collected by non-optical sensing units. Entropy values ​​and instantaneous micro-dynamic responses are calculated, and dynamic attention weight vectors are generated by combining time series analysis models to dynamically adjust advertising content and promotional strategies to improve matching accuracy.

Benefits of technology

It achieves a high degree of adaptation between the vending machine scheduling strategy and the user's real-time needs, improves the user shopping experience and operational efficiency, reduces reliance on historical data, and enhances the real-time performance and accuracy of the scheduling strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121094468B_ABST
    Figure CN121094468B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent scheduling method and system of an automatic vending machine based on artificial intelligence, and relates to the technical field of data processing.The method comprises the following steps: collecting non-semantic physiological and behavioral signals generated when a user interacts with the automatic vending machine through a non-optical sensing unit; calculating an entropy value index representing the uncertainty of the decision state of the user based on the non-semantic physiological and behavioral signals; capturing the instantaneous micro-dynamic response of the non-semantic physiological and behavioral signals under the action of information micro-stimulation; inputting the entropy value index and the instantaneous micro-dynamic response into a trained time series analysis model to generate a dynamic attention weight vector used to describe the current preference degree of the user for each commodity attribute; calculating a comprehensive semantic vector fused by the semantic information of all current schedulable resources in real time based on the dynamic attention weight vector; and dynamically adjusting the advertisement content, the user interface element or the promotion strategy according to the matching result of the dynamic attention weight vector and the comprehensive semantic vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically, to an intelligent scheduling method and system for vending machines based on artificial intelligence. Background Technology

[0002] Against the backdrop of the rapid development of new retail, vending machines, as convenient retail terminals, are widely distributed in various scenarios such as public places and office areas. Their level of intelligent scheduling directly affects the user shopping experience and operational efficiency. Through intelligent scheduling, it is possible to match advertising pushes with user needs, which can not only increase product sales but also reduce operating costs and optimize resource allocation. Therefore, it has become an important direction for the intelligent upgrading of the retail sector.

[0003] However, existing intelligent scheduling technology for vending machines still has certain limitations in practical applications. On the one hand, capturing users' real needs relies heavily on historical purchase data or explicit feedback, making it difficult to reflect the dynamic changes in users' decisions during the interaction process in real time. On the other hand, the adjustment of scheduling strategies is often based on preset rules or static preference models, which do not match the user's current attention distribution and immediate decision-making state well enough, resulting in the scheduling effect not being fully realized.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This application provides an intelligent scheduling method and system for vending machines based on artificial intelligence to solve the above-mentioned technical problems.

[0006] This application provides an intelligent scheduling method for vending machines based on artificial intelligence, comprising: collecting non-semantic physiological and behavioral signals generated when a user interacts with a vending machine through a non-optical sensing unit; calculating an entropy index characterizing the uncertainty of the user's decision-making state based on the non-semantic physiological and behavioral signals; applying one or more information micro-stimuli below a perception threshold to the user and capturing the instantaneous micro-dynamic response of the non-semantic physiological and behavioral signals under the action of the information micro-stimuli; inputting the entropy index and the instantaneous micro-dynamic response into a trained time series analysis model to generate a dynamic attention weight vector describing the user's current preference for each product attribute; calculating a comprehensive semantic vector in real time, which is formed by fusing semantic information from all currently schedulable resources, based on the dynamic attention weight vector; and dynamically adjusting advertising content, user interface elements, or promotional strategies according to the matching result of the dynamic attention weight vector and the comprehensive semantic vector to improve the matching degree between the dynamic attention weight vector and the comprehensive semantic vector.

[0007] This application provides an AI-based intelligent scheduling system for vending machines, comprising: a signal acquisition module for acquiring non-semantic physiological and behavioral signals generated when a user interacts with the vending machine via a non-optical sensing unit; an entropy index calculation module for calculating an entropy index characterizing the uncertainty of the user's decision-making state based on the non-semantic physiological and behavioral signals; a micro-dynamic response capture module for applying one or more information micro-stimuli below a perception threshold to the user and capturing the instantaneous micro-dynamic responses of the non-semantic physiological and behavioral signals under the influence of the information micro-stimuli; an attention weight vector generation module for inputting the entropy index and the instantaneous micro-dynamic responses into a trained time series analysis model to generate a dynamic attention weight vector describing the user's current preference for various product attributes; a comprehensive semantic vector calculation module for calculating, in real time, a comprehensive semantic vector formed by fusing semantic information from all currently schedulable resources based on the dynamic attention weight vector; and an adjustment module for dynamically adjusting advertising content, user interface elements, or promotional strategies according to the matching results of the dynamic attention weight vector and the comprehensive semantic vector, to improve the matching degree between the dynamic attention weight vector and the comprehensive semantic vector.

[0008] Based on the embodiments provided in this application, non-semantic physiological and behavioral signals of users interacting with vending machines are collected by non-optical sensing units. Combined with entropy indicators and instantaneous micro-dynamic responses under the influence of micro-stimuli, this approach can more accurately capture the user's dynamic decision-making state during the interaction process. It eliminates excessive reliance on historical purchase data or explicit feedback and can reflect the uncertainty of the user's current decision-making state in real time, thus providing a more accurate basis for generating dynamic attention weight vectors that fit the user's immediate state. On this basis, a comprehensive semantic vector is calculated based on the dynamic attention weight vector, and the advertising content, user interface elements, or promotional strategies are dynamically adjusted according to the matching results. This ensures that the scheduling strategy is highly adapted to the user's current attention distribution and product attribute preferences, effectively improving the matching degree between the scheduling strategy and the user's immediate needs. This, in turn, better leverages intelligent scheduling to enhance the user's shopping experience and operational efficiency. Attached Figure Description

[0009] The accompanying drawings, which are included to provide a further understanding of embodiments of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0010] Figure 1 This is a flowchart of an optional AI-based intelligent scheduling method for vending machines according to an embodiment of this application;

[0011] Figure 2A flowchart illustrating another optional intelligent scheduling method for vending machines based on artificial intelligence, according to an embodiment of this application;

[0012] Figure 3 This is a structural diagram of an optional AI-based intelligent scheduling system for vending machines according to an embodiment of this application;

[0013] Figure 4 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of this application.

[0014] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0016] According to one aspect of the embodiments of this application, such as Figure 1 As shown, this application provides an intelligent scheduling method for vending machines based on artificial intelligence, including:

[0017] S101 collects non-semantic physiological and behavioral signals generated when a user interacts with a vending machine through a non-optical sensing unit;

[0018] In S101, non-semantic physiological and behavioral signals are collected through non-optical sensing units, avoiding privacy concerns and lighting limitations that may arise from optical sensing. Furthermore, the selection of non-semantic signals (such as physiological reactions and unconscious actions) more realistically reflects the potential state of the user that is not expressed through language or explicit actions. This design provides objective and continuous raw data for subsequent analysis of user decision-making states, laying a reliable perceptual foundation for the entire intelligent scheduling method and ensuring the capture of information that is difficult to convey through explicit feedback during user interaction.

[0019] S102, based on non-semantic physiological and behavioral signals, calculates the entropy index that represents the uncertainty of the user's decision-making state;

[0020] In S102, entropy is used to measure uncertainty, transforming the user's complex decision-making state into a quantifiable indicator. The level of entropy directly reflects the degree of hesitation or clarity in a user's decision, solving the problem that the user's internal decision-making state is difficult to directly observe and quantify. This indicator provides a crucial basis for subsequently judging the user's decision-making stage and understanding their demand tendencies, enabling the analysis of user states to move from qualitative to quantitative.

[0021] S103, apply one or more information microstimuli below the perception threshold to the user, and capture the instantaneous micro-dynamic response of non-semantic physiological and behavioral signals under the action of information microstimuli;

[0022] It should be noted that, based on the subconscious processing theory of cognitive neuroscience, the human brain still processes stimuli that are not consciously perceived (below the perception threshold) subconsciously, triggering subtle physiological responses.

[0023] For example, applying a 50ms visual microstimulus to a user (shorter than the 80ms visual persistence threshold, imperceptible to the user) may subconsciously create a latent association with the stimulus (such as an association with the color of a certain type of product), thereby triggering subtle fluctuations in heart rate variability or microtremors in the hands. The value of this "stimulus-response" causal chain lies in the fact that, without the user's active cooperation, it is possible to capture their latent preferences for product attributes by observing micro-dynamic responses (e.g., a latent preference for sweet products may be reflected in changes in respiratory rhythm under specific microstimuli).

[0024] In S103, microstimuli go unnoticed by the user, avoiding interference with their normal decision-making process, while triggering subconscious physiological and behavioral responses. This design breaks through the limitations of relying on explicit user feedback. By capturing subtle responses to microstimuli, it can uncover latent preferences and reaction patterns that users haven't actively expressed, providing a unique source of information for a deeper understanding of users' true needs and enriching the dimensions of user state analysis.

[0025] S104, input the entropy index and the instantaneous micro-dynamic response into the trained time series analysis model to generate a dynamic attention weight vector that describes the user's current preference for each product attribute;

[0026] In S104, indicators reflecting user decision-making uncertainty and subconscious responses to microstimuli are integrated. A trained model is used to deeply fuse these two types of information, achieving a transformation from raw signals to a quantitative expression of user preferences. The dynamic attention weight vector can reflect the user's level of attention to various product attributes in real time, solving the problem that traditional methods struggle to capture dynamic user preferences accurately and in real time, and providing direct target guidance for subsequent resource scheduling.

[0027] S105, based on dynamic attention weight vector, calculates in real time the comprehensive semantic vector formed by fusing the semantic information of all currently schedulable resources;

[0028] In S105, the schedulable resources of the vending machine (such as advertisements and interface elements) are semantically processed and integrated, so that the resource information on the machine is on the same semantic dimension as the user's attention preferences, providing a comparable basis for matching the two. The comprehensive semantic vector reflects the overall information environment of the current resources in real time, ensuring that resource scheduling can closely follow the user's current attention distribution, making the scheduling clear and adaptive.

[0029] S106, Based on the matching results of the dynamic attention weight vector and the comprehensive semantic vector, dynamically adjust the advertising content, user interface elements or promotional strategies to improve the matching degree between the dynamic attention weight vector and the comprehensive semantic vector.

[0030] It's important to explain that improving the matching degree between the dynamic attention weight vector (user's current preferences) and the comprehensive semantic vector (schedulable resource information) is fundamentally about reducing "cognitive friction" during user decision-making. Specifically, when the matching degree is low, users must sift through irrelevant information to find their target (e.g., users interested in low prices are forced to view ads for high-priced goods), which can easily lead to cognitive conflict and decision fatigue, potentially causing them to abandon their purchase. Conversely, when the matching degree is high (e.g., users interested in health see low-sugar products prioritized), users can quickly identify information that meets their needs, simplifying the decision-making process and increasing the probability of a transaction. This design allows intelligent scheduling to directly serve the business goal of "facilitating transactions and optimizing the user experience."

[0031] In S106, the goal is to improve the matching degree between the two, enabling the scheduling strategy to respond in real time to the user's dynamically changing preferences. This dynamic adjustment avoids the limitations of fixed strategies, and by continuously optimizing advertising, interface, or promotional methods, it ensures that the information received by the user is highly consistent with their own needs, thereby improving the user's shopping experience and enhancing the operational effectiveness of the vending machine.

[0032] Furthermore, the non-optical sensing unit is a millimeter-wave radar; the non-semantic physiological and behavioral signals include the chest cavity undulation waveform detected by the millimeter-wave radar and preprocessed, continuous point cloud data of hand movement trajectory, and frequency domain characteristics of body surface microtremors.

[0033] It's important to clarify that the core reasons for avoiding "semantic" signals (such as images and voice) are twofold: First, to avoid privacy controversies. Semantic signals (such as user facial images and voice commands) are directly linked to personal identity information, which can easily raise privacy concerns in public settings like vending machines. Non-semantic signals (such as slight hand tremors or chest movements) only reflect physiological or behavioral characteristics and do not involve identity recognition, thus reducing privacy risks at the source. Second, to capture subconscious reactions. A user's semantic expressions (such as voice choices) are often the result of rational thinking and may mask true preferences. Non-semantic signals (such as disordered hand movements during hesitation) are more driven by the subconscious and can reflect a user's unexpressed decision-making tendencies (for example, a user may say "just browsing," but their hand movements repeatedly lingering in front of a certain type of product can be captured through non-semantic signals).

[0034] In practical implementation, compared to non-optical sensors such as infrared and ultrasound, millimeter-wave radar has irreplaceable advantages in vending machine scenarios. Specifically, millimeter waves (such as the 24GHz or 77GHz bands) can penetrate thin obstructions such as clothing and plastic bags, accurately detecting chest rise and fall (breathing, heart rate) and hand movements, while infrared is easily affected by ambient temperature, and ultrasound has low resolution (difficult to capture micro-tremors). It only outputs non-image data such as point clouds and waveforms, without involving user facial information, avoiding privacy disputes that may arise from infrared imaging. Its performance is stable in complex environments such as strong light, darkness, and fog, making it suitable for deployment in vending machines in various scenarios such as shopping malls and outdoor locations.

[0035] The preprocessing includes filtering the original radar reflection signal to remove environmental noise interference, and separating the signal components related to vital signs and specific behaviors through digital signal processing algorithms.

[0036] In some embodiments, the steps for converting the raw radar I / Q signal (in-phase / quadrature signal) into a target signal are as follows:

[0037] Filtering and denoising: Environmental electromagnetic interference is removed by bandpass filtering (preserving the 0.1-10Hz frequency band to cover breathing and motion signals), and then high-frequency noise is smoothed by Kalman filtering.

[0038] Target separation: Clustering algorithms (such as DBSCAN) are used to distinguish human body regions from background (such as shelves and the ground) from the original point cloud, and the point cloud subsets corresponding to the chest cavity and hands are extracted.

[0039] Feature transformation: Time-domain analysis of chest cavity point clouds is performed to extract periodic fluctuation waveforms (respiration, heart rate); trajectory reconstruction of hand point clouds is performed (through inter-frame matching) to generate continuous motion trajectories; Fourier transform of body surface microtremor signals is performed to obtain frequency domain features (such as tremor dominant frequency and energy distribution).

[0040] Based on non-semantic physiological and behavioral signals, an entropy index representing the uncertainty of a user's decision-making state is calculated, including:

[0041] Trajectory reconstruction and smoothing are performed on continuous point cloud data of hand motion trajectory to extract the motion change sequence of its speed and direction;

[0042] A windowed Fourier transform is performed on the motion change sequence to obtain the time-frequency spectrum.

[0043] In practical implementation, the processing of continuous point cloud data for hand movement trajectories begins with trajectory reconstruction. Millimeter-wave radar continuously collects point cloud data of the hand. While this data is temporally continuous, individual frames may contain noise or have sparse point clouds. During trajectory reconstruction, a Euclidean distance clustering algorithm is first used to filter out a subset of point clouds belonging to the hand from each frame (excluding background point clouds such as vending machine cabinets and merchandise). Next, Kalman filtering is used for inter-frame point cloud matching. Based on the hand's position, velocity, and other states from the previous moment, the possible position at the current moment is predicted. This prediction is then combined with the point cloud data from the current frame for correction, thus connecting the discrete point cloud data into a continuous hand movement trajectory.

[0044] Next, a smoothing process is performed. Since hand movements may contain slight tremors or noise from radar data, the trajectory may exhibit jagged edges. Here, a moving average filter is used. A sliding window containing 5 frames of point cloud data is set, and the hand position coordinates within the window are averaged to smooth out high-frequency noise in the trajectory, making the trajectory more closely match the actual movement trend of the hand.

[0045] In extracting the motion change sequence of velocity and direction, based on the reconstructed and smoothed trajectory, the displacement of the hand between two adjacent frames is calculated. This displacement is then combined with the time interval between the two frames (the radar sampling interval, such as 10ms) to obtain the velocity corresponding to each frame (displacement divided by the time interval). Simultaneously, the direction of the hand movement is determined by the direction of the displacement vector (e.g., with 0 degrees to the right horizontally, the angle is calculated clockwise). Arranging the velocities and directions at different times in chronological order forms the motion change sequence.

[0046] For windowed Fourier transform of motion sequences, the Hanning window is chosen as the window function because it offers a good balance between frequency and time resolution, making it suitable for processing signals like hand movements that may involve variable frequency motion. The window size is set to 1 second, corresponding to 100 frames of data (based on a 10ms sampling interval). The window overlap rate is 50%, meaning each subsequent window overlaps with the previous one by 50 frames to ensure temporal continuity and avoid information loss.

[0047] A Fourier transform is performed on the motion sequence (including changes in speed and direction over time) within each window, converting the time-domain signal into a frequency-domain signal, thus obtaining the spectrum corresponding to that window. As the window slides successively over the motion sequence (50 frames at a time), a series of spectra are obtained. Arranging these spectra in chronological order forms a time-spectrum graph. The horizontal axis of the time-spectrum graph represents time, and the vertical axis represents frequency. Color or grayscale indicates the energy intensity of the corresponding frequency component, visually demonstrating the changes in hand movement speed and direction at different times and frequencies, providing a foundation for subsequent calculations of Shannon entropy.

[0048] Calculate the Shannon entropy value of the time spectrum within each time window, and use the curve of the change of Shannon entropy value over time as the entropy index. An increase in Shannon entropy value indicates an increase in decision-making hesitation, while a decrease in Shannon entropy value indicates that the decision tends to be clear.

[0049] In some embodiments, the selection criteria for Shannon entropy calculation parameters for hand movement change sequences include:

[0050] Time window size: set to 1 second (balancing time resolution and stability; too short a window is easily affected by noise, and too long a window cannot capture instantaneous changes).

[0051] Overlap rate: 50% (to ensure continuous information between adjacent windows and avoid truncation of key changes).

[0052] Frequency range: 0.5-5Hz (covering the main frequencies of fine human hand movements, excluding high-frequency noise and low-frequency baseline drift).

[0053] During the calculation, the probability distribution of the time-frequency spectrum of each window is first estimated, and then the entropy value is calculated using the Shannon entropy formula, forming an entropy curve that changes over time. The Shannon entropy formula is a well-known technique and will not be elaborated upon in this embodiment.

[0054] It should be explained that, from an information theory perspective, entropy is an indicator of the degree of disorder in a system.

[0055] When a user's decision-making state is uncertain (hesitant), their physiological and behavioral signals exhibit "high disorder." For example, when a hand moves back and forth between multiple products, the speed and direction of the movement change frequently, resulting in numerous and chaotic possible states of the signal. At this time, the entropy value increases, intuitively quantifying the "degree of hesitation." When a user's decision tends to be clear, the signal exhibits "low disorder." For example, when a hand steadily points to a certain product, the movement trajectory is regular, and the possible states of the signal are concentrated. At this time, the entropy value decreases, quantifying the "decision certainty."

[0056] This connection transforms the abstract "decision state" into a calculable numerical indicator, providing an objective basis for subsequent scheduling.

[0057] Based on the embodiments provided in this application, millimeter-wave radar is used as a non-optical sensing unit. This allows for the accurate acquisition of non-semantic physiological and behavioral signals such as chest cavity undulation waveforms, continuous point cloud data of hand movement trajectories, and frequency domain characteristics of micro-tremors on the body surface, without relying on optical means. This avoids the potential impact of lighting conditions or user privacy concerns associated with optical sensing. By processing the continuous point cloud data to obtain a motion change sequence, and then performing a windowed Fourier transform to obtain a time-spectrum graph and calculating the Shannon entropy value, this entropy index cleverly transforms the complex changes in user hand movements into a quantifiable decision-making state indicator. The rise and fall of the Shannon entropy value can intuitively reflect the state of decision-making hesitation and clarity, providing a reliable basis for subsequently capturing user dynamic decisions and improving the accuracy of judging the user's decision-making state.

[0058] Furthermore, information microstimuli below the perception threshold include any one or more of the following stimuli:

[0059] Visual elements of a preset color or shape, encoded, are displayed in a specific area of ​​the display unit for a preset duration; wherein the preset duration is shorter than the time threshold required for visual persistence.

[0060] In practical implementation, the core design principle of preset colors or shapes is to ensure that, under the premise of being "below the perception threshold," they can be quickly associated with product attributes subconsciously, while avoiding triggering the user's conscious attention. The specific selection logic is as follows:

[0061] Preset colors: Prioritize basic colors (not highly saturated colors, to avoid visual impact) that are strongly related to the attributes of frequently purchased items in the vending machine. For example, light cyan corresponds to the attribute of "low sugar" (subconsciously associated with health and refreshment); dark orange corresponds to "energy drinks" (implicitly associated with vitality and calories).

[0062] The brightness and saturation of the colors are controlled at low levels (e.g., brightness 50-60, saturation 30-40, based on the CIE color system) to ensure that when the colors flash for a short time (e.g., 30-50ms), the user's visual system can receive the light signal, but the conscious level cannot recognize the specific meaning of the colors, only triggering the subconscious attribute association.

[0063] Preset shapes: Employ minimalist geometric shapes to avoid conscious recognition due to complex outlines. For example, a slender rectangle for "bottled" products (simplifying the bottle outline); a perfect circle for "canned" products (simplifying the can outline). The lines of the shapes are of uniform thickness (e.g., 2-3 pixels), and the area is controlled to a small area in the corner of the display unit (e.g., 10×10 pixels) to ensure that they are not actively noticed by the user when they flash, serving only as a subconscious cues of the product's form.

[0064] The time threshold required for visual persistence of perception can be set to 50ms, with a preset duration of <50ms. The human visual persistence of perception threshold is approximately 80 to 100ms; flashes shorter than this duration cannot be consciously captured.

[0065] Mechanical vibrations with frequencies below the lower limit of human hearing are generated by a piezoelectric ceramic actuator coupled to the glass panel of the vending machine.

[0066] Among them, the lower limit of the frequency that the human ear can hear can be set to 20Hz. Vibrations below this frequency can only be perceived by the body, but when the amplitude of the piezoelectric ceramic actuator is <0.1mm, the body cannot detect it either.

[0067] The sound energy of the audio signal is focused into a preset space area in front of the vending machine using a directional acoustic generator. The center frequency of the audio signal is pre-selected in a frequency band that is offset from the main energy frequency band of the ambient background noise, and the sound pressure level of the audio signal is set to be lower than the sound pressure level of the ambient background noise at the center frequency band.

[0068] The directional acoustic generator achieves sound energy focusing through phased array acoustic wave superposition technology, specifically as follows:

[0069] Install a linear array of 8-12 miniature speakers (1-2cm in diameter) on the inner front side of the vending machine (near the user interaction area). The spacing between the speakers is an integer multiple of half the wavelength of the sound wave (e.g., for a center frequency of 3kHz, the spacing is about 5.7cm, based on the speed of sound in air of 343m / s).

[0070] The input audio signal (e.g., 1-5kHz frequency band) is phase-modulated, and the output phase and amplitude of each speaker are individually adjusted by an array controller. Within a preset spatial area, the sound waves emitted by each speaker are superimposed and enhanced due to phase consistency (sound pressure level increases by 6-10dB); outside the area, the sound waves are attenuated due to phase cancellation (sound pressure level decreases by 15-20dB), thereby achieving directional focusing of sound energy.

[0071] Built-in environmental noise sensor collects the noise spectrum around the vending machine in real time, dynamically adjusts the frequency (avoiding the main noise frequency) and amplitude of the audio signal (ensuring that the signal in the focused area is below the perception threshold, but the signal-to-noise ratio is ≥3dB), and ensures that the sound waves can be effectively transmitted to the target area.

[0072] The preset space area is a target area defined based on the typical interaction posture of the user and the vending machine. The specific range is as follows: with the center point of the vending machine's dispensing port as the origin, the horizontal direction (left and right) is ±30cm, the vertical direction (up and down) is 120-160cm (corresponding to the height from an adult's chest to their head), and the distance from the front of the machine is 30-80cm (the distance for a user to interact normally while standing).

[0073] Interaction data from 100 users was collected using millimeter-wave radar. The results showed that when users selected products, the vertical distance between their body's midline and the front of the machine was concentrated between 40-60cm, and their head and upper body were within a height range of 120-160cm, with lateral movement not exceeding ±25cm. Therefore, a preset area covering 90% confidence interval of this statistical range was designed to ensure that sound energy was focused on the user's upper body (especially the head and chest area), allowing the radar to effectively capture subtle physiological responses induced by sound waves (such as skull vibrations transmitted to the inner ear, triggering autonomic nervous system reactions).

[0074] If the radar detects that the user's height or position deviates from the preset area (e.g., a child user's height is less than 120cm), the array controller will adjust the phase parameters of the speaker in real time, and vertically shift the focus area downward (e.g., adjust it to 80-120cm) to ensure that the sound energy always acts on the user's effective perception area.

[0075] In some embodiments, the sound pressure level is 3 to 5 dB lower than the ambient background noise (when the signal sound pressure level is more than 3 dB lower than the ambient noise, humans cannot consciously distinguish it).

[0076] It's important to note that the purpose of "avoiding background noise" is to ensure that the physiological responses induced by microstimuli can be effectively detected. Specifically, even if the microstimuli themselves are imperceptible, the physiological signals they trigger (such as muscle tremors or slight changes in heart rate) may be very weak. If the frequency of the audio signal overlaps with environmental noise (such as the 1000-2000Hz background noise in a shopping mall), the noise will mask these weak physiological signals, making it impossible for sensors (such as radar) to distinguish between "stimuli-induced responses" and "environmental interference." For example, setting the center frequency of the audio signal to 3000Hz (avoiding the dominant frequency of environmental noise) allows radar to still detect hand tremors (the frequency of which is related to the stimulus audio) even at low sound pressure levels, ensuring the reliability of subsequent analysis.

[0077] Based on the embodiments provided in this application, the duration of visual element flashing is lower than the visual persistence threshold, the frequency of mechanical vibration is lower than the lower limit of human hearing perception, and the sound pressure level of audio signals is lower than the ambient background noise. These designs can apply stimulation unconsciously to the user, avoiding interference with the user's normal decision-making process. At the same time, the center frequency of the audio signal is offset from the frequency band of ambient noise, ensuring that micro-stimuli can be effectively captured, providing a feasible way to obtain the user's subconscious response and helping to gain a deeper understanding of the user's true preferences.

[0078] Furthermore, capturing the instantaneous micro-dynamic responses of non-semantic physiological and behavioral signals under the influence of information microstimuli includes:

[0079] Record the timestamp of each microstimulus applied;

[0080] Within a preset time window after application, the instantaneous offset of physiological indicators controlled by the autonomic nervous system in non-semantic physiological and behavioral signals relative to their respective baselines is measured.

[0081] The preset time window refers to a specific time interval used to capture instantaneous changes in non-semantic physiological and behavioral signals after the application of information micro-stimulus. Its setting is based on the response characteristics of the autonomic nervous system and the interaction rhythm of the vending machine scenario, as follows:

[0082] The duration is set at 1.5-2 seconds. This is based on the fact that the autonomic nervous system (especially the parasympathetic nervous system) has a delayed physiological response to microstimuli (approximately 300-500 ms), and the peak response usually occurs 1-1.5 seconds after stimulation, with the signal gradually returning to baseline after 2 seconds. This window can fully cover the critical "stimulus-response" period while avoiding the inclusion of too long subsequent signals (which may be mixed with physiological changes caused by the user's active actions, such as raising a hand to select an item).

[0083] If the radar detects that the user is in a state of rapid hand movement (such as continuously swiping the screen to browse products), the time window will be shortened to 1 second (to reduce motion noise interference); if the user is stationary (such as staring at a product), the time window can be extended to 2 seconds (to capture more subtle physiological fluctuations).

[0084] For example, when a vending machine applies mechanical microstimulation through piezoelectric ceramics, it continuously records the waveform of chest cavity fluctuations for the next 1.5 seconds to extract the high-frequency component changes of heart rate variability during that period, ensuring that what is captured is the instantaneous response triggered by the stimulus, rather than the physiological signals of the user's subsequent active decision.

[0085] It should be noted that physiological indicators are regulated by the autonomic nervous system (which is not under conscious control) and can reflect subconscious reactions.

[0086] Specifically, the high-frequency component of heart rate variability is controlled by the parasympathetic nervous system (a branch of the autonomic nervous system) and is associated with relaxation and emotional fluctuations (e.g., an increase in the high-frequency component of heart rate variability when seeing a preferred product reflects subconscious pleasure). Rhythmic interruptions in respiratory waveforms are regulated by the brainstem respiratory center (autonomic nervous system); unconscious apnea (0.5-1 second) may occur during decision-making hesitation and cannot be controlled consciously. Skin conductance is driven by the sympathetic nervous system; changes in sweat gland secretion lead to alterations in skin conductance, which are strongly correlated with subconscious emotions such as tension and interest (e.g., implicit concerns about high-priced products may trigger an increase in skin conductance).

[0087] The baseline is the pre-stimulation physiological benchmark value, which is dynamically established as follows:

[0088] Once a user approaches the vending machine (human detection by radar), a 5-second "baseline acquisition period" is initiated. During this period, no stimulation is applied, and physiological indicators are continuously recorded. The data within these 5 seconds are averaged (with a 0.5-second window) to remove transient spikes and noise, yielding the average values ​​of heart rate variability, respiratory rate, and skin conductance, which serve as the user's "baseline." If the user's interaction time exceeds 30 seconds, the baseline is recalculated every 10 seconds (to accommodate natural fluctuations in physiological state).

[0089] Among them, the physiological indicators include the high-frequency component of heart rate variability extracted from the chest cavity undulation waveform, the duration of rhythmic interruption of the respiratory waveform, and the amplitude of the simulated skin conductance response signal calculated by analyzing the sensitivity of radar echo phase to skin surface displacement.

[0090] Heart rate variability (HRV) refers to the minute fluctuations in the heart rate interval (RR interval). Its high-frequency component is a core indicator reflecting parasympathetic (vagus nerve) activity and is directly related to the user's subconscious state of relaxation and concentration. Specifically, high frequency is defined as 0.15-0.4 Hz (Hertz). This range is based on the spectral analysis standard of heart rate variability (referencing the TaskForce international guidelines). Fluctuations in this frequency band are mainly driven by respiratory rhythm (normal respiratory rate is approximately 0.2-0.3 Hz) and are entirely regulated by the parasympathetic nervous system, independent of conscious control. For example, when a user subconsciously develops an interest in a product, breathing will unconsciously slow down, and the power of the high-frequency component of HRV will significantly increase (manifested as increased energy in this frequency band).

[0091] After extracting the RR interval sequence from the chest cavity fluctuation waveform, a Fast Fourier Transform (FFT) is used for spectral analysis, and the signal energy (unit: ms² / Hz) in the 0.15-0.4Hz frequency band is used as the quantized value of the high-frequency component. To ensure accuracy, the RR interval sequence is first linearly interpolated (sampling rate set to 4Hz) to eliminate nonlinear interference in heart rate fluctuations.

[0092] For example, when directional acoustic microstimuli related to "low-priced goods" are applied, if the amplitude of the high-frequency component of HRV increases by more than 20% from the baseline within the preset time window, it indicates that the user's subconscious mind has a positive response to the "low-priced" attribute. This signal will be incorporated into the instantaneous micro-dynamic response and used to generate a dynamic attention weight vector in the future.

[0093] It needs to be explained that when skin conductivity increases, sweat gland secretion increases, and the skin surface will produce tiny expansions (about 0.1-1 micrometers). This tiny displacement will cause a change in the phase of the radar echo (the phase is sensitive to micrometer-level displacement).

[0094] By continuously acquiring radar echo phase data, a phase unwrapping algorithm is used to remove phase ambiguity, and then an inverse Fourier transform is applied to convert the phase change into a displacement signal in the time domain. The amplitude of this displacement signal is positively correlated with changes in skin conductance, and therefore can be used as a "skin conductance response simulation signal" to indirectly reflect the user's subconscious emotional fluctuations.

[0095] Based on the embodiments provided in this application, the instantaneous offset of physiological indicators relative to the baseline is measured within a preset time window, focusing on physiological indicators controlled by the autonomic nervous system, such as the high-frequency component of heart rate variability and rhythmic interruptions in respiratory waveforms. These indicators can accurately reflect the user's physiological response to microstimulation. The amplitude of the simulated skin conductance response signal is calculated through radar echo phase analysis, further enriching the methods for acquiring physiological indicators. This design accurately captures the instantaneous micro-dynamic response triggered by microstimulation, providing specific and quantifiable evidence for subsequent analysis of the user's subconscious feedback, and improving the targeting and effectiveness of response capture.

[0096] Furthermore, such as Figure 2 As shown, the time series analysis model is trained based on the following steps:

[0097] S201, Collect training dataset. The training dataset consists of time series data of user entropy index, time series data of instantaneous micro-dynamic response, and corresponding final purchase behavior records.

[0098] S202, Construct a multi-branch deep neural network model;

[0099] S203, construct a multi-task learning framework, using entropy index time series data and instantaneous micro-dynamic response time series data as input, the final purchased product category as the first supervision signal, and the user's preference weights for each product attribute parsed from the purchased products as the second supervision signal;

[0100] In the context of vending machines, "product category" is a structured classification system based on the core functions, form, and user consumption scenarios of products. This system clarifies the attribution of the user's final purchase behavior and specifically includes:

[0101] Primary categories: Classified by function, such as beverages, snacks, instant foods (bread, sandwiches), daily necessities (tissues, masks), etc.; Secondary categories: Subdivided under the primary categories, such as beverages can be divided into carbonated drinks, juices, functional drinks (containing caffeine / electrolytes), mineral water, etc.; snacks can be divided into puffed foods, nuts, chocolate, etc.

[0102] This classification not only conforms to the product display logic of vending machines (usually divided into categories), but also provides clear supervisory signals to the model through the category tags purchased by users (such as "energy drinks") (reflecting users' preferences for attributes such as "energizing" and "replenishing energy").

[0103] It should be noted that the core of analyzing preference weights from product purchases is establishing a mapping relationship between "product-attribute-weight," using the product's inherent attribute tags to infer the user's level of attention to each dimension, as detailed below:

[0104] For vending machine scenarios, four core attribute dimensions are defined: price sensitivity (low price / high price), functional requirements (such as refreshing, healthy, convenient), taste preference (sweet / salty / unflavored), and brand preference (well-known brands / niche brands).

[0105] If a user purchases a product with a unit price 20% lower than the average price of similar products (e.g., 3 yuan for bottled water, while the average price of similar products is 4 yuan), then the "price sensitivity" weight is assigned 0.8-1.0 (out of 1.0); if the user purchases a beverage labeled "sugar-free" or "low-calorie," then the "functional need - health" weight is assigned 0.7-0.9; if the user purchases a rice ball labeled "ready to eat," then the "functional need - convenience" weight is assigned 0.8-1.0; if the user purchases potato chips from a well-known brand (rather than a niche brand at the same price), then the "brand preference - well-known brand" weight is assigned 0.6-0.8.

[0106] Finally, the weights of each dimension are normalized (summing up to 1) and used as a second supervisory signal to enable the model to learn the association between "signal and preference" (e.g., "increased high-frequency component of heart rate variability" may correspond to "high weight of health attribute").

[0107] S204 employs the time backpropagation algorithm to iteratively optimize the weight parameters of the multi-branch deep neural network model with the goal of minimizing the weighted sum of the classification task loss corresponding to the first supervision signal and the regression task loss corresponding to the second supervision signal.

[0108] It should be noted that the Backpropagation Time (BPTT) algorithm is used in this scheme to process the temporal input (entropy index temporal data, instantaneous micro-dynamic response temporal data) and optimize the parameters of the multi-branch deep neural network. The specific implementation is as follows:

[0109] The input time series data (such as the entropy curve within 10 seconds, micro-dynamic response sequence) is split into continuous input segments according to time steps (such as every 0.5 seconds as one step), and input into the decision hesitation analysis branch (LSTM) and the subconscious feedback analysis branch (LSTM) to obtain the hidden state of each time step.

[0110] For classification tasks (first supervision signal: product category), calculate the cross-entropy loss between the output of the last time step and the true category label; for regression tasks (second supervision signal: preference weight), calculate the mean squared error loss between the weight vector output at each time step and the true weight vector.

[0111] The loss is propagated backward along the time step by the BPTT algorithm, and the gating parameters (forget gate, input gate, output gate) of the LSTM and the weights of the fully connected layer are updated, so that the model gradually learns the mapping pattern of "time series signal → category → preference".

[0112] Each iteration uses interaction data from 128 users (including time-series signals, purchase categories, and preference weights) as a batch, and repeats the iteration for 50-100 rounds until the loss converges.

[0113] In one specific implementation, let the total loss function be... The loss for the classification task is (Cross-entropy loss), the regression task loss is (Mean squared error loss), then the weighted summation formula is:

[0114]

[0115] in, The classification loss weights (ranging from 0.6 to 0.7) are used because product category is a direct result of user decision-making and is more crucial for the model to learn the association between "signal and behavior". The regression loss weights (ranging from 0.3 to 0.4) are used to ensure that the model simultaneously learns the preference weights.

[0116]

[0117] in, This represents the total number of product categories. The true category label (0 or 1); The predicted class probabilities by the model; An index for product categories;

[0118]

[0119] in, The number of attribute dimensions. Weights representing true preferences. The weights predicted by the model; For the index of the attribute dimension.

[0120] S205, the optimized multi-branch deep neural network model is determined as the time series analysis model.

[0121] Furthermore, multi-branch deep neural network models include:

[0122] The decision hesitation analysis branch, consisting of a one-dimensional convolutional neural network (CNN) and a long short-term memory network (LSTM), is used to process time-series data of entropy indicators.

[0123] The subconscious feedback analysis branch consists of another independent one-dimensional convolutional neural network and long short-term memory network, used to process transient micro-dynamic response time series data;

[0124] The dynamic feature fusion module integrates a signal-to-noise ratio (SNR) sensing gating unit, which includes a parallel evaluation sub-network used to calculate the confidence score of the input signal.

[0125] It should be noted that the signal-to-noise ratio sensing gating unit is used to dynamically adjust the fusion weights of the decision hesitation encoding vector and the subconscious feedback encoding vector. Its core function is to determine the reliability of the input signal through the evaluation sub-network, as detailed below:

[0126] The evaluation subnetwork consists of two parallel small neural networks (each processing one of the two encoded vectors). Each subnetwork consists of two fully connected layers (with the number of hidden neurons being half the dimension of the encoded vector) and a sigmoid activation function, with an output value range of [0,1] (i.e., the confidence score).

[0127] In some embodiments, the process of evaluating the sub-network includes:

[0128] Input: Decision hesitation encoding vector h1 (from a one-dimensional CNN+LSTM), subconscious feedback encoding vector h2 (from another one-dimensional CNN+LSTM);

[0129] Credibility assessment: Evaluation subnetwork 1 calculates the signal-to-noise ratio (signal energy / noise energy) for h1 and outputs a credibility score s1 (the higher the value, the less noise interference h1 is subjected to); Evaluation subnetwork 2 outputs s2 for h2 in the same way;

[0130] Weighted fusion: Normalized weights ensure a higher proportion of reliable signals. For example, if a user's hand movement signal is chaotic (h1 has high noise), but the heart rate response triggered by microstimuli is stable (h2 is reliable), then h2 will dominate in the fusion, improving feature reliability.

[0131] A cognitive state smoothing layer is used to apply constraints based on the principles of nonlinear dynamic systems; and a fully connected output layer;

[0132] The training of the cognitive state smoothing layer is constrained by a regularization term based on smoothing priors. This regularization term penalizes the higher-order differences of the output vector so that the cognitive state vector output by this layer conforms to the continuously changing physical laws.

[0133] It should be noted that nonlinear dynamical systems emphasize that the continuous changes in the system state are constrained by inherent laws and will not suddenly change. The cognitive state smoothing layer constrains the output of the user's cognitive state accordingly, as follows:

[0134] Users' preferences for products change continuously (e.g., the transition from "hesitant" to "definite" is a gradual process), rather than abruptly. The smoothing layer simulates the characteristics of a "damped system," making the current cognitive state vector... Compared to the previous moment satisfy ≈ +Δx (Δx is a small change) to avoid abrupt changes in the output vector (such as a sudden increase in the weight of an attribute from 0.2 to 0.8). A state memory unit is introduced within the layer to store the output vector at the previous moment. The current output must pass the similarity check with the memory vector (cosine similarity ≥ 0.8). Otherwise, a smooth adjustment is triggered (approaching the memory vector by a ratio of 0.2) to ensure that the cognitive state conforms to the continuity of human decision-making.

[0135] In some embodiments, this regularization term forces smooth output by penalizing higher-order changes in the cognitive state vector, as shown in the following formula:

[0136]

[0137] in, For regularization terms; This is the regularization coefficient (ranging from 0.01 to 0.05), used to control the intensity of the penalty; This refers to the number of time steps (e.g., 10 steps, corresponding to a 10-second interaction process). For the first The output vector of the cognitive state smoothing layer at each step; For the first The output vector of the cognitive state smoothing layer at each step; For the first The output vector of the cognitive state smoothing layer at each step; It represents the second-order difference (reflecting the curvature of a vector change), and a larger value indicates a more drastic change (such as a jump).

[0138] This represents the L2 norm.

[0139] It needs to be explained that in the formula It is the second difference of the vector, intuitively reflecting the change in the rate of change of the cognitive state vector. For example, if the user's weight on "health attributes" changes from... At time 0.3, to At time 0.4, and then... At time 0.5, the second difference is approximately 0 (the change is gradual). The contribution is small; if the weight suddenly increases from 0.3 to 0.8 ( Time 0.3, At time 0.8, the second-order difference increases significantly (the change is drastic). It will increase significantly, triggering a penalty.

[0140] When users interact with vending machines, sudden shifts in cognitive state are often caused by model noise or misjudgment, rather than actual decision-making. For example, a hand tremor accidentally detected by millimeter-wave radar might be misjudged as a "clear decision," causing the weight of "functional requirement" to jump from 0.2 to 0.9. This sudden shift does not conform to the user's actual thought process (in real-world scenarios, users need at least 1-2 seconds to determine whether the product's function matches their needs). By penalizing such mutations, the cognitive state vector output by the model is forced to be closer to the real decision-making rhythm (e.g., the weight change rate does not exceed 0.1 / second), ensuring that the dynamically generated attention weight vectors are reliable.

[0141] This design is more accurate than the first-order difference (which only reflects the change in adjacent moments) – the human decision-making process of “hesitation → clarity” is a gradual process, and the core is not “the magnitude of the change”, but “whether there is a sudden jump” (such as the shift from “not paying attention to prices” to “paying close attention to prices” usually requires a few seconds of weighing, rather than an instantaneous switch).

[0142] Its function is to balance "smoothing constraints" and "normal change requirements." If... If it's too large, it will excessively suppress reasonable changes (such as a rapid increase in the "price sensitivity" weight after a user sees promotional information); if If the value is too small, it cannot effectively punish sudden noise. In the vending machine scenario, the duration of a single user interaction is usually within 30 seconds. A value of 0.01 to 0.05 can ensure that the weight can increase from 0.2 to 0.6 within 10 seconds (a reasonable change), but it is prohibited to jump from 0.2 to 0.8 within 2 seconds (an unreasonable sudden change).

[0143] In conclusion, Essentially, by using mathematical constraints to "simulate the temporal continuity of human decision-making," the cognitive state output by the model is made more consistent with real-world interaction scenarios, providing a more reliable basis for subsequent resource scheduling.

[0144] In some embodiments, during training, the total loss function is: During gradient descent, if an update to a parameter causes an increase in the second-order difference, This will significantly increase, forcing the model to adjust its parameters, ultimately affecting the output. Presenting a continuous and smooth change (such as the attribute weight gradually increasing from 0.3 to 0.6, rather than a sudden increase), is more in line with the user's actual cognitive process.

[0145] Based on the embodiments provided in this application, a multi-branch deep neural network is used to process entropy indices and instantaneous micro-dynamic response time-series data respectively, achieving targeted analysis of different types of data. The multi-task learning framework uses the final purchased product category and product attribute preference weights as supervision signals, and optimizes them by combining the weighted sum of classification task loss and regression task loss, enabling the model to simultaneously learn user purchasing behavior and preference characteristics. The signal-to-noise ratio sensing gating unit of the dynamic feature fusion module can calculate the input signal credibility score, and the cognitive state smoothing layer is constrained by a smoothing prior regularization term to ensure that the output conforms to a continuous change pattern. These designs enable the model to generate dynamic attention weight vectors more accurately, improving the model's adaptability and accuracy.

[0146] Furthermore, the entropy index and the instantaneous micro-dynamic response are input into the trained time series analysis model to generate a dynamic attention weight vector describing the user's current preference for each product attribute, including:

[0147] The entropy index is input into the trained decision hesitation analysis branch to obtain the decision hesitation encoding vector;

[0148] The instantaneous micro-dynamic response is input into the trained subconscious feedback analysis branch to obtain the subconscious feedback encoding vector;

[0149] The decision hesitation encoding vector and the subconscious feedback encoding vector are input into a trained dynamic feature fusion module;

[0150] The signal-to-noise ratio sensing gating unit performs a weighted fusion of the decision hesitation encoding vector and the subconscious feedback encoding vector based on the confidence scores of the real-time calculated decision hesitation encoding vector and the subconscious feedback encoding vector to obtain a fused feature vector.

[0151] It should be explained that during the application phase, the evaluation sub-network performs a rapid credibility assessment of the real-time signal based on the trained parameters. The specific process is as follows:

[0152] During training, the evaluation subnetworks (two parallel small fully connected networks) have learned the mapping of "encoded vector features → signal-to-noise ratio → confidence score" through a large number of samples (e.g., "high proportion of high-frequency noise in the encoded vector → low signal-to-noise ratio → low confidence score"). During application, the pre-trained weight parameters are directly loaded, eliminating the need for retraining.

[0153] In some embodiments, the real-time computing step includes:

[0154] Input: Decision hesitation encoding vector h1 and subconscious feedback encoding vector h2 (both 64-dimensional);

[0155] Feature extraction: Evaluation subnetwork 1 performs two fully connected layer calculations on h1 (32 hidden layer neurons, consistent with training) to extract noise features (such as numerical fluctuation amplitude and outlier ratio) from the vector.

[0156] Signal-to-noise ratio estimation: The features are mapped to a confidence score s1 in the [0,1] interval by the Sigmoid activation function (the higher the value, the less noise interference h1 is, such as when the hand movement trajectory is stable, s1 is close to 0.8).

[0157] Similarly, the confidence score s2 is calculated for subnetwork 2 on h2 (s2 is close to 0.9 when the heart rate response induced by microstimuli is stable).

[0158] The evaluation subnetwork has fewer parameters and less computation time per frame, meeting the real-time requirements of vending machine interaction (the weight vector needs to be updated 5-10 times per second in a single user interaction).

[0159] It should be noted that the weighted fusion based on confidence scores is the core of the dynamic feature fusion module. By allowing reliable signals to dominate the fusion result, feature quality is improved. The specific steps are as follows: First, the two confidence scores are normalized. The confidence score of the decision hesitation encoding vector and the confidence score of the subconscious feedback encoding vector are added together to obtain a sum. Then, each confidence score is divided by this sum to obtain two normalized weight values. The sum of these two weight values ​​is 1, ensuring a reasonable distribution of signal proportions during fusion. Next, the decision hesitation encoding vector is multiplied by its corresponding normalized weight to obtain its contribution to the fusion. Similarly, the subconscious feedback encoding vector is multiplied by its corresponding normalized weight to obtain its contribution to the fusion. Finally, these two contributions are added together to form the fused feature vector.

[0160] The fused feature vector is input into a trained cognitive state smoothing layer to obtain a smooth feature vector that represents the continuous changes in the user's cognitive state.

[0161] The smoothed feature vector is input into the trained fully connected output layer and mapped to generate a dynamic attention weight vector.

[0162] It should be noted that the inference process corresponds to the training multi-branch deep neural network structure. Data flows through each module according to the path of "input → branch processing → feature fusion → smoothing constraint → output", as follows:

[0163] Input layer: Real-time acquired entropy indicators (such as the Shannon entropy time series within 10 seconds) are input to the decision hesitation analysis branch (consistent with the structure during training: one-dimensional CNN+LSTM); real-time captured instantaneous micro-dynamic responses (such as time series data such as high-frequency components of heart rate variability and duration of respiratory interruption within 10 seconds) are input to the subconscious feedback analysis branch (another independent one-dimensional CNN+LSTM).

[0164] Branch processing:

[0165] The decision hesitation analysis branch extracts local temporal features of the entropy sequence (such as the time point of entropy change) through a one-dimensional CNN, and then captures long-term dependence (such as the trend of entropy from high to low) through LSTM, and outputs a decision hesitation encoding vector (with a dimension of 64, which condenses the temporal features of user decision uncertainty).

[0166] Similarly, the subconscious feedback analysis branch processes the micro-dynamic response sequence through a one-dimensional CNN+LSTM to output a subconscious feedback encoding vector (with the same dimension of 64, which condenses the user's subconscious response characteristics to micro-stimuli).

[0167] Dynamic feature fusion module: Receives two encoded vectors, calculates confidence scores through the signal-to-noise ratio sensing gating unit (including evaluation sub-network), and weights and fuses them to output a fused feature vector (64 dimensions).

[0168] Cognitive State Smoothing Layer: Applies nonlinear dynamic system constraints to the fused feature vector (consistent with the smooth prior during training) and outputs a smooth feature vector (ensuring that the cognitive state changes continuously over time).

[0169] Fully connected output layer: Maps smooth feature vectors to dynamic attention weight vectors (the dimensions are consistent with the product attribute dimensions, such as 4 dimensions: price sensitivity, functional requirements, taste preferences, brand preference).

[0170] Based on the embodiments provided in this application, entropy indicators and instantaneous micro-dynamic responses are input into corresponding branches to obtain encoded vectors. These vectors are then weighted and fused according to confidence scores through the signal-to-noise ratio sensing gating unit of the dynamic feature fusion module. This cleverly utilizes the reliability differences of different signals, making the fused feature vectors more reflective of the user's true state. The smoothed feature vectors generated by the cognitive state smoothing layer conform to the continuous changes in user cognition. The fully connected output layer maps and generates dynamic attention weight vectors. The entire process is logically coherent, fully leveraging the advantages of the processing results of each branch and the fusion mechanism, thus improving the accuracy and rationality of the dynamic attention weight vectors.

[0171] Furthermore, based on the dynamic attention weight vector, a comprehensive semantic vector, fused from the semantic information of all currently schedulable resources, is calculated in real time, including:

[0172] Iterate through all currently active schedulable resources;

[0173] Whether a resource is active depends on whether it "interacts with the user within the current time window," specifically including:

[0174] Spatial association: The physical location of the resource display is within the user's current line of sight or operation range (such as the advertising screen area that the user is looking at, or the UI interface element that the user is operating).

[0175] Time-related: Resources are actively invoked or naturally presented by the system at the current stage of user interaction (such as browsing, hesitating, or making a selection) (e.g., when a user browses the beverage section, the advertisements in that section play automatically).

[0176] Functional relevance: The content of the resources is directly related to the user's current decision-making scenario (e.g., when the user focuses on price, resources such as "special offer tag" and "price comparison button" are activated).

[0177] Specific examples are as follows:

[0178] Advertising resources: When a user stands in front of a beverage shelf (the millimeter-wave radar detects the user's location), the functional beverage advertisement playing on the screen above the shelf is active; while the advertisements in the snack area are inactive because the user is not paying attention to them.

[0179] UI elements: When a user clicks the "Beverage Category" button, the sub-category labels such as "Carbonated Beverages" and "Juice", the "Sort" button, and the "Special Offer" label displayed on the interface are activated; the "Snack Category" related elements that are not displayed are deactivated.

[0180] Information prompts: When a user hesitates for more than 5 seconds, the "Hot Recommendations" text prompt that pops up automatically is activated; the "Insufficient Stock" prompt that is not triggered is deactivated.

[0181] In some embodiments, the activation state switches in real time based on user behavior. For example, when a user moves from the beverage section to the snack section, the advertising resources in the beverage section change from an active state to an inactive state, while the advertising resources in the snack section change from an inactive state to an active state. After the user closes the "Price Comparison" interface, all tags, charts, and other resources within that interface change from an active state to an inactive state. This dynamism ensures that the comprehensive semantic vector only incorporates resource information that currently has a real impact on the user's decision, avoiding interference from irrelevant resources and enabling the vector to accurately reflect the current information environment.

[0182] Extract the multidimensional semantic label vector and its confidence weight pre-assigned to each schedulable resource;

[0183] The multidimensional semantic tag vectors of all schedulable resources are superimposed according to their confidence weights and then normalized to generate a multidimensional comprehensive semantic vector representing the overall tone of the current information environment.

[0184] It should be explained that the process of fusing the semantic tag vectors of multiple schedulable resources into a comprehensive semantic vector consists of two steps: weighted summation and normalization, as detailed below:

[0185] Weighted summation: For each active schedulable resource, multiply its semantic label vector by its own confidence weight to obtain the weighted semantic vector of that resource; then sum the weighted semantic vectors of all resources to obtain the initial comprehensive vector. For example, if there are two active resources, a mineral water advertisement image (confidence weight 0.8, semantic vector [0.8, 0.6, 0.5, 0.7]) and a "New Product Recommendation" button (confidence weight 0.85, semantic vector [0.3, 0.7, 0.2, 0.4]), then the weighted semantic vector of the mineral water advertisement = 0.8 × [0.8, 0.6, 0.5, 0.7] = [0.64, 0.48, 0.65]. 4,0.56];Weighted semantic vector of the “New Product Recommendation” button = 0.85×[0.3,0.7,0.2,0.4]=[0.255,0.595,0.17,0.34];Initial comprehensive vector =[0.64+0.255,0.48+0.595,0.4+0.17,0.56+0.34]=[0.895,1.075,0.57,0.9].

[0186] Normalization: Divide the initial composite vector by the sum of the confidence weights of all resources to obtain the "composite semantic vector," ensuring that the values ​​of each dimension of the vector are within the range [0,1], thus avoiding excessive influence of high-confidence resources on the results. For example, the sum of the confidence weights of the two resources mentioned above = 0.8 + 0.85 = 1.65;

[0187] The comprehensive semantic vector = [0.895 / 1.65, 1.075 / 1.65, 0.57 / 1.65, 0.9 / 1.65] ≈ [0.542, 0.652, 0.345, 0.545]. This vector reflects the prominent features of the current resource as a whole in the "emotion" dimension (0.652) and the "functional" dimension (0.542).

[0188] The schedulable resources include advertising video or image content, graphic elements and their layout on the user interface, real-time changing discount or promotional information, and the display order and highlighting method of products in the digital interface.

[0189] Each type of schedulable resource is pre-assigned a set of multi-dimensional semantic tags and corresponding confidence weights through a resource description framework. The semantic tags are used to describe the attributes of the resource in multiple dimensions such as product function, emotional appeal, price level and brand tone.

[0190] It should be noted that the resource description framework is a rule system for semantically tagging the schedulable resources of vending machines (such as advertising content, interface elements, etc.). It achieves a quantitative description of resource attributes through multi-dimensional semantic tag vectors and confidence weights. A specific example is as follows:

[0191] Example 1: An advertisement image for a certain brand of mineral water:

[0192] Multidimensional semantic tag vector (containing 4 core dimensions: function, emotion, price, and brand): Function dimension: 0.8 (tags are "thirst-quenching, natural water source," directly related to the core function of mineral water); Emotion dimension: 0.6 (tags are "pure, comfortable," image background is snow-capped mountains and lakes, conveying a sense of nature); Price dimension: 0.5 (tags are "affordable," image corner shows "3 yuan / bottle," but not emphasized); Brand dimension: 0.7 (tags are "regionally well-known brand," image includes brand logo, but influence is limited). Confidence weight: 0.8 (This tag vector was manually labeled based on product attributes and validated through 100 user browsing data sessions, showing high consistency with user perception).

[0193] Example 2: The "New Arrivals" button on the user interface:

[0194] Multidimensional semantic tag vector (same as 4 dimensions): Functional dimension: 0.3 (no specific function, only indicates new product information); Emotional dimension: 0.7 (tags are "curiosity, novelty", the button uses a blinking animation to attract attention); Price dimension: 0.2 (unrelated to price); Brand dimension: 0.4 (tags are "new brand product", implying brand promotion intention). Confidence weight: 0.85 (button design purpose is clear, semantic tags are unambiguous).

[0195] Using this framework, vending machines can transform various resources into standardized semantic vectors, enabling them to be matched and analyzed in the same dimension as the user's dynamic attention weight vector.

[0196] Based on the embodiments provided in this application, schedulable resources in the active state are traversed, multi-dimensional semantic tag vectors and confidence weights are extracted, and a comprehensive semantic vector is generated by superimposing and normalizing the weights. This method comprehensively integrates the semantic information of various schedulable resources. The resource description framework endows resources with multi-dimensional semantic tags, covering product functions, emotional appeals, etc., so that the comprehensive semantic vector can fully reflect the tone of the current information environment. Schedulable resources include various types such as advertising content, interface elements, and promotional information, ensuring the comprehensiveness of the comprehensive semantic vector. This design achieves effective fusion of the semantic information of schedulable resources, providing a comprehensive and accurate reference for subsequent matching of dynamic attention weight vectors, and improving the targeting of scheduling.

[0197] Furthermore, based on the matching results between the dynamic attention weight vector and the comprehensive semantic vector, the advertising content, user interface elements, or promotional strategies are dynamically adjusted, including:

[0198] The system calculates the cosine distance between the dynamic attention weight vector and the comprehensive semantic vector in each preset semantic dimension in real time; based on the preset attention threshold and conflict threshold, it diagnoses the matching status of the dynamic attention weight vector and the comprehensive semantic vector in each preset semantic dimension; the matching status includes good matching, potential mismatch or clear semantic conflict.

[0199] Among them, the attention threshold and conflict threshold are used to classify the matching status (good match, potential mismatch, explicit semantic conflict) between the dynamic attention weight vector and the comprehensive semantic vector. Their settings are combined with experimental data and adaptive adjustment.

[0200] For example, note that the threshold is set to 0.7 (a "good match" is determined when the cosine distance is ≤0.7). This is based on the following: Through 500 user interaction experiments, it was found that when the cosine distance between two vectors is ≤0.7, the average user browsing time is reduced by 20%, and the willingness to click on recommended content increases by 35%, indicating a high degree of matching between resource information and user preferences. When the distance exceeds 0.7, users begin to show hesitation, hence this is defined as the critical point of "potential mismatch".

[0201] The conflict threshold was set at 0.85 (a cosine distance ≥ 0.85 is considered a "clear semantic conflict"). Basis: In the experiment, when the distance was ≥ 0.85, 65% of users showed obvious confusion (e.g., repeatedly swiping the interface, frowning), and 50% ultimately abandoned the purchase, indicating a significant contradiction between resource information and user preferences, requiring urgent adjustment.

[0202] Every 300 interactions, the purchase conversion rate for different distance ranges will be recalculated. If the actual purchase rate in the original "good match" range (≤0.7) drops by 15%, the attention threshold will be automatically lowered by 0.05 (e.g., from 0.7 to 0.65) to ensure that the threshold is synchronized with changes in user behavior.

[0203] In response to the diagnostic results of the matching status, a hierarchical scheduling strategy is executed, including:

[0204] If a potential mismatch exists, adjust the display intensity or confidence weight of the relevant schedulable resources on the mismatch dimension;

[0205] If a clear semantic conflict exists, a conflict resolution strategy is initiated. The conflict resolution strategy includes: selecting a new resource with a higher degree of matching in the conflict dimension from the alternative resource library to replace the existing resource, and / or generating and presenting explanatory information to bridge the semantic conflict.

[0206] It's important to explain that the overall cosine distance between the dynamic attention weight vector (reflecting the user's attention to each dimension) and the comprehensive semantic vector (reflecting the resource's characteristics in each dimension) can only determine "whether there is a match," but cannot indicate "which dimension is mismatched." However, calculating the cosine distance for each dimension (such as calculating the distance for the price dimension and the function dimension separately) can directly pinpoint specific gaps. For example, the cosine distance for the price dimension is 0.9 (far higher than the conflict threshold of 0.85), but the distance for the function dimension is 0.6 (lower than the attention threshold of 0.7), indicating that the core conflict lies in the "price" dimension (users focus on low prices, but resources emphasize high prices), while the "function" dimension matches well. This accurate positioning avoids the blind approach of "making comprehensive adjustments for overall mismatches," allowing subsequent resource optimization to focus on specific gaps and improve adjustment efficiency.

[0207] Different semantic dimensions have varying weights in influencing user decisions (e.g., price is more critical for price-sensitive users), and dimensional distance calculations can support differentiated scheduling logic. If the emotional dimension distance is too high (users prefer "simple" but the resource style is "flashy"), then the visual style of the resource should be adjusted first (e.g., simplifying the interface, reducing color saturation); if the functional dimension distance is too high (users focus on "low sugar" but the resource does not highlight this attribute), then "low sugar" labels or ingredient descriptions should be added to the resource first. This targeted optimization is more in line with users' actual needs than full-dimensional adjustments and can reduce unnecessary resource consumption (e.g., no need to modify already matched price dimension information).

[0208] User attention dynamically changes during interaction (e.g., shifting from "focusing on price" to "focusing on features"), and resource characteristics may also be updated in real time due to scheduling operations (e.g., changes in feature dimension characteristics after adding a "low sugar" tag). Real-time calculation of cosine distance across dimensions allows for real-time tracking of the matching status changes for each dimension. For example, if the initial price dimension distance is 0.8 (potential mismatch), and after the system pushes price comparison information, the distance drops to 0.6 (good match), indicating the strategy is effective. If the feature dimension distance increases from 0.6 to 0.75 (entering potential mismatch), new optimizations need to be triggered promptly (e.g., supplementing feature descriptions). This real-time tracking creates a closed loop of "monitoring-adjustment-re-monitoring" in scheduling, ensuring that the matching degree between resources and user attention remains at a high level.

[0209] It's important to explain that the directed graph of conflict propagation analyzes propagation relationships based on the conflict status (distance exceeding a threshold) of each dimension. The cosine distance across dimensions is the direct basis for determining whether a conflict exists in a particular dimension. If the price dimension distance exceeds the conflict threshold (0.85), and historical data shows that "price conflicts easily trigger functional conflicts" (propagation weight 0.7), the system can proactively optimize resource information in the functional dimension (e.g., actively stating "higher prices correspond to higher functions") to prevent conflict spread. This predictive adjustment based on dimensional data transforms conflict resolution from "passive response" to "proactive prevention," further reducing user decision-making fatigue.

[0210] The scheduling strategy is optimized based on a dynamically updated conflict propagation directed graph, which is used to model the conditional propagation relationship of conflicts between different semantic dimensions in order to predict and block conflict chains.

[0211] It should be noted that the directed graph of conflict propagation is used to model the propagation relationship of conflicts between different semantic dimensions and to guide the priority of conflict resolution, as follows:

[0212] Topology: Nodes: Represent 4 semantic dimensions (function, sentiment, price, brand); Directed edges: Pointing from node A to node B, indicating that "a conflict in dimension A may trigger a conflict in dimension B", and the weight of the edge (0-1) represents the probability of propagation (the higher the weight, the greater the probability of propagation).

[0213] Initial setup: Based on historical conflict data statistics, for example, after a price-dimensional conflict (users focus on low prices but resources highlight high prices), the occurrence rate of a function-dimensional conflict (users question whether "high prices really have high functions") is 60%, so the initial weight of the "price → function" edge is set to 0.6; after a brand-dimensional conflict (users are unfamiliar with the brand but resources strongly promote the brand), the occurrence rate of an emotional-dimensional conflict (users develop distrust) is 30%, so the initial weight of the "brand → emotion" edge is set to 0.3.

[0214] Update mechanism (based on online learning): Each time a new conflict event occurs (such as a conflict in dimension A followed by a conflict in dimension B), the weight of the edge is updated in the following way: record the historical observation count N and the number of propagation occurrences M of the edge (initially M / N is the weight); if a new observation is made that "conflict A triggers conflict B", the weight is updated to (M+1) / (N+1); if a new observation is made that "conflict A does not trigger conflict B", the weight is updated to M / (N+1).

[0215] For example, the initial weight of the "price → function" edge is 0.6 (based on 100 observations and 60 propagations). If a new observation of "price conflict triggering function conflict" is found, the weight is updated to (60+1) / (100+1)≈0.604, making the propagation relationship more realistic.

[0216] Based on the embodiments provided in this application, the cosine distance between the dynamic attention weight vector and the comprehensive semantic vector in each semantic dimension is calculated. Matching status is diagnosed based on preset thresholds, providing a clear basis for adjusting scheduling strategies. The hierarchical scheduling strategy takes different measures for potential mismatches and explicit semantic conflicts, adjusting display intensity, replacing resources, or generating explanatory information, demonstrating scheduling flexibility. The conflict propagation-based directed graph optimization strategy selection can predict and block conflict chains, preventing conflict escalation. This design makes scheduling adjustments more targeted and forward-looking, effectively improving dynamic matching and optimizing user experience.

[0217] Furthermore, generating and presenting explanatory information to bridge conflicting semantics includes:

[0218] Model multiple pre-defined interpretive communication strategies as selectable arms in a multi-armed slot machine;

[0219] The combined features, including the current dynamic attention weight vector and conflict dimension information, are used as the context feature vector;

[0220] The contextual multi-armed slot machine algorithm is adopted. The contextual feature vector is input into the strategy value evaluation model of the contextual multi-armed slot machine algorithm to calculate the expected reward value of each optional arm.

[0221] Select the option arm with the highest expected reward value as the current optimal explanatory communication strategy and execute it;

[0222] The positive or negative feedback generated by the user's interaction after execution will be transformed into reward signals;

[0223] Based on reward signals, the parameters of the strategy value assessment model are updated to optimize the selection of explanatory communication strategies under different contextual features.

[0224] In some embodiments, the following four explanatory communication strategies (i.e., "arms") are designed to address semantic conflicts in vending machine scenarios.

[0225] Arm 1: Functional Adaptation Explanation Strategy:

[0226] Applicable scenario: Users are interested in "features" (such as "low sugar") but the current resources do not highlight this attribute.

[0227] Specific operation: The UI will display the text "This product contains ≤5g / 100ml of sugar, which meets the low sugar standard" and highlight the relevant data in the ingredient list.

[0228] Arm 2: Price Justification Explanation Strategy

[0229] Applicable scenario: Users are interested in "low price" but the current resource is a high-priced product.

[0230] Specific operation: Display "This product uses imported raw materials, which improves the quality by 40% compared with similar ordinary products and has a better cost performance", and attach a raw material comparison chart.

[0231] Arm 3: Emotional Compensation Strategy

[0232] Applicable scenarios: When there is a conflict between user emotional preferences (such as "minimalist") and resource style (such as "flashy").

[0233] Specific operation: The interface switches to minimalist mode, and the text prompts "You have been adjusted to a minimalist style for easy and quick selection".

[0234] Arm 4: Brand Endorsement Strategy

[0235] Applicable scenario: Users have low brand awareness, but resources are being used to strongly promote the brand.

[0236] Specific instructions: The system will display "This brand has received quality certification for 5 consecutive years. Click to view the quality inspection report" and provide a clickable link to the certification mark.

[0237] In one specific implementation, the LinUCB algorithm is used as the core of the contextual multi-armed slot machine, and the model evaluates the policy value in the following way:

[0238] Context feature vector: includes the current conflict dimension (e.g., "price"), user dynamic attention weight vector (e.g., [price: 0.9, function: 0.1]), and comprehensive semantic vector (e.g., [price: 0.3, function: 0.8]), for a total of 8 features (4 dimension identifiers + 4 weight values), used to describe the current conflict scenario.

[0239] The linear model for each arm: A weight vector is assigned to each of the four policies. This vector is learned from historical data and used to predict the expected performance (expected reward) of the policy in the current scenario. The expected reward is calculated as the inner product of the context feature vector and the weight vector (i.e., the sum of corresponding element-wise multiplication).

[0240] Policy selection rule: Calculate the upper confidence bound (UCB) for each policy and select the policy with the highest UCB for execution. The formula for calculating the upper confidence bound is:

[0241]

[0242] in, It is the upper confidence bound of the k-th strategy; It is the expected reward of the k-th strategy; It is the exploration coefficient (set to 0.1), used to balance "choosing known effective strategies" and "trying new strategies"; It is a context feature vector; It is the historical feature covariance matrix of the k-th strategy; yes The inverse matrix; the square root term is used to measure the uncertainty of the prediction (the smaller the value, the lower the uncertainty).

[0243] For example, user interaction behavior can be converted into a numerical reward ranging from -2 to +2 to reflect the effect of the strategy. For example, +2 points: the user completes a purchase after the strategy is executed; +1 point: the user clicks on strategy-related content (such as ingredient list or quality inspection report) after the strategy is executed; +0.5 points: the user stays on the platform for more than 3 seconds longer after the strategy is executed; -1 point: the user closes the strategy prompt after the strategy is executed; -2 points: the user leaves the vending machine directly after the strategy is executed.

[0244] For example, if a user triggers the "price reasonableness explanation strategy" due to a price conflict, and then clicks on the raw material comparison chart and purchases the product, the final reward is +1 (click) +2 (purchase) = +3 (normalized to +2 proportionally).

[0245] In some embodiments, after each reward signal is obtained, the weight vector of each policy is updated using stochastic gradient descent, as follows:

[0246] Suppose the k-th policy is executed, the actual reward is r, and the expected reward predicted by the model is... ;

[0247] Calculate the loss (the difference between the predicted value and the actual reward): Loss = (r - )²;

[0248] Calculate the gradient of the loss with respect to the weight vector (reflecting the direction of weight adjustment): Gradient = -2 × (r - )×Context feature vector;

[0249] Update the weight vector: New weight = Old weight - Learning rate × Gradient (set the learning rate to 0.01 to control the adjustment range).

[0250] For example, if the predicted reward of the "price rationality explanation strategy" is 0.6 and the actual reward is +1, then the gradient is -2×(1-0.6)×context feature vector = -0.8×context feature vector. The weight vector will be adjusted in the direction of reducing loss, so that the prediction is more accurate when encountering similar scenarios next time.

[0251] Based on the embodiments provided in this application, the explanatory communication strategy is modeled as the selectable arms of a multi-armed slot machine. Combining contextual feature vectors, a contextual multi-armed slot machine algorithm is used to select the optimal strategy. This cleverly leverages the advantages of reinforcement learning, enabling the dynamic selection of the most effective explanation method based on different scenarios. The strategy value evaluation model parameters are updated through user interaction feedback, allowing the model to continuously learn and optimize. This improves the adaptability and effectiveness of the explanatory communication strategy in different conflict situations, better bridging conflict semantics and promoting positive interaction between users and vending machines.

[0252] According to another aspect of the embodiments of this application, an intelligent scheduling system for vending machines based on artificial intelligence is also provided. For example... Figure 3 As shown, the system includes:

[0253] The signal acquisition module 301 is used to acquire non-semantic physiological and behavioral signals generated when a user interacts with a vending machine through a non-optical sensing unit.

[0254] Entropy index calculation module 302 is used to calculate the entropy index that represents the uncertainty of the user's decision-making state based on non-semantic physiological and behavioral signals.

[0255] The micro-dynamic response capture module 303 is used to apply one or more information micro-stimuli below the perception threshold to the user and capture the instantaneous micro-dynamic response of non-semantic physiological and behavioral signals under the action of information micro-stimuli.

[0256] The attention weight vector generation module 304 is used to input the entropy index and the instantaneous micro-dynamic response into the trained time series analysis model to generate a dynamic attention weight vector that describes the user's current preference for each product attribute.

[0257] The comprehensive semantic vector calculation module 305 is used to calculate in real time the comprehensive semantic vector formed by fusing the semantic information of all currently schedulable resources based on the dynamic attention weight vector.

[0258] The adjustment module 306 is used to dynamically adjust the advertising content, user interface elements or promotional strategies based on the matching results of the dynamic attention weight vector and the comprehensive semantic vector, so as to improve the matching degree between the dynamic attention weight vector and the comprehensive semantic vector.

[0259] It should be noted that the embodiments implemented on the side of the AI-based intelligent scheduling system for vending machines in this application can be referenced to each other, and will not be described in detail here.

[0260] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described intelligent scheduling method for vending machines based on artificial intelligence is also provided. This electronic device may be... Figure 4 The terminal device or server shown. This embodiment uses this electronic device as an example of a server. Figure 4 As shown, the electronic device includes a memory 402, a processor 404, and a transmission device 406. The memory 402 stores a computer program, and the processor 404 is configured to execute the steps of any of the above method embodiments through the computer program.

[0261] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.

[0262] Optionally, the transmission device 406 is used to receive or send data via a network. Specific examples of the network described above may include wired and wireless networks. In one example, the transmission device 406 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 406 is a Radio Frequency (RF) module used to communicate with the Internet wirelessly. Furthermore, the electronic device also includes a display 408 and a connection bus 410, which connects the various module components within the electronic device.

[0263] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. An intelligent scheduling method for vending machines based on artificial intelligence, characterized in that, include: Non-semantic physiological and behavioral signals generated when users interact with vending machines are collected using non-optical sensing units; Based on the aforementioned non-semantic physiological and behavioral signals, an entropy index representing the uncertainty of the user's decision-making state is calculated. Apply one or more information microstimuli below the perception threshold to the user and capture the instantaneous micro-dynamic response of the non-semantic physiological and behavioral signals under the action of the information microstimuli; The entropy index and the instantaneous micro-dynamic response are input together into a trained time series analysis model to generate a dynamic attention weight vector that describes the user's current preference for each product attribute. Based on the dynamic attention weight vector, a comprehensive semantic vector composed of the semantic information of all currently schedulable resources is calculated in real time. Based on the matching result between the dynamic attention weight vector and the comprehensive semantic vector, the advertising content, user interface elements, or promotional strategies are dynamically adjusted to improve the matching degree between the dynamic attention weight vector and the comprehensive semantic vector. The non-optical sensing unit is a millimeter-wave radar; the non-semantic physiological and behavioral signals include the chest cavity undulation waveform, continuous point cloud data of hand movement trajectory, and frequency domain features of microtremors on the body surface, which are obtained by the millimeter-wave radar and preprocessed. The calculation of the entropy index characterizing the uncertainty of the user's decision-making state based on the non-semantic physiological and behavioral signals includes: The continuous point cloud data of the hand movement trajectory is reconstructed and smoothed to extract the motion change sequence of its speed and direction; Perform a windowed Fourier transform on the motion change sequence to obtain the time-frequency spectrum. Calculate the Shannon entropy value of the time spectrum graph within each time window, and use the curve of the change of Shannon entropy value over time as the entropy value index. The increase of Shannon entropy value indicates the enhancement of decision hesitation, and the decrease of Shannon entropy value indicates that the decision tends to be clear. The information microstimuli below the perception threshold include any one or more of the following stimuli: Encoded visual elements of a preset color or shape are flashed in a specific area of ​​the display unit for a preset duration; wherein the preset duration is shorter than the time threshold required for visual persistence detection. Mechanical vibrations with frequencies below the lower limit of human hearing are generated by a piezoelectric ceramic actuator coupled to the glass panel of the vending machine. A directional acoustic generator focuses the acoustic energy of an audio signal onto a preset spatial area in front of the vending machine; wherein the center frequency of the audio signal is pre-selected in a frequency band that is offset from the main energy frequency band of the ambient background noise, and the sound pressure level of the audio signal is set to be lower than the sound pressure level of the ambient background noise at the center frequency band.

2. The intelligent scheduling method for vending machines based on artificial intelligence according to claim 1, characterized in that, The capture of the instantaneous micro-dynamic response of the non-semantic physiological and behavioral signals under the influence of the information micro-stimulation includes: Record the timestamp of each microstimulus applied; Within a preset time window after application, the instantaneous offset of the physiological indicators controlled by the autonomic nervous system in the non-semantic physiological and behavioral signals relative to their respective baselines is measured. The physiological indicators include the high-frequency component of heart rate variability extracted from the chest cavity undulation waveform, the duration of rhythmic interruption of the respiratory waveform, and the amplitude of the simulated skin conductance response signal calculated by analyzing the sensitivity of radar echo phase to skin surface displacement.

3. The intelligent scheduling method for vending machines based on artificial intelligence according to claim 2, characterized in that, The time series analysis model is trained based on the following steps: Collect a training dataset, which consists of time-series data of user entropy index, time-series data of instantaneous micro-dynamic response, and corresponding final purchase behavior records; Construct a multi-branch deep neural network model; A multi-task learning framework is constructed, with the entropy index time series data and the instantaneous micro-dynamic response time series data as inputs, the final purchased product category as the first supervision signal, and the user's preference weights for each product attribute parsed from the purchased products as the second supervision signal. The weight parameters of the multi-branch deep neural network model are iteratively optimized by employing the time backpropagation algorithm with the goal of minimizing the weighted sum of the classification task loss corresponding to the first supervision signal and the regression task loss corresponding to the second supervision signal. The optimized multi-branch deep neural network model is determined as the time series analysis model; The multi-branch deep neural network model includes: The decision hesitation analysis branch, composed of a one-dimensional convolutional neural network and a long short-term memory network, is used to process the time-series data of the entropy index. The subconscious feedback analysis branch, composed of another independent one-dimensional convolutional neural network and long short-term memory network, is used to process the instantaneous micro-dynamic response time-series data; The dynamic feature fusion module integrates a signal-to-noise ratio (SNR) sensing gating unit, which includes a parallel evaluation sub-network for calculating the confidence score of the input signal. A cognitive state smoothing layer is used to apply constraints based on the principles of nonlinear dynamic systems; and a fully connected output layer.

4. The intelligent scheduling method for automated vending machines based on artificial intelligence according to claim 3, characterized in that, The step of inputting the entropy index and the instantaneous micro-dynamic response into a trained time series analysis model to generate a dynamic attention weight vector describing the user's current preference for each product attribute includes: The entropy index is input into the trained decision hesitation analysis branch to obtain the decision hesitation encoding vector; The instantaneous micro-dynamic response is input into the trained subconscious feedback analysis branch to obtain the subconscious feedback encoding vector; The decision hesitation encoding vector and the subconscious feedback encoding vector are input into the trained dynamic feature fusion module; The signal-to-noise ratio sensing gating unit performs a weighted fusion of the decision hesitation encoding vector and the subconscious feedback encoding vector based on the confidence score of the decision hesitation encoding vector and the confidence score of the subconscious feedback encoding vector calculated in real time, to obtain a fused feature vector. The fused feature vector is input into the trained cognitive state smoothing layer to obtain a smooth feature vector representing the continuous changes in the user's cognitive state. The smoothed feature vector is input into the trained fully connected output layer to generate the dynamic attention weight vector.

5. The intelligent scheduling method for automated vending machines based on artificial intelligence according to claim 4, characterized in that, The process of calculating a comprehensive semantic vector in real time, based on the dynamic attention weight vector and fused from the semantic information of all currently schedulable resources, includes: Iterate through all currently active schedulable resources; Extract the multidimensional semantic label vector and its confidence weight pre-assigned to each schedulable resource; The multidimensional semantic tag vectors of all schedulable resources are superimposed according to their confidence weights and then normalized to generate a multidimensional comprehensive semantic vector representing the overall tone of the current information environment. The schedulable resources include advertising video or image content, graphic elements and their layout on the user interface, real-time changing discount or promotional information, and the display order and highlighting method of goods in the digital interface. Each type of schedulable resource is pre-assigned a set of multi-dimensional semantic tags and corresponding confidence weights through a resource description framework. The semantic tags are used to describe the attributes of the resource in multiple dimensions such as product function, emotional appeal, price level, and brand tone.

6. The intelligent scheduling method for vending machines based on artificial intelligence according to claim 5, characterized in that, The step of dynamically adjusting advertising content, user interface elements, or promotional strategies based on the matching result of the dynamic attention weight vector and the comprehensive semantic vector includes: Real-time calculation of the cosine distance between the dynamic attention weight vector and the comprehensive semantic vector in each preset semantic dimension; Based on preset attention thresholds and conflict thresholds, the matching status of the dynamic attention weight vector and the comprehensive semantic vector in each preset semantic dimension is diagnosed; wherein, the matching status includes good matching, potential mismatch or clear semantic conflict. In response to the diagnostic results of the matching status, a hierarchical scheduling strategy is executed, including: If a potential mismatch exists, adjust the display intensity or confidence weight of the relevant schedulable resources on the mismatch dimension; If a clear semantic conflict exists, a conflict resolution strategy is initiated; wherein, the conflict resolution strategy includes: selecting a new resource with a higher matching degree in the conflict dimension from the candidate resource library to replace the existing resource, and / or generating and presenting explanatory information to bridge the semantics of the conflict; The scheduling strategy is optimized based on a dynamically updated conflict propagation directed graph, which is used to model the conditional propagation relationship of conflicts between different semantic dimensions in order to predict and block conflict chains.

7. The intelligent scheduling method for vending machines based on artificial intelligence according to claim 6, characterized in that, The generation and presentation of interpretive information to bridge conflicting semantics includes: Model multiple pre-defined interpretive communication strategies as selectable arms in a multi-armed slot machine; The combined features, including the current dynamic attention weight vector and conflict dimension information, are used as the context feature vector; The context feature vector is input into the strategy value evaluation model of the context multi-armed slot machine algorithm to calculate the expected reward value of each optional arm. Select the option arm with the highest expected reward value as the current optimal explanatory communication strategy and execute it; The positive or negative feedback generated by the user's interaction after execution will be transformed into reward signals; Based on the reward signal, the parameters of the strategy value assessment model are updated to optimize the selection of explanatory communication strategies under different contextual features.

8. An intelligent scheduling system for automated vending machines based on artificial intelligence, characterized in that: include: The signal acquisition module is used to acquire non-semantic physiological and behavioral signals generated when a user interacts with a vending machine through a non-optical sensing unit; The entropy index calculation module is used to calculate the entropy index, which represents the uncertainty of the user's decision-making state, based on the non-semantic physiological and behavioral signals. The micro-dynamic response capture module is used to apply one or more information micro-stimuli below the perception threshold to the user and capture the instantaneous micro-dynamic response of the non-semantic physiological and behavioral signals under the action of the information micro-stimuli. The attention weight vector generation module is used to input the entropy index and the instantaneous micro-dynamic response into a trained time series analysis model to generate a dynamic attention weight vector that describes the user's current preference for each product attribute. The comprehensive semantic vector calculation module is used to calculate in real time the comprehensive semantic vector formed by fusing the semantic information of all currently schedulable resources based on the dynamic attention weight vector. The adjustment module is used to dynamically adjust the advertising content, user interface elements or promotional strategies based on the matching result of the dynamic attention weight vector and the comprehensive semantic vector, so as to improve the matching degree between the dynamic attention weight vector and the comprehensive semantic vector. The non-optical sensing unit is a millimeter-wave radar; the non-semantic physiological and behavioral signals include the chest cavity undulation waveform, continuous point cloud data of hand movement trajectory, and frequency domain features of microtremors on the body surface, which are obtained by the millimeter-wave radar and preprocessed. The calculation of the entropy index characterizing the uncertainty of the user's decision-making state based on the non-semantic physiological and behavioral signals includes: The continuous point cloud data of the hand movement trajectory is reconstructed and smoothed to extract the motion change sequence of its speed and direction; Perform a windowed Fourier transform on the motion change sequence to obtain the time-frequency spectrum. Calculate the Shannon entropy value of the time spectrum graph within each time window, and use the curve of the change of Shannon entropy value over time as the entropy value index. The increase of Shannon entropy value indicates the enhancement of decision hesitation, and the decrease of Shannon entropy value indicates that the decision tends to be clear. The information microstimuli below the perception threshold include any one or more of the following stimuli: Encoded visual elements of a preset color or shape are flashed in a specific area of ​​the display unit for a preset duration; wherein the preset duration is shorter than the time threshold required for visual persistence detection. Mechanical vibrations with frequencies below the lower limit of human hearing are generated by a piezoelectric ceramic actuator coupled to the glass panel of the vending machine. A directional acoustic generator focuses the acoustic energy of an audio signal onto a preset spatial area in front of the vending machine; wherein the center frequency of the audio signal is pre-selected in a frequency band that is offset from the main energy frequency band of the ambient background noise, and the sound pressure level of the audio signal is set to be lower than the sound pressure level of the ambient background noise at the center frequency band.