Fragrance display interaction control system based on camera visual input

By using a camera-based visual input-based fragrance display and interactive control system, user actions are identified in stages and confirmed using multiple features. This solves the problem of misjudgment of user behavior in existing technologies and achieves a more stable fragrance display control and user interaction experience.

CN122347770APending Publication Date: 2026-07-07黄奕文
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
黄奕文
Filing Date
2026-04-13
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

In existing fragrance display interaction solutions, the continuous behavioral process of users from approaching, probing, confirming candidates to leaving is not effectively understood in stages, leading to misjudgment and unstable control boundaries.

Method used

The fragrance display and interactive control system based on camera visual input includes a trajectory construction module, an interaction recognition module, a fragrance release confirmation module, and a threshold adjustment module. By generating an outer approach zone, a trial zone, and a core confirmation zone, it identifies user actions in stages and combines convergence features, stability features, alignment features, and stage confidence features to make confirmation judgments and output graded visual pre-response and atomization execution modes.

Benefits of technology

It improves the accuracy and stability of fragrance display control, reduces the probability of false triggering, enhances the continuity and perceptibility of user interaction, and achieves more accurate fragrance release control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122347770A_ABST
    Figure CN122347770A_ABST
Patent Text Reader

Abstract

This invention discloses a fragrance display and interactive control system based on camera visual input, relating to the field of computer vision technology. The system calibrates the display screen and camera to determine the object center point and object interaction area of ​​a virtual tree object. It extracts the hand area, original interaction points, and smoothed interaction points to generate a basic trajectory description sequence, including object distance, displacement velocity, and local jitter amplitude. Based on this basic trajectory description sequence, a stage recognition model is established, outputting stages such as approach, probing, candidate confirmation, and departure, and correspondingly outputting graded visual pre-responses. Subsequently, convergence, stabilization, alignment, dwell, and stage confidence features are extracted. A confirmation score is calculated using a confirmation model to determine the fogging execution mode. Finally, based on short-term subsequent behavior after execution, a lower confirmation threshold is narrowed and corrected, thus forming a closed-loop control process from trajectory construction, stage recognition, confirmation judgment to threshold correction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and more specifically, to a fragrance display and interactive control system based on camera visual input. Background Technology

[0002] In existing fragrance display and interaction solutions, a visual interaction link is usually formed by a display screen and a camera. The user's hand movements are used as trigger inputs. Under single-camera conditions, motion change detection, hand area extraction, and position dwell judgment are used to identify whether the user has triggered a functional response associated with the display object. Interaction recognition often revolves around whether the hand enters the vicinity of the object, whether it dwells, or whether a preset threshold is met, and then decides whether to execute subsequent fragrance release or corresponding display control.

[0003] The existing technology has the following shortcomings: Existing technologies fail to provide a phased understanding of the continuous behavioral process of users from approaching, probing, candidate confirmation to leaving, and also lack a closed-loop adjustment mechanism that combines phase information, confirmation characteristics, and post-execution behavioral feedback. When users merely pass by the periphery of an object, briefly linger at its edge, make tentative movements, or experience slight fluctuations near the candidate confirmation boundary, they are prone to misinterpreting observation or probing lingering as genuine confirmation actions, thereby affecting the judgment accuracy and execution boundary stability of fragrance display control.

[0004] To address the above problems, this invention proposes a solution. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a fragrance display interactive control system based on camera visual input to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: A fragrance display and interactive control system based on camera visual input includes: a trajectory construction module, an interaction recognition module, a fragrance release confirmation module, and a threshold adjustment module, with signal connections between the modules; The trajectory construction module obtains the calibration relationship between the display screen and the camera, determines the object center point and the main interaction area of ​​the virtual tree object, generates the outer approach area, the probing area and the core confirmation area, extracts the main hand area and the original interaction points, smooths the original interaction points, calculates the object distance, displacement speed and local jitter amplitude, and forms the basic trajectory description sequence. Interactive recognition module: Based on the basic trajectory description sequence, a phase recognition model is established, which constructs phase features including object distance, displacement velocity, velocity change, local jitter amplitude, alignment deviation and core confirmation area dwell count, and outputs standby state, approach state, probing state, candidate confirmation state and departure state, and outputs hierarchical visual pre-response by combining phase continuity judgment and phase back-off judgment. Aroma release confirmation module: When the candidate confirmation state meets the continuous condition, it extracts convergence features, stability features, alignment features, dwell features and stage confidence features, establishes a confirmation model and calculates the confirmation score, and combines the confirmation score with the number of continuous frames in the core confirmation area to determine the atomization execution mode; Threshold adjustment module: Records confirmation score, number of frames in the core confirmation area, fogging execution mode and short-term subsequent behavior after execution, filters samples in the range of the lower confirmation threshold, and performs amplitude limiting correction on the lower confirmation threshold based on the comparison results of suspected premature triggering samples and suspected overly strict samples.

[0007] In a preferred embodiment, the trajectory construction module includes the following steps: The specific method for generating the outer approach area, probing area, and core confirmation area around the main interaction area of ​​the object is as follows: The outer approach area is obtained by proportionally expanding the outer boundary of the main interaction area of ​​the object, and the probing area and the core confirmation area are obtained by proportionally shrinking the inner boundary of the main interaction area of ​​the object. The current frame is compared with the background reference to obtain motion candidate regions. Skin color screening is performed within the motion candidate regions. The continuity is judged by combining the hand position of the previous frame. The connected region with the largest area and the highest continuity with the position of the previous frame is retained as the main hand region. The original interaction points are smoothed using an exponential smoothing method, and the object distance, displacement velocity and local jitter amplitude are calculated based on the smoothed interaction points. The displacement velocity is obtained by dividing the distance between two adjacent smooth interaction points by the time interval, and the local jitter amplitude is obtained by statistically analyzing the average deviation of the smooth interaction points in the most recent few frames from the local mean center.

[0008] In a preferred embodiment, the interaction recognition module includes the following steps: The graded visual pre-response includes Level 1 visual pre-response, Level 2 visual pre-response, and Level 3 visual pre-response; When in a near-term state, a first-level visual pre-response is output, causing the canopy of the virtual tree object to sway slightly. When in a probing state, a secondary visual pre-response is output, which enhances the brightness of the canopy edge or produces particle light spots; When in the candidate confirmation state, a level three visual pre-response is output, causing a breathing-like brightness change in the center of the canopy. The stage persistence determination is that a certain stage is only identified as the current stage if it remains the highest probability stage for a consecutive number of frames. The stage rollback is determined when the candidate confirmation state rolls back to the trial zone from the interaction point in the subsequent window and the displacement velocity increases and the local jitter amplitude increases.

[0009] In a preferred embodiment, the aroma release confirmation module includes the following steps: The convergence feature is the velocity change trend before and after entering the candidate confirmation state; the stability feature is the average deviation of smooth interaction points within the candidate confirmation window from the local mean center; the alignment feature is the average deviation of interaction points within the candidate confirmation window from the object center point; the dwell feature is the number of frames in the core confirmation area; and the stage confidence feature is the candidate confirmation stage probability output by the interaction recognition module. The atomization execution modes include no execution mode, light fog mode, standard fog mode, and enhanced fog mode; When the confirmation score is lower than the lower confirmation threshold, it is determined to be in non-execution mode; When the confirmation score is not lower than the lower confirmation threshold and lower than the higher confirmation threshold, and the number of consecutive frames in the core confirmation area is lower than the number of consecutive frames, it is determined to be light fog mode. When the confirmation score is not lower than the lower confirmation threshold and is lower than the higher confirmation threshold, and the number of consecutive frames in the core confirmation area is not lower than the number of consecutive frames, it is judged as standard fog mode. When the confirmation score is not lower than the higher confirmation threshold, it is determined to be enhanced fog mode.

[0010] In a preferred embodiment, the threshold adjustment module includes the following steps: The nearest neighbor interval refers to the absolute difference between the confirmation score and the current lower confirmation threshold not exceeding a preset small range; The suspected premature triggering samples are those where, after being identified as being in light fog mode, users returned to the probing area or core confirmation area within a short period of time and continued to probe. The suspected overly strict sample is the previous interaction sample where the user did not enter the light fog mode but formed a more stable candidate confirmation state again in a short period of time and finally entered the light fog mode or standard fog mode. A small-step amplitude limiting correction strategy is adopted for the lower confirmation threshold: when the number of suspected premature triggering samples in the neighboring interval is higher than the number of suspected overly strict samples, the lower confirmation threshold is adjusted upward by a small step. When the number of suspected overly strict samples exceeds the number of suspected prematurely triggered samples, the lower confirmation threshold will be adjusted downward by a small step. The revised lower confirmation threshold is limited to the preset upper and lower bounds.

[0011] The technical effects and advantages of the fragrance display and interactive control system based on camera visual input of the present invention are as follows: Compared to methods that directly use a single pause or instantaneous position as the trigger, this invention first constructs a standardized trajectory input for a virtual tree object, then performs phased recognition of user actions, and combines convergence features, stability features, alignment features, pause features, and phase confidence features to make a confirmation judgment based on the candidate confirmation stage, so that the atomization execution no longer depends on a single pause condition. At the same time, this invention also continuously accumulates temporal evidence for subsequent confirmation judgments through hierarchical visual pre-response, and finely corrects the lower confirmation threshold based on short-term subsequent behavior after execution, thereby giving the entire fragrance display control process better phase coherence, judgment targeting, and boundary adjustment capabilities. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of the structure of a fragrance display and interactive control system based on camera visual input according to the present invention.

[0013] Figure 2 This is a diagram showing the state transition of visual pre-response for stage recognition and hierarchical classification. Detailed Implementation

[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0015] Example: Please refer to Figures 1-2 As shown, the present invention discloses a fragrance display interactive control system based on camera visual input, including: a trajectory construction module, an interaction recognition module, a fragrance release confirmation module, and a threshold adjustment module, with each module connected by a signal. The trajectory construction module obtains the calibration relationship between the display screen and the camera, determines the object center point and the main interaction area of ​​the virtual tree object, generates the outer approach area, the probing area and the core confirmation area, extracts the main hand area and the original interaction points, smooths the original interaction points, calculates the object distance, displacement speed and local jitter amplitude, and forms the basic trajectory description sequence. Interactive recognition module: Based on the basic trajectory description sequence, a phase recognition model is established, which constructs phase features including object distance, displacement velocity, velocity change, local jitter amplitude, alignment deviation and core confirmation area dwell count, and outputs standby state, approach state, probing state, candidate confirmation state and departure state, and outputs hierarchical visual pre-response by combining phase continuity judgment and phase back-off judgment. Aroma release confirmation module: When the candidate confirmation state meets the continuous condition, it extracts convergence features, stability features, alignment features, dwell features and stage confidence features, establishes a confirmation model and calculates the confirmation score, and combines the confirmation score with the number of continuous frames in the core confirmation area to determine the atomization execution mode; Threshold adjustment module: Records confirmation score, number of frames in the core confirmation area, fogging execution mode and short-term subsequent behavior after execution, filters samples in the range of the lower confirmation threshold, and performs amplitude limiting correction on the lower confirmation threshold based on the comparison results of suspected premature triggering samples and suspected overly strict samples.

[0016] In the trajectory construction module, the calibration relationship between the display screen and the camera is obtained, the object center point and main interaction area of ​​the virtual tree object are determined, the peripheral approach area, the probing area, and the core confirmation area are generated, the main hand area and the original interaction points are extracted, the original interaction points are smoothed, the object distance, displacement velocity, and local jitter amplitude are calculated, and a basic trajectory description sequence is formed. Specific content includes: Establish a basic interaction link, using a fixed rectangular panel as the interaction panel, and setting the virtual tree object displayed on the screen as the interaction object, transforming the hand gesture changes around the virtual tree object into the core input for subsequent judgment; In the fragrance display scenario, the system determines whether the user actually wants to trigger the release of the fragrance associated with the virtual tree object based on the visual input, transforms the interactive object from a physical rectangular board into a virtual tree object, and transforms the output result from a normal function response into a standardized trajectory input for fragrance execution. The loading of the natural scene display interface includes a forest background, virtual tree objects, visual pre-response animation resources associated with the virtual tree objects, and hidden coordinate parameters for object semantic calibration. During the startup phase, it first enters a static initialization state. In the static initialization state, the camera captures an empty scene sequence to build a background reference when no user enters. Synchronously output the current virtual tree object's bounding box, approximate area of ​​the tree crown outline, center point of the tree trunk, and center point of the tree crown in the screen coordinate system, and perform a single-image calibration on the display screen plane; Preferably, by detecting the positions of the four corners of the display screen in the camera image, the homography mapping relationship from screen coordinates to camera image coordinates is calculated, and the semantic points of the virtual tree object on the screen are projected into the camera image according to the following formula: ;in, These are the coordinates of the virtual tree object's point in the display coordinate system. These are the projected coordinates of the point in the camera image coordinate system. This is the homography matrix from the screen coordinate system to the camera image coordinate system. This is the scaling factor for homogeneous coordinates; The above mapping can map the center point of the tree canopy to the center point of the object. And form the main interactive area of ​​the object based on the canopy projection outline. ; Main interaction area around the object Automatically generate peripheral proximity zone Exploration Zone and core confirmation area ; Preferably, the peripheral approach zone is obtained by expanding the outer boundary of the main object interaction area by 8% to 15%, with an initial value of 10%; the probing zone is obtained by shrinking the inner boundary of the main object interaction area by 10% to 18%, with an initial value of 15%; and the core confirmation zone is obtained by shrinking the inner boundary of the main object interaction area by 30% to 40%, with an initial value of 35%. This interprets user actions as a progressive process involving three levels: spatial approach, edge probing, and center confirmation relative to the virtual tree object, providing a clear spatial semantic basis for the phased visual interaction recognition of the subsequent interaction recognition module. After the object region is established, the process of hand region extraction and candidate interaction point generation begins. The current frame is compared with the background reference to obtain motion change candidate regions. A skin color filter is used for secondary segmentation. Under single-camera conditions, combining motion constraints and color constraints is an feasible basic path. The formula for obtaining the motion candidate region is: ;in, For the first Frame image in pixels The grayscale or brightness value at that location. The value at the corresponding pixel position serves as a background reference. For motion candidate thresholds, preferably, The value can be between 12 and 25, and the initial value can be 18. After obtaining the motion candidate region, color filtering is performed only within the candidate region. It is preferred to use color space to constrain skin color candidates, while brightness and saturation limits are applied to exclude large areas of gray-white reflection or pure screen bright spots. Median filtering and morphological closing operations are performed on the filtering results to repair small cracks in the hand contour and remove scattered noise. Based on the hand position in the previous frame, the connected region with the largest area and the highest continuity with the previous frame position is selected as the main hand region. The highest continuity is determined by whether the distance between the center of the current connected region and the center of the main hand region in the previous frame is the smallest. This can avoid frequent switching between multiple candidate regions when the screen background brightness changes or there is reflection from the physical scene, and ensure the continuity of the subsequent trajectory. After obtaining the primary hand area, extract the original interaction points from this primary hand area to describe the user's current operation position; Preferably, either of the following two methods can be used: The first method is to take the point on the outline of the main hand area that points towards the center of the object. The foremost point on one side is used as the original interaction point; The second method is to take the centroid of the main hand area and the center point of the object. The intersection of the line connecting the two points and the boundary of the hand contour is taken as the original interaction point. Considering that the original interaction point is easily affected by finger opening and closing, posture changes, and local occlusion; Perform a smoothing calculation on it to obtain the smoothed interaction point, which is obtained by the following formula: ; where, = Let t be the original interaction point. Let t be the smooth interaction point in frame t. The smoothing coefficient is preferably... The value can be between 0.25 and 0.45, with an initial value of 0.35. If the smoothing coefficient is too small, the response to real actions will be sluggish, while if the smoothing coefficient is too large, the jitter suppression will be insufficient. Therefore, the above range is preferred. By smoothing the interaction points, a stable trajectory can be output without significantly increasing the amount of computation, which is convenient for subsequent temporal feature extraction. Based on smooth interaction points, basic trajectory quantities such as object distance, displacement velocity and local jitter amplitude are further generated. The object distance represents the distance between the user's current operation position and the center of the virtual tree object. Displacement velocity indicates whether the user is currently passing by quickly, approaching slowly, pausing briefly, or leaving. The amplitude of local jitter indicates whether the user is still frequently probing while pausing. Among the basic trajectory quantities mentioned above, the displacement velocity is calculated using the following formula: ;in, and These are the smooth interaction points between two adjacent frames. The time interval between two adjacent frames. It is a two-dimensional Euclidean norm, preferably at 25 frames per second. It can be approximated as 0.04 seconds; The amplitude of localized jitter is determined by statistical analysis of recent data. The average deviation of frame smoothing interaction points from the local mean center is obtained, and the optimal value is obtained. The value can be 8 to 12, and the initial value can be 10. The three results of object distance, displacement velocity and local jitter amplitude will be cached in a unified manner as the basic trajectory description sequence, which will be used as the direct input of the interactive recognition module. It should be noted that this step does not immediately determine whether it is triggered, but rather establishes a standardized trajectory input system for virtual tree objects: that is, the input is the camera video stream, and the output is the object center point, interaction area, smooth interaction point, object distance, displacement speed, and local jitter amplitude.

[0017] In the interactive recognition module, a phase recognition model is established based on the basic trajectory description sequence. This model constructs phase features including object distance, displacement velocity, velocity change, local jitter amplitude, alignment deviation, and dwell count in the core confirmation area. It outputs standby state, approach state, probing state, candidate confirmation state, and departure state. Furthermore, it combines phase continuity judgment and phase regression judgment to output a graded visual pre-response. Specific content includes: This step involves performing phased visual interaction recognition of user actions around a virtual tree object. In the fragrance display scenario, the behavior around the virtual tree object is not a simple binary trigger or non-trigger, but usually goes through a continuous process of approaching the object, observing and probing around the edge of the object, converging the action to the center of the object, forming a short-term stable confirmation, and leaving the object. If this process is skipped and conclusions are drawn only based on whether the hand stops at a certain moment, the subsequent fragrance release is very likely to be mistried. This step converts the continuous trajectory output by the trajectory construction module into a phase sequence with clear behavioral semantics, providing context for the confirmation judgment of the fragrance release confirmation module. Main interaction area around the object Automatically generate peripheral proximity zone Exploration Zone and core confirmation area The construction phase state set includes: standby state, approach state, probing state, candidate confirmation state, and departure state; The standby state indicates that the trajectory has not yet formed an effective approach. The approach state indicates that the interaction point is continuously moving from the periphery of the object towards the object's interaction area. The probing state indicates that the interaction point pauses briefly, travels back and forth in a small range, corrects its direction, or explores the edge after entering the vicinity of the object. The candidate confirmation state indicates that the interaction point has further converged from the probing area to the vicinity of the core confirmation area, and its speed has decreased significantly and its jitter has decreased. The departure state indicates that the interaction point leaves the object area after completing one action. It is preferable to use a lightweight temporal model rather than a complex large visual language model or a three-dimensional skeleton model. In terms of model input, this step does not directly use the original image, but uses the basic trajectory description sequence already generated by the trajectory construction module. The model input is not high-dimensional pixel data, but low-dimensional temporal feature sequence. For frame t, the current feature vector can be constructed, including: object distance, displacement velocity, velocity change rate, local jitter amplitude, alignment deviation, and core confirmation area dwell count. The alignment deviation represents the normalized offset of the smooth interaction point relative to the object center point, which can be obtained by dividing the distance between the current smooth interaction point and the object center point by the equivalent radius of the main interaction region of the object. The core confirmation area dwell count represents the number of frames in which the interaction point falls into the core confirmation area in the most recent frames. The features of the most recent frames are combined in chronological order and used as the input window of the stage recognition model. The model output is the predicted probability of the five stages, and the stage with the highest probability is taken as the current candidate stage. For model training, it is preferable to use real interactive video data collected during prototype operation to construct an offline training set; Specifically, video data of hand movements from multiple participants is continuously collected, and the trajectory construction module automatically converts the video data into a basic trajectory description sequence. The basic trajectory description sequence includes the object distance change sequence, displacement velocity change sequence, local jitter amplitude sequence, alignment deviation sequence, and core confirmation area dwell time sequence; The basic trajectory description sequence is labeled with stage labels according to the time window to obtain training samples corresponding to the standby state, approach state, probing state, candidate confirmation state and departure state. To improve the generalization ability of the stage recognition model to different interaction rhythms and different operation styles, the training samples are preferably selected to include fast approach trajectory, slow approach trajectory, short stay and then leave trajectory and long trial trajectory. Preferably, all samples can be divided into a training set, a validation set, and a test set in a ratio of 7:2:1; During training, the weighted cross-entropy loss function is preferred, and higher weights are assigned to the categories corresponding to candidate confirmation states to reduce recognition bias caused by the small proportion of such samples. After training, the parameters of the model that performs well on the validation set are selected, solidified, and deployed into the online recognition process; Since stage boundaries are easily affected by instantaneous jitter in real operation, it is preferable to set the stage minimum duration rule and stage rollback rule after the model output; The minimum duration rule of a stage means that a stage is only truly recognized as the current stage if it remains the stage with the highest probability for several consecutive frames, and the stage will not be switched immediately due to changes in prediction in a single frame. The phase rollback rule is that if the current window is judged as a candidate confirmation state, but the interaction point obviously rolls back to the probing area in a very short window, and the displacement speed increases and the local jitter amplitude increases, it means that the candidate confirmation is only a temporary focus rather than a real confirmation. The phase should be rolled back to the probing state instead of continuing to advance towards confirmation judgment. By using minimum persistence rules and backoff rules, stage tag jitter can be significantly reduced, so that the aroma confirmation module obtains a more stable stage sequence instead of loose frame-by-frame tags. To avoid interpreting any brief pause by the user around the virtual tree object as a confirmation of fragrance release, after completing the minimum duration rule and the stage rollback rule, a graded visual pre-response is output for different stages. The graded visual pre-response does not directly trigger the atomizer to execute, but is used to provide feedback to the user on the current recognition result of the interaction state, while continuously accumulating temporal evidence for the fragrance release confirmation judgment in the subsequent fragrance release confirmation module. When in a near-term state, only cause the virtual tree object's canopy to sway slightly, the leaves to move slightly, or the light and shadow to flicker slightly, as a first-level visual pre-response; The role of the first-level visual pre-response is to indicate to the user that a hand movement has been detected entering the outer area of ​​the object interaction, thereby establishing a primary correspondence between the user's action and the virtual object's feedback, reducing the user's uncertainty about whether the interaction has been perceived; at the same time, the first-level visual pre-response only causes minor changes to the tree object, does not change the fragrance execution state, and does not issue a fogging start command, so it can complete object locking and trajectory continuity observation in the early stage of the interaction, avoiding subsequent false triggers when the user only passes by the outer area of ​​the object; When in the exploratory stage, the brightness of the tree crown edge can be further enhanced, and particle light spots or tree trunk texture can be highlighted as a secondary visual pre-response; The role of secondary visual pre-response is to convey to the user that the current action has changed from simple approach to active probing around the object, so that the user can perceive that the perception has transitioned from peripheral perception to object-level attention. The secondary visual pre-response can distinguish the probing phase from the approach phase on the software side, allowing the system to continue observing whether the user exhibits probing behaviors such as moving back and forth along the object edge, pausing briefly, or correcting their orientation during this phase. These behaviors are then used as the preceding context for the subsequent fragrance release confirmation module to make a confirmation judgment. When in the candidate confirmation state, the center of the tree canopy can produce a breathing-like brightness change, a slight enlargement of the central area, or a local fog visual cue, as a level three visual pre-response; The role of the three-level visual pre-response is to provide feedback to the user that the interaction trajectory has been identified and is focusing on the center of the object, and that the conditions for entering the fragrance release confirmation judgment have been initially met, so that the user can get a clearer visual prompt that it is about to be triggered. The Level 3 visual pre-response is not equivalent to the final execution result. Its essence is to mark the current trajectory as a candidate confirmation trajectory, so that the trajectory can be observed for a very short time window to see if it maintains low speed, low jitter, center alignment and no obvious backtracking. Only when the candidate confirmation state corresponding to the level 3 visual pre-response continues to meet the confirmation conditions in subsequent windows will the fragrance release confirmation module further calculate the fragrance release confirmation score and decide whether to output the atomization execution command. Therefore, the three-level visual pre-response logically serves as a bridge between stage recognition and confirmation judgment, which can enhance the user's continuous interactive perception and provide clear candidate confirmation entry points for subsequent execution control. It should be noted that the first-level visual pre-response, second-level visual pre-response, and third-level visual pre-response are in a progressive relationship in terms of function: the first-level visual pre-response is used to confirm that the user's action has entered the object interaction range, the second-level visual pre-response is used to confirm that the user is actively exploring around the object, and the third-level visual pre-response is used to confirm that the user's action has the characteristics of a candidate confirmation. The interaction recognition module forms a visual feedback hierarchy from weak to strong and from coarse to fine, enabling the virtual tree object to change gradually with the interaction state, rather than outputting the result all at once when finally confirmed. This progressive pre-response method not only improves the perceptibility of the interaction process, but also avoids directly equating the stage recognition result with the fragrance execution result, thereby achieving a balance between user experience and control accuracy. Based on the standardized trajectory input output by the trajectory construction module, the hand movements around the virtual tree object are divided into temporal stages, and hierarchical visual pre-responses are output for different stages. This allows for the gradual identification of the evolution of user actions from approach and probing to candidate confirmation without prematurely activating the atomizer. The direct result is that the fragrance release confirmation module no longer receives loose frame-by-frame position labels, but rather a stable sequence of candidate interactions with clear stage semantics, stage duration information, and pre-response level information. This enables a more accurate distinction between observation and lingering and actual fragrance release confirmation actions in subsequent confirmation decisions. Actions around virtual tree objects are analyzed into multiple consecutive stages, such as approach state, probing state, and candidate confirmation state. A hierarchical visual pre-response mechanism corresponding to each stage is introduced, transforming the interaction process from a trigger-based response to a progressive understanding and feedback. Its advantages are twofold: firstly, it can significantly reduce the probability of false triggers caused by brief pauses, probing swings, or accidental entry into the object area; secondly, it can enhance the user's perception and continuity of the recognition process, allowing the user to receive progressively enhanced visual feedback before actually entering the confirmation action, thereby improving the stability, interpretability, and immersive experience of the entire fragrance display interaction control.

[0018] In the aroma release confirmation module, when the candidate confirmation state meets the persistence condition, convergence features, stability features, alignment features, dwell features, and stage confidence features are extracted to establish a confirmation model and calculate the confirmation score. The confirmation score is combined with the number of frames in the core confirmation area to determine the atomization execution mode. Specific content includes: After the interaction recognition module has output a stable phase sequence and clearly identified the current action as a candidate confirmation state, it comprehensively utilizes the convergence degree, stability degree, jitter degree, alignment degree and phase probability within the candidate confirmation window to construct a trainable, interpretable and easy-to-implement fragrance release confirmation model. Furthermore, the confirmation result is mapped to three execution modes: light fog, standard fog or enhanced fog, to achieve closed-loop control between visual understanding and fragrance execution. This step is initiated only after the candidate confirmation status output by the interaction recognition module continuously meets the minimum number of consecutive frames. Confirmation features are extracted from the candidate confirmation window. Confirmation features include the following categories: convergence features, stability features, alignment features, dwell features, and stage confidence features. Convergence features are the speed change trends before and after entering the candidate confirmation state, used to describe whether the user has experienced a natural deceleration convergence from fast to slow. Stability features refer to the average deviation of smooth interaction points within the candidate confirmation window from the local mean center, which is used to describe whether the pause is stable. Alignment feature is the average deviation of the interaction point within the candidate confirmation window relative to the center point of the object, used to describe whether the user is truly focused on the central area of ​​the virtual tree object. The dwell feature, or core confirmation zone duration frame count, describes whether the current action maintains a center dwell for a sufficient amount of time. The stage confidence feature, which is the candidate confirmation stage probability output by the interaction recognition module, is used to reflect whether the time series model is stable enough to consider the action as a candidate confirmation rather than a short-term mistaken entry. These five features allow us to extract a relatively complete behavioral summary from a single window: whether the user has converged, whether the stop is stable enough, whether the stop is in the correct position, whether the stop lasts long enough, and the confidence level of the time series model in the candidate confirmation. Preferably, a lightweight confirmation model in the form of logistic regression is used, and the confirmation feature vector of the current window is denoted as... And calculate the fragrance release confirmation score according to the following formula. : ;in, This is the confirmation feature vector corresponding to the current candidate confirmation window. To confirm the model weight vector, For bias terms, For the Sigmoid function, This is the probability score for the current action being a confirmation action of genuine fragrance release; like The closer the value is to 1, the more likely the current action is a real fragrance release confirmation action formed by the user in front of the virtual tree object; if A lower value indicates that the current action is more likely to be a lingering or tentative stop around the object. The advantage of logistic regression is that the weight of each confirming feature has a clear meaning. In terms of model training, binary classification training is preferred based on candidate confirmation window samples collected during the actual operation of the prototype. Specifically, video data of hand movements of multiple users in the interactive area in front of the display screen is continuously collected, and the video data is converted into candidate confirmation window samples by the trajectory construction module and the interaction recognition module. Each candidate confirmation window sample corresponds to a set of confirmation feature parameters, which include the velocity change before and after entering the candidate confirmation state, the average object distance within the candidate confirmation window, the average jitter amplitude, the alignment deviation, the number of frames in the core confirmation area, and the probability of the candidate confirmation stage. The candidate confirmation window samples are processed offline and labeled as real confirmation samples or non-real confirmation samples based on the trajectory convergence characteristics, dwell stability characteristics and subsequent behavior characteristics corresponding to the candidate confirmation window; the label value of real confirmation samples is 1, and the label value of non-real confirmation samples is 0. Preferably, in order to improve the model's adaptability to different interaction rhythms and different operation styles, the training samples should simultaneously include fast-converging samples, slow-converging samples, short-term pause-and-exit samples, and long-term exploratory samples. All samples can be divided into training set, validation set and test set in a 7:2:1 ratio; Furthermore, the ratio of real confirmed samples to non-real confirmed samples is preferably controlled between 1:1 and 1:2 to mitigate the impact of class imbalance on training results. After training, the model parameters that show the best overall recognition performance on the validation set are selected, solidified, and deployed into the online recognition process; It should be noted that, since the confirmation model used in this step is a lightweight binary classification model, the model input is only the low-dimensional confirmation feature parameters corresponding to the candidate confirmation window. After receiving the confirmed score Then, based on the confirmed score The combination relationship with the number of frames continuously in the core confirmation area determines the fogging execution mode corresponding to the current interaction; Two confirmation thresholds and one duration frame threshold are set. The lower confirmation threshold is used to distinguish between actions that have not reached a valid confirmation and actions that have entered the range of executable confirmations. The higher confirmation threshold is used to distinguish between normal confirmation actions and high-confidence confirmation actions. The duration frame threshold is used to distinguish between actions with short confirmation durations and actions with long confirmation durations. Specifically, when the confirmation score is lower than the lower confirmation threshold, regardless of whether the core confirmation area lasts for a long time, it is determined that the current action has not yet met the effective fragrance release confirmation conditions. Only the current visual pre-response level is maintained, and no fogging start command is issued. When the confirmation score itself is low, it means that the current action is still insufficient to support the true confirmation conclusion in terms of convergence, stability, center alignment, or candidate confirmation stage probability. At this time, even if the lasting frame is long, it is more likely to be a viewing pause or hesitation pause, and it is not appropriate to directly trigger fragrance release. When the confirmation score is not lower than the lower confirmation threshold but lower than the higher confirmation threshold, and the number of frames in the core confirmation area is lower than the number of frames in the continuous frame threshold, the current action is determined to be a light confirmation action, and a light fog mode is output. This indicates that the current action has certain confirmation characteristics, but the duration is still relatively short, or that although the user has clear approach and focus behavior, the stay has not yet developed to a high level of confidence confirmation. Therefore, it is suitable to output a weak fragrance feedback to indicate to the user that the initial confirmation intention has been recognized, while avoiding the overall display rhythm due to excessive release. When the confirmation score is not lower than the lower confirmation threshold but lower than the higher confirmation threshold, and the number of frames in the core confirmation area is not lower than the number of frames in the core confirmation area, the current action is determined to be a stable confirmation action, and the standard fog mode is output. Although the confirmation score has not yet reached the highest level, since the user has maintained a stable stay in the core confirmation area for a long time, it indicates that the current action has clearly shown the intention to confirm. Therefore, the standard intensity of fragrance release can be used to obtain a more complete physical feedback effect. When the confirmation score is not lower than the higher confirmation threshold, regardless of whether the number of frames in the core confirmation area has reached the number of frames in the core confirmation area, the current action is determined to be a high-confidence confirmation action, and the enhanced fog mode is output. When the confirmation score has reached the higher confirmation threshold, it means that the current action has formed strong comprehensive confirmation characteristics in terms of speed change before and after entering the core confirmation area, stability of stay, degree of center alignment and semantic consistency of stage. Even if the stay time is relatively short, it is enough to show that the user has a strong intention to release fragrance. Therefore, the enhanced fog mode can be directly entered to reflect the rapid response to the high-confidence confirmation action. The preferred atomization execution modes include four types: the first is the non-execution mode, which means that only the current visual pre-response is maintained without releasing fragrance; the second is the light fog mode, which means that fragrance is released for a shorter duration and at a lower intensity; the third is the standard fog mode, which means that fragrance is released for a normal duration and at a normal intensity; and the fourth is the enhanced fog mode, which means that fragrance is released for a longer duration or at a higher intensity. Preferably, the lower confirmation threshold can be 0.72, the higher confirmation threshold can be 0.88, and the continuous frame count threshold can be 10 frames; under the video acquisition condition of 25 frames / second, the 10 frames correspond to a core confirmation area continuous dwell time of approximately 0.4 seconds; Furthermore, the execution time of the light fog mode can be 0.6 seconds, the execution time of the standard fog mode can be 1.0 seconds, the execution time of the enhanced fog mode can be 1.8 seconds, and the uniform cooldown time after execution can be 5 seconds; It should be noted that the above four scenarios cover all intervals corresponding to the combination of confirmation score and core confirmation area continuous frame count: when the confirmation score is lower than the lower confirmation threshold, it is uniformly classified into non-execution mode; when the confirmation score is between the lower and higher confirmation thresholds, it is further classified into light fog mode or standard fog mode based on whether the continuous frame count reaches the continuous frame count threshold; when the confirmation score is not lower than the higher confirmation threshold, it is uniformly classified into enhanced fog mode. In this way, the problem of some intervals not being clearly covered in the original segmentation rules is avoided, and the execution mapping relationship is made simpler and clearer, which is convenient for software implementation and manual disclosure. Based on the above mapping rules, different intensity fogging modes are output according to the degree of confirmation and the degree of dwell time. Even if both actions are real confirmations, if one of them is only a light and short confirmation, only a light fog mode is output; if the other action is more stable and dwells longer, a standard fog or enhanced fog mode is output, which better meets the needs of display-type applications for delicate feedback. Transforming the traditional single-camera interaction's pause trigger into a multi-level control process involving stage sequence context, multi-confirmation feature fusion, and segmented atomization mode selection, using lightweight features and logistic regression models, without relying on complex hardware and large neural networks, and capable of being deployed directly on ordinary host computers, is not only the core innovation of the entire solution but also the key bridge that truly transforms the output of the first two steps into fragrance execution control.

[0019] In the threshold adjustment module, the confirmation score, the number of frames in the core confirmation zone, the fogging execution mode, and the short-term subsequent behavior are recorded. Samples located in the vicinity of the lower confirmation threshold are filtered. Based on the comparison results between suspected premature triggering samples and suspected overly strict samples, the lower confirmation threshold is narrowed and corrected. The specific content includes: After the aroma release confirmation module has determined the fogging execution mode corresponding to the current interaction based on the confirmation score and the number of frames in the core confirmation area, this step no longer adjusts the object center point, the minimum number of frames in the stage, or other multiple parameters at the same time. Instead, it only makes fine-grained corrections to the lower confirmation threshold in the aroma release confirmation module, so as to make more accurate corrections to the boundary between actions that have not achieved effective confirmation and mild confirmation actions that enter the light fog mode, while keeping the main structure from the trajectory construction module to the aroma release confirmation module unchanged. A lower confirmation threshold directly determines whether the confirmation score enters the executable confirmation range and directly affects the trigger boundary of the fog mode. Compared with adjusting multiple parameters at the same time, making small-step corrections only to the lower confirmation threshold is more conducive to maintaining the overall stability of the mapping relationship between the main area of ​​object interaction, candidate confirmation status, confirmation score calculation method and fog execution mode. Specifically, after each candidate confirmation action is completed, the confirmation score, the number of frames in the core confirmation area, the final fogging execution mode, and the short-term follow-up behavior are recorded for that interaction. The short-term follow-up behavior after execution is used to determine whether the action, which is near a low confirmation threshold, has been correctly interpreted. Preferably, when an interaction is determined to be in light fog mode by the fragrance confirmation module, if the user leaves the main interaction area of ​​the object naturally within a short period of time and does not return to the probing area or the core confirmation area to continue probing, then the interaction is recorded as a valid light confirmation sample. If an interaction is determined to be in light fog mode by the fragrance confirmation module, and the user returns to the probing area or core confirmation area within a short period of time and continues to make similar probing actions as before, then the interaction will be recorded as a suspected premature triggering sample. If an interaction does not enter the light fog mode, but the user forms a more stable candidate confirmation state again in a short period of time and eventually enters the light fog mode or standard fog mode, then the previous interaction is recorded as a suspected strict sample. Furthermore, it is preferable to only count samples whose confirmation scores fall within the range of the lower confirmation threshold, rather than uniformly correcting all samples; The adjacent interval refers to the range within which the difference between the confirmation score and the current lower confirmation threshold falls, and the absolute difference between the confirmation score and the lower confirmation threshold does not exceed 0.05. Actions with confirmation scores significantly below the lower confirmation threshold are more likely to be viewed or ignored; actions with confirmation scores significantly above the higher confirmation threshold are already considered high-confidence actions; the actions most prone to boundary misjudgment are those with confirmation scores just around the lower confirmation threshold. In terms of correction methods, a small step-limit correction strategy is preferred. Specifically, when the number of suspected premature triggering samples in the neighboring interval samples is higher than the number of suspected overly strict samples in a recent period, it indicates that the current lower confirmation threshold is too low. At this time, the lower confirmation threshold is corrected upward by a small step. When the number of suspected overly strict samples in the neighboring interval is higher than the number of suspected premature trigger samples, it indicates that the current lower confirmation threshold is too high. At this time, the lower confirmation threshold is adjusted downward by a small step. When the number of samples in the two categories is similar, the lower confirmation threshold is maintained. Preferably, the step size for each correction can be 0.005 to 0.02, the initial value can be 0.01, and the lower confirmation threshold after correction is always limited to the preset upper and lower bounds, for example, the lower bound is 0.60 and the upper bound is 0.90. This step only fine-tunes the lower confirmation threshold parameter in the aroma release confirmation module, without changing the existing definitions of the higher confirmation threshold, the continuous frame threshold, and the fog execution mode. Therefore, it can finely optimize the trigger boundary of the light fog mode without destroying the overall structure of the first three steps. It should be noted that the threshold adjustment module does not retrain the stage recognition model in the interaction recognition module, nor does it change the confirmation score calculation model in the fragrance confirmation module. Instead, it only makes fine corrections to the lower confirmation threshold in the fragrance confirmation module while keeping the existing terminology, parameter system, and processing flow from the trajectory construction module to the fragrance confirmation module unchanged.

[0020] In the context of fragrance display and interaction based on camera visual input, the system no longer directly interprets a user's brief pause around a virtual tree object as a fragrance release confirmation action. Instead, it accurately distinguishes between observation pauses, tentative pauses, and actual confirmation actions through standardized trajectory input, phased visual interaction recognition, confirmation score judgment, and fine correction with a low confirmation threshold. This reduces the probability of atomizer false activation and improves the stability, interpretability, and feasibility of fragrance display and interaction control.

[0021] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0022] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0023] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and inventive constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0024] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0025] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0026] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A fragrance display and interactive control system based on camera visual input, characterized in that, Signal connections between modules; The trajectory construction module obtains the calibration relationship between the display screen and the camera, determines the object center point and the main interaction area of ​​the virtual tree object, generates the outer approach area, the probing area and the core confirmation area, extracts the main hand area and the original interaction points, smooths the original interaction points, calculates the object distance, displacement speed and local jitter amplitude, and forms the basic trajectory description sequence. Interactive recognition module: Based on the basic trajectory description sequence, a phase recognition model is established, which constructs phase features including object distance, displacement velocity, velocity change, local jitter amplitude, alignment deviation and core confirmation area dwell count, and outputs standby state, approach state, probing state, candidate confirmation state and departure state, and outputs hierarchical visual pre-response by combining phase continuity judgment and phase back-off judgment. Aroma release confirmation module: When the candidate confirmation state meets the continuous condition, it extracts convergence features, stability features, alignment features, dwell features and stage confidence features, establishes a confirmation model and calculates the confirmation score, and combines the confirmation score with the number of continuous frames in the core confirmation area to determine the atomization execution mode; Threshold adjustment module: Records confirmation score, number of frames in the core confirmation area, fogging execution mode and short-term subsequent behavior after execution, filters samples in the range of the lower confirmation threshold, and performs amplitude limiting correction on the lower confirmation threshold based on the comparison results of suspected premature triggering samples and suspected overly strict samples.

2. The fragrance display and interactive control system based on camera visual input according to claim 1, characterized in that, In the trajectory construction module, the specific method for generating the outer approach zone, probing zone, and core confirmation zone around the main interaction area of ​​the object is as follows: The outer approach area is obtained by proportionally expanding the outer boundary of the main interaction area of ​​the object, and the probing area and the core confirmation area are obtained by proportionally shrinking the inner boundary of the main interaction area of ​​the object.

3. The fragrance display and interactive control system based on camera visual input according to claim 2, characterized in that, Extracting the main hand area includes: The current frame is compared with the background reference to obtain motion candidate regions. Skin color filtering is performed within the motion candidate regions. The continuity is judged by combining the hand position of the previous frame. The connected region with the largest area and the highest continuity with the position of the previous frame is retained as the main hand region.

4. The fragrance display and interactive control system based on camera visual input according to claim 2, characterized in that, The original interaction points are smoothed using an exponential smoothing method, and the object distance, displacement velocity, and local jitter amplitude are calculated based on the smoothed interaction points. The displacement velocity is obtained by dividing the distance between two adjacent smooth interaction points by the time interval, and the local jitter amplitude is obtained by statistically analyzing the average deviation of the smooth interaction points in the most recent few frames from the local mean center.

5. The fragrance display and interactive control system based on camera visual input according to claim 1, characterized in that, In the interactive recognition module, the hierarchical visual pre-response includes a first-level visual pre-response, a second-level visual pre-response, and a third-level visual pre-response; When in a near-term state, a level one visual pre-response is output, causing the canopy of the virtual tree object to sway slightly. When in a probing state, a secondary visual pre-response is output, which enhances the brightness of the canopy edge or produces particle light spots; When in the candidate confirmation state, a level 3 visual pre-response is output, causing a breathing-like brightness change in the center of the tree canopy.

6. The fragrance display and interactive control system based on camera visual input according to claim 5, characterized in that, The stage persistence determination is that a certain stage is only identified as the current stage if it remains the highest probability stage for a consecutive number of frames. The stage rollback is determined when the candidate confirmation state rolls back to the test area at the interaction point in the subsequent window and the displacement velocity increases and the local jitter amplitude increases.

7. The fragrance display and interactive control system based on camera visual input according to claim 1, characterized in that, In the fragrance release confirmation module, the convergence feature is the velocity change trend before and after entering the candidate confirmation state, the stability feature is the average deviation of the smooth interaction points in the candidate confirmation window around the local mean center, the alignment feature is the average deviation of the interaction points in the candidate confirmation window relative to the center point of the object, the dwell feature is the number of frames in the core confirmation area, and the stage confidence feature is the candidate confirmation stage probability output by the interaction recognition module.

8. The fragrance display and interactive control system based on camera visual input according to claim 7, characterized in that, The atomization execution modes include no execution mode, light fog mode, standard fog mode, and enhanced fog mode; When the confirmation score is lower than the lower confirmation threshold, it is determined to be in non-execution mode; When the confirmation score is not lower than the lower confirmation threshold and lower than the higher confirmation threshold, and the number of consecutive frames in the core confirmation area is lower than the number of consecutive frames, it is determined to be light fog mode. When the confirmation score is not lower than the lower confirmation threshold and is lower than the higher confirmation threshold, and the number of consecutive frames in the core confirmation area is not lower than the number of consecutive frames, it is judged as standard fog mode. When the confirmation score is not lower than the higher confirmation threshold, it is determined to be enhanced fog mode.

9. A fragrance display and interactive control system based on camera visual input according to claim 1, characterized in that, In the threshold adjustment module, the nearest range refers to the absolute difference between the confirmation score and the current lower confirmation threshold not exceeding a preset small range; The suspected premature triggering samples are those where, after being identified as being in light fog mode, users returned to the probing area or core confirmation area within a short period of time and continued to probe. The suspected overly strict sample is the previous interaction sample where the user did not enter the light fog mode but formed a more stable candidate confirmation state again in a short period of time and finally entered the light fog mode or standard fog mode.

10. A fragrance display and interactive control system based on camera visual input according to claim 9, characterized in that, A small-step amplitude limiting correction strategy is adopted for the lower confirmation threshold: when the number of suspected premature triggering samples in the neighboring interval is higher than the number of suspected overly strict samples, the lower confirmation threshold is adjusted upward by a small step. When the number of suspected overly strict samples exceeds the number of suspected prematurely triggered samples, the lower confirmation threshold will be adjusted downward by a small step. The revised lower confirmation threshold is limited to the preset upper and lower bounds.