Work scene visual data analysis system and method
By monitoring head rotation parameters to generate freeze commands, defining depth ranges and recording object information in stages, and adjusting parameters based on user feedback, the problem of missed and false detections of objects caused by changes in field of view in existing technologies has been solved, achieving high-precision visual analysis of the work scene.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN UNICAIR COMM TECH CO LTD
- Filing Date
- 2026-05-26
- Publication Date
- 2026-06-26
AI Technical Summary
Existing visual analysis technologies for operational scenarios are unable to adapt to changes in field of vision caused by head rotation, resulting in missed or false detections of objects. Furthermore, they lack tiered processing for objects of different sizes, failing to meet the requirements for high-precision operations.
Freeze commands are generated by monitoring the head rotation angular velocity, duration, and static duration. The effective depth range is defined and object information is recorded in stages. Parameters are dynamically adjusted in conjunction with user feedback to eliminate operation occlusion and out-of-depth offset, ensuring accurate determination of disappearance events.
It improves the accuracy of object recognition and the determination of disappearance events, reduces the false alarm rate, and enhances operational efficiency and safety.
Smart Images

Figure CN122293980A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of augmented reality technology, specifically to a visual data analysis system and method for work scenarios. Background Technology
[0002] In industrial operations and precision manufacturing scenarios, visual data analysis is crucial for ensuring operational safety and accuracy. Existing visual analysis technologies for these scenarios often employ fixed parameters for field-of-view capture and object recognition, making it difficult to adapt to changes in the field of view caused by operator head movements. This leads to issues such as inappropriate snapshot freezing timing and missed or false detections due to fixed depth-of-field ranges. Furthermore, current technologies lack tiered processing for object size, offer no differentiated design for recognition accuracy across different object sizes, and lack an adaptive parameter correction mechanism based on user feedback. This results in low accuracy in detecting missing objects and inadequate alert strategies, failing to meet the visual analysis needs of high-precision operational scenarios and hindering improvements in operational efficiency and safety. Summary of the Invention
[0003] The purpose of this invention is to provide a visual data analysis system and method for work scenarios to solve the problems raised in the prior art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: a method for visual data analysis of a work scene, the method comprising: S100: Monitor the peak angular velocity, duration, and rest time after head rotation. When the peak angular velocity exceeds a first preset threshold, the duration of rotation exceeds a second preset threshold, and the rest time after rotation exceeds a third preset threshold, generate a freeze command to freeze the visual field semantic snapshot and record the rotation amplitude simultaneously. Head rotation is monitored in real time through a wearable visual acquisition device. The visual field semantic snapshot is a comprehensive feature set of object category, position, and depth information within the current visual field. The rotation amplitude is quantified by the change in head rotation angle. S200. Centered on the depth of the reference object on the work surface, the near-end boundary is defined by the near-end offset, and the far-end boundary is defined by the far-end offset. The depth range between the near-end boundary and the far-end boundary constitutes the effective depth range. The reference object on the work surface is a calibrated object with a relatively fixed position in the work scene. The effective depth range is the depth range that is effectively focused on during the visual analysis process. The recording accuracy is reflected in the completeness of object feature extraction and the accuracy of position positioning. When S300 detects visual field regression, it sequentially determines the global motion of the image, gaze point drift, and overlap with the reference object. When all of these meet the corresponding stability threshold, it initiates new frame analysis, spatially matches the new frame with the visual field semantic snapshot, and verifies the number of consecutive frames of disappearance confirmation to check for undetected targets, marking them as disappearance candidates. The global motion of the image reflects the overall displacement of the acquired image, the gaze point drift reflects the shift of the visual attention center, and the overlap with the reference object reflects the positioning accuracy after visual field regression. Spatial matching achieves consistency comparison based on coordinate position and feature information. S400: Obtain a list of objects marked as disappearance candidates, exclude operation occlusion and depth-of-field offset, analyze the disappearance event and assign confidence, and output visual cues; confidence is used to characterize the credibility of the disappearance event judgment result, operation occlusion is the visual occlusion of the target object by the human body or tool during the operation, and depth-of-field offset is the visual undetected state formed by the target object leaving the effective depth-of-field range. S500 determines the authenticity of prompts based on user responses, and corrects near-end offset, far-end offset, number of consecutive frames for disappearance confirmation, and stability threshold. User responses include confirmation and rejection of prompt results. The correction process dynamically adjusts the corresponding parameter values to make the subsequent visual analysis process more in line with actual operating scenarios.
[0005] According to the above scheme, step S100 includes: S110. Obtain the peak angular velocity, duration, and static duration during head rotation; the peak angular velocity is the maximum value of the angular velocity during head rotation, the duration of rotation is the duration during which the angular velocity exceeds the reference value, and the static duration after rotation is the duration during which the image remains stable after the rotation ends. S120. Compare the peak angular velocity with the first preset threshold, compare the rotation duration with the second preset threshold, and compare the stationary duration after rotation with the third preset threshold. S130. When the peak angular velocity exceeds the first preset threshold, the rotation duration exceeds the second preset threshold, and the static duration after rotation exceeds the third preset threshold, a freeze command is generated to freeze the semantic snapshot of the field of view. The value of the third preset threshold is correlated with the value of the peak angular velocity in the same direction. When the peak angular velocity changes in the positive direction, the third preset threshold is adjusted in the positive direction synchronously, and when the peak angular velocity changes in the negative direction, the third preset threshold is adjusted in the negative direction synchronously. The freeze command is used to control the visual acquisition device to stop updating the field of view feature information and fix the semantic data at the current moment. S140. Synchronously record the amplitude of head rotation during the rotation process.
[0006] According to the above scheme, step S200 includes: S210. Identify the reference object of the work surface in the current field of view, and obtain the visual depth value of the reference object of the work surface as the depth center; the visual depth value is used to characterize the distance information between the reference object of the work surface and the vision acquisition device; S220. Using the depth center as a reference, the near-end boundary is defined using the near-end offset, and the far-end boundary is defined using the far-end offset. The near-end offset is adaptively adjusted based on the rotation amplitude, and the far-end offset is determined by the current object size attribute, which is a pre-configured process parameter. The depth range between the near-end boundary and the far-end boundary constitutes the effective depth range. S230. For objects located within the effective depth-of-field range, record them in a hierarchical manner according to their size. Specifically: objects within the first size range are recorded with two-dimensional coordinates and contour feature points; objects within the second size range are recorded with two-dimensional coordinates and category labels; and objects within the third size range are recorded with existence markers. The first size range is smaller than the second size range, and the second size range is smaller than the third size range. Two-dimensional coordinates represent the object's planar position in the captured image; contour feature points are key positioning points of the object's outline; category labels are identification information of the object's type; and existence markers are indicators of whether the object exists. The size classification threshold used for hierarchical recording is correlated in the same direction as the width of the effective depth-of-field range. When the range width changes positively, the upper limit of the first size range is adjusted positively in sync. The size classification threshold is used to distinguish objects of different size levels, and changes in the range width directly alter the criteria for determining object size classification.
[0007] According to the above scheme, step S300 includes: S310. Acquire the image data after the field of view regression, and determine the global motion state of the image, the gaze point drift state and the overlap between the gaze point and the reference object on the work surface in sequence. The field of view regression is the state in which the visual acquisition device returns to the target work field of view after the head rotation ends. The global motion state of the image is determined by the inter-frame pixel displacement amount, and the gaze point drift state is determined by the displacement deviation amount of the attention center. S320, Preset a first stability judgment threshold, a second stability judgment threshold, and a third stability judgment threshold; the values of the first stability judgment threshold and the second stability judgment threshold are inversely related to the rotation amplitude value; the value of the third stability judgment threshold is in the same direction as the rotation amplitude value. S330. When the global motion state of the screen meets the first stability judgment threshold, the gaze point drift state meets the second stability judgment threshold, and the overlap between the gaze point and the reference object on the work surface meets the third stability judgment threshold, start the semantic analysis of the new frame. S340. Perform spatial dimension matching between the semantic analysis results of the new frame and the visual semantic snapshot, and retrieve objects in the snapshot that were not detected in the new frame. Spatial dimension matching is achieved by combining planar coordinates and depth information. Objects in the snapshot that were not detected are targets that exist in the snapshot but are not identified in the new frame. S350. Perform continuous frame verification on undetected objects. Only if the object is not detected in any new frame of the consecutive disappearance confirmation frame number is the object marked as a disappearance candidate. The value of the consecutive disappearance confirmation frame number is inversely correlated with the field of view stability.
[0008] According to the above scheme, step S400 includes: S410. Obtain a list of objects marked as disappearance candidates, and perform operation occlusion exclusion judgment for each disappearance candidate object. The operation occlusion exclusion judgment includes determining whether the degree of overlap between the expected coordinate position of the disappearance candidate object in the snapshot and the bounding box of the large object identified in the semantic analysis result of the new frame meets the coverage judgment tolerance. If it meets the tolerance and the category of the covered object is a tool or a hand, then the disappearance candidate object is removed from the list and judged as operation occlusion. The bounding box is the area surrounded by the shape of the identified object in the new frame. The coverage judgment tolerance is the allowable range of positional deviation of the occlusion relationship. The delineation accuracy is reflected in the accuracy of the boundary definition of the depth range. S420. For disappearing candidate objects that remain in the list after occlusion exclusion, perform out-of-field offset determination. Out-of-field offset determination includes retrieving object features in the semantic analysis results of the new frame that are located in the vicinity of the effective depth range boundary. If there are contour features or category labels that match the disappearing candidate object, the disappearing candidate object is removed from the list and determined to be out-of-field offset. The vicinity region is the extended detection region outside the effective depth range boundary, used to determine whether the object deviates only from the effective depth range. S430. After occlusion exclusion and depth-of-field offset determination, the disappearing candidate objects that remain in the list are confirmed as disappearance events and assigned a confidence level, and a visual cue is output. The confidence level and the size range of the disappearing candidate objects are used to determine the intensity and direction of the visual cue. The visual cue presents the disappearance event information to the user through a visual identifier, and the cue points to the expected location of the corresponding disappearing object. The intensity of the cue is distinguished by the way the identifier is displayed.
[0009] According to the above scheme, step S500 includes: S510. Obtain the user's response to the visual cues, and determine whether the result of the disappearance event is true based on the response behavior; S520. When the disappearance event is not determined to be true and the disappearance candidate object has matching features in the area near the boundary of the effective depth range in the depth-of-field offset determination, the near-end offset and far-end offset are corrected and the coverage of the effective depth range is adjusted. S530. When an object within the first size range is determined to be a disappearance event but has not actually disappeared, the number of consecutive frames for disappearance confirmation is corrected. S540. When the semantic analysis of a new frame is not started normally after the view regression, the first stability judgment threshold, the second stability judgment threshold and the third stability judgment threshold are corrected and the stability judgment conditions are adjusted. S550, the magnitude of the correction is correlated in the same direction as the degree of deviation of the judgment result.
[0010] A visual data analysis system for a work scene, comprising: a snapshot freezing module, a depth-of-field segmentation module, a stability determination module, an occlusion removal module, and a feedback correction module; The snapshot freeze module is used to monitor the peak value of the head rotation angular velocity, its duration, and the duration of stillness after rotation. When all three values meet the corresponding thresholds, a semantic snapshot of the visual field is frozen, and the rotation amplitude is recorded simultaneously. The depth of field segmentation module is used to define the effective depth of field range centered on the depth of the reference object on the working surface, and to record object information within the range according to object size. The stability determination module is used to sequentially determine the global motion of the screen, the gaze point drift, and the degree of overlap with the reference object. After all these conditions meet the corresponding stability determination thresholds, a new frame analysis is initiated. The new frame is matched with the snapshot space. The number of consecutive frames of disappearance confirmation is used to verify the detection of undetected targets and mark disappearance candidates. The number of consecutive frames of disappearance confirmation decreases adaptively with the stability of the field of view. The occlusion exclusion module is used to obtain a list of objects marked as disappearance candidates, exclude operation occlusion and out-of-field offset, confirm the disappearance event and assign confidence level; The feedback correction module is used to determine the authenticity of the prompts based on the user's response, and to correct the near-end offset, far-end offset, number of consecutive frames of disappearance confirmation, and stability judgment threshold accordingly.
[0011] According to the above scheme, the depth-of-field segmentation module includes an offset adjustment unit and a hierarchical recording unit; The offset adjustment unit is used to define the near boundary using the near offset and the far boundary using the far offset, with the depth range between the near boundary and the far boundary constituting the effective depth range. The hierarchical recording unit is used to record objects located within the effective depth range according to their size. The first size range records the two-dimensional coordinates and contour feature points of objects, the second size range records the two-dimensional coordinates and category labels of objects, and the third size range records the existence markers of objects.
[0012] According to the above scheme, the stability determination module includes a threshold adjustment unit and a candidate label unit; The threshold adjustment unit is used to preset a first stability judgment threshold, a second stability judgment threshold, and a third stability judgment threshold; the first stability judgment threshold is used to determine the global motion state of the screen, the second stability judgment threshold is used to determine the gaze point drift state, and the third stability judgment threshold is used to determine the overlap. The candidate marking unit is used to start semantic analysis of the new frame when the global motion state, gaze point drift state and overlap degree of the image all meet the corresponding thresholds. The semantic analysis result of the new frame is matched with the visual semantic snapshot space, and continuous frame verification is performed on undetected objects. The object is marked as a disappearance candidate only when it has not been detected in the new frames of the consecutive disappearance confirmation.
[0013] According to the above scheme, the occlusion removal module includes an occlusion determination unit and a prompt confirmation unit; The occlusion determination unit is used to perform operation occlusion exclusion determination for each disappearing candidate object. It determines whether the degree of overlap between the expected coordinate position of the object in the snapshot and the bounding box of the large object identified in the semantic analysis result of the new frame meets the coverage determination tolerance. If it meets the tolerance and the covered object category is a tool or a hand, it is removed from the list and determined as operation occlusion. The unit then performs out-of-field offset determination on the objects that remain after operation occlusion exclusion. It searches for features in the region near the boundary of the effective depth range in the new frame. If there are matching contour features or category labels, the objects are removed from the list and determined as out-of-field offset. The prompt confirmation unit is used to identify objects that remain after double exclusion as disappearance events and assign them a confidence level.
[0014] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention establishes a multi-parameter joint triggering logic for head rotation angular velocity, duration, and static duration after rotation, and adjusts the static duration threshold in the same direction as the peak angular velocity, effectively distinguishing between brief saccades and real line-of-sight departure. 2. This invention sequentially determines that the global motion of the screen, the gaze point drift, and the overlap with the reference object on the work surface all meet the corresponding thresholds before starting the semantic analysis of the new frame. This ensures that the disappearing object comparison is only performed after the screen is fully stable and the gaze is re-anchored to the work surface, thus improving the accuracy of the comparison timing and the reliability of the comparison results. 3. The present invention adopts a progressive elimination mechanism of operation occlusion and out-of-field offset. Before confirming the disappearance event, it first eliminates the situation of normal occlusion caused by tools and hands and the situation of objects being offset out of the depth of field, thus avoiding misjudging normal operation actions as objects disappearing and greatly reducing the false alarm rate. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating the steps of a visual data analysis method for a work scenario according to the present invention. Figure 2This is a schematic diagram of the structure of a visual data analysis system for work scenarios according to the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Example: Figures 1-2 As shown, the present invention provides a technical solution, a method for visual data analysis of a work scene, the method comprising the following steps: S100: Monitor the peak angular velocity, duration, and rest time after head rotation. When the peak angular velocity exceeds a first preset threshold, the duration of rotation exceeds a second preset threshold, and the rest time after rotation exceeds a third preset threshold, generate a freeze command to freeze the visual field semantic snapshot and record the rotation amplitude simultaneously. Head rotation is monitored in real time through a wearable visual acquisition device. The visual field semantic snapshot is a comprehensive feature set of object category, position, and depth information within the current visual field. The rotation amplitude is quantified by the change in head rotation angle. S200. Centered on the depth of the reference object on the work surface, the near-end boundary is defined by the near-end offset, and the far-end boundary is defined by the far-end offset. The depth range between the near-end boundary and the far-end boundary constitutes the effective depth range. The reference object on the work surface is a calibrated object with a relatively fixed position in the work scene. The effective depth range is the depth range that is effectively focused on during the visual analysis process. The recording accuracy is reflected in the completeness of object feature extraction and the accuracy of position positioning. When S300 detects visual field regression, it sequentially determines the global motion of the image, gaze point drift, and overlap with the reference object. When all of these meet the corresponding stability threshold, it initiates new frame analysis, spatially matches the new frame with the visual field semantic snapshot, and verifies the number of consecutive frames of disappearance confirmation to check for undetected targets, marking them as disappearance candidates. The global motion of the image reflects the overall displacement of the acquired image, the gaze point drift reflects the shift of the visual attention center, and the overlap with the reference object reflects the positioning accuracy after visual field regression. Spatial matching achieves consistency comparison based on coordinate position and feature information. S400: Obtain a list of objects marked as disappearance candidates, exclude operation occlusion and depth-of-field offset, analyze the disappearance event and assign confidence, and output visual cues; confidence is used to characterize the credibility of the disappearance event judgment result, operation occlusion is the visual occlusion of the target object by the human body or tool during the operation, and depth-of-field offset is the visual undetected state formed by the target object leaving the effective depth-of-field range. S500 determines the authenticity of prompts based on user responses, and corrects near-end offset, far-end offset, number of consecutive frames for disappearance confirmation, and stability threshold. User responses include confirmation and rejection of prompt results. The correction process dynamically adjusts the corresponding parameter values to make the subsequent visual analysis process more in line with actual operating scenarios.
[0018] Specifically, step S100 includes: S110. Obtain the peak angular velocity, duration, and static duration during head rotation; the peak angular velocity is the maximum value of the angular velocity during head rotation, the duration of rotation is the duration during which the angular velocity exceeds the reference value, and the static duration after rotation is the duration during which the image remains stable after the rotation ends. Specifically, in this embodiment, the wearable acquisition device collects the head angular velocity at 100Hz; a 0.5-second sliding window is set, all angular velocity sampling points are traversed, the maximum angular velocity within each window is calculated, and the maximum angular velocity within all windows is compared to obtain a peak angular velocity of 65° / s; with 3° / s as a reference, timing starts when the real-time angular velocity exceeds the reference value and stops when the angular velocity falls back below the reference value, resulting in a rotation duration of 0.45s; after the rotation ends, frame images are continuously acquired, the average global pixel grayscale difference between adjacent frames is calculated, and if the difference is lower than a stable threshold for 3 consecutive frames, the device is determined to be stationary, resulting in a stationary duration of 0.3s after rotation; this is only an example and is not a limitation. S120. Compare the peak angular velocity with the first preset threshold, compare the rotation duration with the second preset threshold, and compare the stationary duration after rotation with the third preset threshold. Specifically, in this embodiment, the first preset threshold is 50° / s; the second preset threshold is 0.2s; and the third preset threshold is 0.15s. The comparison results are: 65° / s > 50° / s, 0.45s > 0.2s, and 0.3s > 0.15s, all of which are satisfied. This is only an example and is not intended to be limiting. S130. When the peak angular velocity exceeds the first preset threshold, the rotation duration exceeds the second preset threshold, and the static duration after rotation exceeds the third preset threshold, a freeze command is generated to freeze the semantic snapshot of the field of view. The value of the third preset threshold is correlated with the value of the peak angular velocity in the same direction. When the peak angular velocity changes in the positive direction, the third preset threshold is adjusted in the positive direction synchronously, and when the peak angular velocity changes in the negative direction, the third preset threshold is adjusted in the negative direction synchronously. The freeze command is used to control the visual acquisition device to stop updating the field of view feature information and fix the semantic data at the current moment. Specifically, in this embodiment, a freeze command is generated to fix the object category, position, and depth information of the current image, forming a semantic snapshot of the field of view; the third preset threshold is linearly correlated with the peak angular velocity in the same direction, using the formula: T3=k×ω max+b; where T3 represents the third preset threshold; k represents the proportional coefficient, which is 0.002 in this embodiment; ω max 'b' represents the peak angular velocity; 'b' represents the base offset, which is 0.02 in this embodiment; peak angular velocity = 65° / s, dynamically adjusting the third preset threshold to obtain T3 = 0.002 × 65 + 0.02 = 0.15s; when the angular velocity increases, T3 increases synchronously; when the angular velocity decreases, T3 decreases synchronously; this is only an example and is not a limitation. S140. Synchronously record the amplitude of head rotation during the rotation process; Specifically, the total rotation angle is calculated using a three-axis angle synthesis algorithm, with the formula: θ = (Δθ) x 2 +Δθ y 2 +Δθ z 2 ) 1 / 2 Where θ represents the rotation amplitude, Δθ x Δθ represents the change in the head's angle around the X-axis. y Δθ represents the change in the head's angle around the Y-axis. z This represents the change in the head's angle around the Z-axis. In this embodiment, Δθ x 2 =8°, Δθ y 2 =20°, Δθ z 2 =3°, the calculated rotation amplitude θ≈21.75°, and this value is saved.
[0019] Specifically, step S200 includes: S210. Identify the reference object of the work surface in the current field of view, and obtain the visual depth value of the reference object of the work surface as the depth center; the visual depth value is used to characterize the distance information between the reference object of the work surface and the vision acquisition device; Specifically, to identify a fixed reference object in the field of view, this embodiment uses a monocular depth estimation algorithm to extract the image features of the reference object on the work surface in the image. Based on the depth mapping table, the pixel grayscale values are converted into depth values. The average depth of the reference object area is calculated, and a depth value of 2.2m is obtained, which is used as the depth center. S220. Using the depth center as a reference, the near-end boundary is defined using the near-end offset, and the far-end boundary is defined using the far-end offset. The near-end offset is adaptively adjusted based on the rotation amplitude, and the far-end offset is determined by the current object size attribute, which is a pre-configured process parameter. The depth range between the near-end boundary and the far-end boundary constitutes the effective depth range. Specifically, the proximal offset is adaptively calculated with respect to the rotation amplitude, and the formula is: d near=α×θ; where, d near denoted by α; denoted by α; denoted by 0.01 m / ° in this embodiment; θ; denoted by θ; and calculated as d. near =0.01×21.75=0.2175m; the near edge boundary is 2.2-0.2175=1.9825m; the far edge offset is 0.35m based on the object size, and the far edge boundary is 2.2+0.35=2.55m; therefore, the effective depth of field range is 1.9825m~2.55m; this is only an example and is not a limitation. S230. For objects located within the effective depth-of-field range, record them in a hierarchical manner according to their size. Specifically: objects within the first size range are recorded with two-dimensional coordinates and contour feature points; objects within the second size range are recorded with two-dimensional coordinates and category labels; and objects within the third size range are recorded with existence markers. The first size range is smaller than the second size range, and the second size range is smaller than the third size range. Two-dimensional coordinates represent the object's planar position in the captured image; contour feature points are key positioning points of the object's outline; category labels are identification information of the object's type; and existence markers are indicators of whether the object exists. The size classification threshold used for hierarchical recording is correlated in the same direction as the width of the effective depth-of-field range. When the range width changes positively, the upper limit of the first size range is adjusted positively in sync. The size classification threshold is used to distinguish objects of different size levels, and changes in the range width directly alter the criteria for determining object size classification. Specifically, the depth of field width is approximately 0.57m, and the size threshold is correlated with the width in the same direction; the first size is <0.12m, which records coordinates and contour feature points; the second size is 0.12m-0.4m, which records coordinates and category labels; the third size is >0.4m, which only records the existence status; as the width of the depth of field range increases, the upper limit of the first size increases synchronously; this is only an example and is not a limitation.
[0020] Specifically, step S300 includes: S310. Acquire the image data after the field of view regression, and determine the global motion state of the image, the gaze point drift state and the overlap between the gaze point and the reference object on the work surface in sequence. The field of view regression is the state in which the visual acquisition device returns to the target work field of view after the head rotation ends. The global motion state of the image is determined by the inter-frame pixel displacement amount, and the gaze point drift state is determined by the displacement deviation amount of the attention center. Specifically, in this embodiment, an inter-frame pixel displacement algorithm is used to select uniform feature points in the image, calculate the average pixel displacement of corresponding feature points in adjacent frames, and obtain a global motion value of 1.2 pixels; a focus center centroid algorithm is used to extract feature points in the focus area of the image; the focus center is obtained by calculating the average coordinates of the feature points; the focus center is obtained by calculating the offset between the focus center and the initial center, and the gaze point drift is obtained as 2.5 pixels; the reference object overlap is obtained as 4.5 pixels based on the pixel distance between the current gaze point and the center of the reference object; this is only an example and is not a limitation. S320, Preset a first stability judgment threshold, a second stability judgment threshold, and a third stability judgment threshold; the values of the first stability judgment threshold and the second stability judgment threshold are inversely related to the rotation amplitude value; the value of the third stability judgment threshold is in the same direction as the rotation amplitude value. Specifically, in this embodiment, the rotation amplitude θ = 21.75°; The first stability threshold S1 is inversely related to the rotation amplitude. That is, the larger the rotation amplitude, the more violent the head movement, and the higher the probability of slight image jitter after regression. Therefore, stricter judgment conditions are needed for the overall image displacement, meaning the allowable pixel displacement should be smaller. The calculation formula is: S1 = a1 - b1 × θ; where a1 represents the global motion baseline threshold, which is taken as 5 pixels in this embodiment; b1 represents the proportional coefficient of the global motion threshold changing with the rotation amplitude, which is taken as 0.05 pixels / degree in this embodiment; θ represents the head rotation amplitude; thus, S1 = 5 - 0.05 × 21.75 ≈ 3.91 pixels; the judgment condition is that the global motion pixel displacement ≤ S1. The second stability threshold S2 is inversely related to the rotation amplitude. That is, the larger the rotation amplitude, the more difficult it is for the gaze point to re-lock onto the target after regression. Therefore, the allowable gaze center offset should be smaller to force the gaze to fully return to center. The calculation formula is: S2 = a2 - b2 × θ; where a2 represents the gaze point drift baseline threshold, which is taken as a2 = 8 pixels in this embodiment; b2 represents the proportional coefficient of the gaze point drift threshold changing with the rotation amplitude, which is taken as b2 = 0.1 pixels / degree in this embodiment; θ represents the magnitude of the head rotation amplitude; thus, S2 = 8 - 0.1 × 21.75 ≈ 5.82 pixels; the judgment condition is that the gaze point drift pixel deviation ≤ S2. The third stability threshold S3 is correlated with the rotation amplitude. That is, the larger the rotation amplitude, the more likely the aiming and positioning of the reference object will deviate angularly after the head turns. If extremely high coincidence accuracy is still required at this point, the system will struggle to start analysis. Therefore, the requirement for coincidence accuracy should be relaxed, meaning the maximum allowable distance between the gaze point and the center of the reference object should be greater. Here, the pixel distance between the gaze point and the center of the reference object is used as the quantitative indicator of coincidence. The calculation formula is: S3 = c3 × θ; where c3 represents the proportional coefficient of the coincidence threshold as a function of the rotation amplitude, and in this embodiment, it is taken as 0.2 pixels / degree; thus, S3 = 0.2 × 21.75 ≈ 4.35 pixels. The judgment condition is: the pixel distance between the gaze point and the center of the reference object ≤ S3. The larger the rotation amplitude θ, the smaller S1 and S2, and the larger S3; in this embodiment, the threshold values are S1≈3.91 pixels, S2≈5.82 pixels, and S3≈4.35 pixels; this is only for illustrative purposes and is not a limitation. S330. When the global motion state of the screen meets the first stability judgment threshold, the gaze point drift state meets the second stability judgment threshold, and the overlap between the gaze point and the reference object on the work surface meets the third stability judgment threshold, start the semantic analysis of the new frame. Specifically, 1.2 pixels < 3.91 pixels, 2.5 pixels < 5.82 pixels, and 4.5 pixels > 4.35 pixels do not meet the start-up conditions; when all three conditions are met, semantic analysis of the new frame is started. S340. Perform spatial dimension matching between the semantic analysis results of the new frame and the visual semantic snapshot, and retrieve objects in the snapshot that were not detected in the new frame. Spatial dimension matching is achieved by combining planar coordinates and depth information. Objects in the snapshot that were not detected are targets that exist in the snapshot but are not identified in the new frame. Specifically, extract the two-dimensional coordinates of objects in the semantic analysis results of the new frame and the semantic snapshot of the field of view, compare the corresponding object depth values, and compare the object category and contour features. When the two-dimensional coordinate deviation, depth difference and category or contour feature similarity are all within the preset matching tolerance range, they are determined to be the same object; otherwise, they are recorded as undetected targets. S350. Perform continuous frame verification on undetected objects. Only if the object is not detected in any new frame of the consecutive disappearance confirmation frame number is the object marked as a disappearance candidate. The value of the consecutive disappearance confirmation frame number is inversely related to the field of view stability. Specifically, the number of consecutive frames required to confirm the disappearance is calculated using the formula: N = ⌈12 / S stable ⌉; where N represents the number of consecutive frames of disappearance confirmation; S stableThis represents the overall stability of the field of view, which is calculated by weighting the global motion of the image and the gaze point drift. The value range is usually 1-10, with higher stability resulting in a larger value. ⌈⌉ represents the rounding up sign. In this embodiment, the current overall stability of the field of view is rated as 6, and substituting it into the formula, we get N=⌈12 / 6⌉=2 frames. If the target is not detected for 2 consecutive frames, it is marked as a candidate for disappearance. This is only an example and is not a limitation.
[0021] Specifically, step S400 includes: S410. Obtain a list of objects marked as disappearance candidates, and perform operation occlusion exclusion judgment for each disappearance candidate object. The operation occlusion exclusion judgment includes determining whether the degree of overlap between the expected coordinate position of the disappearance candidate object in the snapshot and the bounding box of the large object identified in the semantic analysis result of the new frame meets the coverage judgment tolerance. If it meets the tolerance and the category of the covered object is a tool or a hand, then the disappearance candidate object is removed from the list and judged as operation occlusion. The bounding box is the area surrounded by the shape of the identified object in the new frame. The coverage judgment tolerance is the allowable range of positional deviation of the occlusion relationship. The delineation accuracy is reflected in the accuracy of the boundary definition of the depth range. Specifically, the Intersection over Union (IoU) calculation formula is: IoU = S overlap / S union Where IoU represents the intersection-union ratio of the bounding boxes; S overlap S represents the overlap area between the candidate object and the bounding box of the occluder; union This represents the total area of the bounding box between the candidate object and the occluder; the coverage tolerance is set to 0.7. In this embodiment, IoU=0.2<0.7, which is determined to be no tool or hand occlusion and not operation occlusion. S420. For disappearing candidate objects that remain in the list after occlusion exclusion, perform out-of-field offset determination. Out-of-field offset determination includes retrieving object features in the semantic analysis results of the new frame that are located in the vicinity of the effective depth range boundary. If there are contour features or category labels that match the disappearing candidate object, the disappearing candidate object is removed from the list and determined to be out-of-field offset. The vicinity region is the extended detection region outside the effective depth range boundary, used to determine whether the object deviates only from the effective depth range. Specifically, contour feature matching includes extracting object contours in a 0.2m neighborhood outside the depth of field boundary, calculating the similarity of the contour features of candidate objects and objects within the region, and determining that the similarity is higher than a threshold; in this embodiment, if no matching features are found, it is not determined to be an out-of-depth offset; this is only an example and is not a limitation. S430. After occlusion exclusion and depth-of-field offset determination, the disappearing candidate objects that remain in the list are identified as disappearing events and given a confidence level, and a visual cue is output. The confidence level and the size range of the disappearing candidate objects are used to determine the intensity and direction of the visual cue. The visual cue presents the disappearing event information to the user through a visual identifier, and the cue points to the expected position of the corresponding disappearing object. The intensity of the cue is distinguished by the way the identifier is displayed. Specifically, the confidence level is jointly determined by the basic confidence level of the size range and the stability factor for the disappearance event determination. In this embodiment, the basic confidence level of the first size range is 0.9, the second size is 0.7, and the third size is 0.5. The stability factor is adjusted according to the detection fluctuation of the disappearance candidate object in consecutive verification frames. Here, the stability factor is 1.0, so the final confidence level is the basic confidence level of the size. In this embodiment, the object is the second size with a confidence level of 0.7. A medium-brightness flashing indicator is used to output a visual cue. This is only an example and is not a limitation.
[0022] Specifically, step S500 includes: S510. Obtain the user's response to the visual cues, and determine whether the result of the disappearance event is true based on the response behavior; Specifically, in this embodiment, the user performs a negative operation, marking the current disappearance judgment as a misjudgment; S520. When the disappearance event is not determined to be true and the disappearance candidate object has matching features in the area near the boundary of the effective depth range in the depth-of-field offset determination, the near-end offset and far-end offset are corrected and the coverage of the effective depth range is adjusted. Specifically, the depth-of-field deviation factor is defined as the ratio of the distance by which the actual depth of an object exceeds the boundary to the width of the depth-of-field interval. In this embodiment, the depth-of-field deviation factor is calculated to be 0.3. The correction magnitude is correlated with the degree of deviation and a linear correction formula is used: ΔD=k×ξ×L0; where ΔD represents the offset correction increment; k represents the correction ratio coefficient, which is taken as k=1.0 in this embodiment; ξ represents the depth-of-field deviation factor, characterizing the degree to which the object deviates from the depth-of-field interval; L0 represents the baseline adjustment step size, which is taken as L0=0.1m in this embodiment; the calculated correction increment ΔD=1.0×0.3×0.1=0.03m increases the near-end offset and far-end offset by 0.03m each, thereby expanding the effective depth-of-field interval. S530. When an object within the first size range is determined to be a disappearance event but has not actually disappeared, the number of consecutive frames for disappearance confirmation is corrected. Specifically, the object was incorrectly identified as disappearing despite being of the first size. The number of consecutive frames for confirming disappearance was increased from 2 to 3 to improve the stability of the judgment. S540. When the semantic analysis of a new frame is not started normally after the view regression, the first stability judgment threshold, the second stability judgment threshold and the third stability judgment threshold are corrected and the stability judgment conditions are adjusted. Specifically, if the analysis is not initiated due to a strict threshold, reduce the values of S1 and S2, increase the value of S3, and relax the stability judgment conditions. S550, The magnitude of the correction is correlated in the same direction as the degree of deviation of the judgment result; Specifically, the magnitude of all parameter corrections changes in the same direction as the degree of deviation in the judgment. The greater the degree of deviation, the greater the parameter adjustment magnitude, so that the system can adapt to the actual operation scenario.
[0023] This invention provides another technical solution: a visual data analysis system for work scenarios, which includes: a snapshot freezing module, a depth-of-field segmentation module, a stability determination module, an occlusion elimination module, and a feedback correction module. The snapshot freeze module monitors the peak value, duration, and static duration of head rotation angular velocity. When all three meet the corresponding thresholds, a semantic snapshot of the visual field is frozen, and the rotation amplitude is recorded simultaneously. The depth-of-field segmentation module delineates the effective depth-of-field range centered on the depth of the reference object on the work surface and records object information within the range according to object size. The stability determination module determines that after the global motion of the screen, gaze point drift, and overlap with the reference object all meet the corresponding stability determination thresholds, a new frame analysis is initiated. The new frame is matched with the snapshot space, and the number of consecutive frames of disappearance confirmation is used to verify the detection of undetected targets and mark them as disappearance candidates. The number of consecutive frames of disappearance confirmation decreases adaptively with the stability of the visual field. The occlusion exclusion module obtains a list of objects marked as disappearance candidates, excludes operation occlusion and out-of-field offset, confirms the disappearance event, and assigns a confidence level. The feedback correction module determines the authenticity of the prompt based on the user's response and corrects the near-end offset, far-end offset, number of consecutive frames of disappearance confirmation, and stability determination threshold accordingly.
[0024] Specifically, the depth-of-field segmentation module includes an offset adjustment unit and a hierarchical recording unit. The offset adjustment unit is used to define the near boundary using the near offset and the far boundary using the far offset, with the depth range between the near and far boundaries constituting the effective depth-of-field interval. The hierarchical recording unit is used to record objects located within the effective depth-of-field interval according to their size. Objects within the first size range are recorded with two-dimensional coordinates and contour feature points, objects within the second size range are recorded with two-dimensional coordinates and category labels, and objects within the third size range are recorded with existence markers.
[0025] Specifically, the stability determination module includes a threshold adjustment unit and a candidate marking unit. The threshold adjustment unit is used to preset a first stability determination threshold, a second stability determination threshold, and a third stability determination threshold. The first stability determination threshold is used to determine the global motion state of the screen, the second stability determination threshold is used to determine the gaze point drift state, and the third stability determination threshold is used to determine the overlap. The candidate marking unit is used to start new frame semantic analysis when the global motion state, gaze point drift state, and overlap all meet the corresponding thresholds. The new frame semantic analysis results are matched with the visual field semantic snapshot space, and continuous frame verification is performed on undetected objects. Only when the object is not detected in any new frame of the consecutive disappearance confirmation frame number is it marked as a disappearance candidate.
[0026] Specifically, the occlusion exclusion module includes an occlusion determination unit and a prompt confirmation unit. The occlusion determination unit performs operation occlusion exclusion determination on each candidate object that has disappeared. It determines whether the overlap between the expected coordinate position of the object in the snapshot and the bounding box of the large object identified in the semantic analysis result of the new frame meets the coverage determination tolerance. If it meets the tolerance and the covered object category is a tool or a hand, it is removed from the list and determined to be an operation occlusion. The unit also performs out-of-field offset determination on the objects that remain after operation occlusion exclusion. It searches for features in the area near the boundary of the effective depth range in the new frame. If there are matching contour features or category labels, the objects are removed from the list and determined to be out-of-field offsets. The prompt confirmation unit confirms the objects that remain after double exclusion as disappearance events and assigns a confidence level.
[0027] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A method for visual data analysis of a work scenario, characterized in that: The method includes: S100: Monitor the peak angular velocity, duration, and rest duration after head rotation. When the peak angular velocity exceeds a first preset threshold, the duration of rotation exceeds a second preset threshold, and the rest duration after rotation exceeds a third preset threshold, generate a freeze command to freeze the visual semantic snapshot and record the rotation amplitude simultaneously. S200. Using the depth of the reference object on the working surface as the center, the near-end boundary is defined by the near-end offset, and the far-end boundary is defined by the far-end offset. The depth range between the near-end boundary and the far-end boundary constitutes the effective depth range. S300 When visual field regression is detected, the global motion of the screen, gaze point drift and overlap with the reference object are determined in sequence. When all of them meet the corresponding stability judgment threshold, a new frame analysis is started. The new frame is matched with the visual field semantic snapshot space. The target that is not detected is checked based on the number of consecutive frames of disappearance confirmation and marked as a disappearance candidate. S400: Obtain a list of objects marked as disappearance candidates, exclude operation occlusion and depth-of-field offset, analyze disappearance events and assign confidence scores, and output visual cues; S500 determines the authenticity of prompts based on user responses and corrects near-end offset, far-end offset, number of consecutive frames for disappearance confirmation, and stability judgment threshold.
2. The method for visual data analysis of a work scene according to claim 1, characterized in that: Step S100 includes: S110. Obtain the peak angular velocity, duration, and duration of stillness during head rotation. S120. Compare the peak angular velocity with a first preset threshold, compare the rotation duration with a second preset threshold, and compare the stationary duration after rotation with a third preset threshold. S130. When the peak angular velocity exceeds the first preset threshold, the duration of rotation exceeds the second preset threshold, and the duration of stillness after rotation exceeds the third preset threshold, a freeze command is generated to freeze the visual semantic snapshot; the value of the third preset threshold is correlated in the same direction as the value of the peak angular velocity. S140. Simultaneously record the rotation amplitude value during the head rotation process.
3. The method for visual data analysis of a work scene according to claim 1, characterized in that: Step S200 includes: S210. Identify the reference object of the work surface in the current field of view, and obtain the visual depth value of the reference object of the work surface as the depth center; S220. Using the depth center as a reference, the near-end boundary is defined using the near-end offset, and the far-end boundary is defined using the far-end offset. The near-end offset is adaptively adjusted based on the rotation amplitude, and the far-end offset is determined by the current object size attribute, which is a pre-configured process parameter. The depth range between the near-end boundary and the far-end boundary constitutes the effective depth range. S230. For objects located within the effective depth range, record them in a hierarchical manner according to their size, wherein: objects within the first size range are recorded with two-dimensional coordinates and contour feature points, objects within the second size range are recorded with two-dimensional coordinates and category labels, and objects within the third size range are recorded with existence markers. The first size range is smaller than the second size range, and the second size range is smaller than the third size range. The size division threshold used for the hierarchical recording is correlated in the same direction as the width of the effective depth range.
4. The method for visual data analysis of a work scene according to claim 1, characterized in that: Step S300 includes: S310. Obtain the image data after the field of view regression, and determine the global motion state of the image, the gaze point drift state, and the overlap between the gaze point and the reference object on the work surface in sequence. S320, Preset a first stability determination threshold, a second stability determination threshold, and a third stability determination threshold; the values of the first stability determination threshold and the second stability determination threshold are inversely related to the rotation amplitude value; the value of the third stability determination threshold is in the same direction as the rotation amplitude value; S330. When the global motion state of the screen meets the first stability judgment threshold, the gaze point drift state meets the second stability judgment threshold, and the overlap between the gaze point and the reference object on the work surface meets the third stability judgment threshold, start the new frame semantic analysis. S340. Perform spatial dimension matching between the semantic analysis results of the new frame and the visual semantic snapshot to retrieve objects in the snapshot that were not detected in the new frame. S350. Perform continuous frame verification on undetected objects. Only if the object is not detected in any new frame of the consecutive number of consecutive disappearance confirmation frames, mark the object as a disappearance candidate. The value of the consecutive number of disappearance confirmation frames is inversely related to the field of view stability.
5. The method for visual data analysis of a work scene according to claim 1, characterized in that: Step S400 includes: S410. Obtain a list of objects marked as disappearance candidates, and perform an operation occlusion exclusion judgment for each disappearance candidate object; the operation occlusion exclusion judgment includes determining whether the degree of overlap between the expected coordinate position of the disappearance candidate object in the snapshot and the bounding box of the large object identified in the semantic analysis result of the new frame meets the coverage judgment tolerance. If it meets the tolerance and the category of the covered object is a tool or a hand, then remove the disappearance candidate object from the list and determine it as operation occlusion. S420. For the missing candidate objects that remain in the list after the occlusion exclusion judgment, perform the depth-of-field offset judgment; the depth-of-field offset judgment includes searching for object features in the region near the boundary of the effective depth-of-field interval in the semantic analysis results of the new frame. If there are contour features or category labels that match the missing candidate objects, the missing candidate objects are removed from the list and judged as depth-of-field offset. S430. After the occlusion exclusion judgment and the depth-of-field offset judgment, the disappearing candidate objects that are still retained in the list are confirmed as disappearance events and assigned a confidence level, and a visual cue is output; the confidence level and the size range recorded by the disappearing candidate objects jointly determine the intensity and direction of the visual cue.
6. The method for visual data analysis of a work scene according to claim 1, characterized in that: Step S500 includes: S510. Obtain the user's response to the visual cues, and determine whether the result of the disappearance event is true based on the response behavior; S520. When the disappearance event is not determined to be true and the disappearance candidate object has matching features in the area near the boundary of the effective depth range in the depth-of-field offset determination, the near-end offset and the far-end offset are corrected to adjust the coverage of the effective depth range. S530. When an object within the first size range is determined to be a disappearance event but has not actually disappeared, the number of consecutive frames for disappearance confirmation is corrected. S540. When the semantic analysis of a new frame is not started normally after the view regression, the first stability judgment threshold, the second stability judgment threshold and the third stability judgment threshold are corrected and the stability judgment conditions are adjusted. S550, The magnitude of the correction is correlated in the same direction as the degree of deviation of the judgment result.
7. A visual data analysis system for a work scene, applied to the visual data analysis method for a work scene as described in any one of claims 1-6, characterized in that: The system includes: a snapshot freeze module, a depth-of-field segmentation module, a stabilization determination module, an occlusion removal module, and a feedback correction module; The snapshot freezing module is used to monitor the peak value of the head rotation angular velocity, the duration of the rotation, and the duration of stillness after rotation. When all three values meet the corresponding thresholds, the visual semantic snapshot is frozen, and the rotation amplitude is recorded simultaneously. The depth-of-field segmentation module is used to define the effective depth-of-field range with the reference object depth on the work surface as the center, and to record object information within the range according to object size. The stability determination module is used to sequentially determine that the global motion of the screen, the gaze point drift, and the overlap with the reference object all meet the corresponding stability determination thresholds before starting a new frame analysis. The new frame is matched with the snapshot space, and the undetected target is marked as a disappearance candidate based on the number of consecutive frames of disappearance confirmation. The number of consecutive frames of disappearance confirmation decreases adaptively with the stability of the field of view. The occlusion exclusion module is used to obtain a list of objects marked as disappearance candidates, exclude operation occlusion and out-of-field offset, confirm the disappearance event and assign confidence level; The feedback correction module is used to determine the authenticity of the prompt based on the user's response, and to correct the near-end offset, far-end offset, number of consecutive frames of disappearance confirmation, and stability judgment threshold accordingly.
8. The visual data analysis system for a work scene according to claim 7, characterized in that: The depth-of-field segmentation module includes an offset adjustment unit and a hierarchical recording unit; The offset adjustment unit is used to define the near boundary using the near offset and the far boundary using the far offset, with the depth center as the reference. The depth range between the near boundary and the far boundary constitutes the effective depth range. The hierarchical recording unit is used to record objects located within the effective depth range according to their size. The objects within the first size range are recorded with two-dimensional coordinates and contour feature points, the objects within the second size range are recorded with two-dimensional coordinates and category labels, and the objects within the third size range are recorded with existence markers.
9. A visual data analysis system for a work scene according to claim 7, characterized in that: The stability determination module includes a threshold adjustment unit and a candidate labeling unit; The threshold adjustment unit is used to preset a first stability determination threshold, a second stability determination threshold, and a third stability determination threshold; the first stability determination threshold is used to determine the global motion state of the screen, the second stability determination threshold is used to determine the gaze point drift state, and the third stability determination threshold is used to determine the overlap. The candidate marking unit is used to start new frame semantic analysis when the global motion state of the screen, the gaze point drift state and the overlap degree all meet the corresponding thresholds, match the new frame semantic analysis results with the visual semantic snapshot space, and perform continuous frame verification on undetected objects. Only when the object is not detected in any new frame of the consecutive disappearance confirmation frame number is it marked as a disappearance candidate.
10. A visual data analysis system for a work scene according to claim 7, characterized in that: The occlusion removal module includes an occlusion determination unit and a prompt confirmation unit; The occlusion determination unit is used to perform operation occlusion exclusion determination for each disappearing candidate object, and to determine whether the degree of overlap between its expected coordinate position in the snapshot and the bounding box of the large object identified in the semantic analysis result of the new frame meets the coverage determination tolerance. If it meets the tolerance and the covered object category is a tool or a hand, it is removed from the list and determined to be operation occlusion. The unit also performs depth-of-field offset determination on the objects retained after operation occlusion exclusion, and searches for features in the area near the boundary of the effective depth-of-field range in the new frame. If there are matching contour features or category labels, they are removed from the list and determined to be depth-of-field offset. The prompt determination unit is used to identify objects that remain after double exclusion as disappearance events and assign them a confidence level.