Deviation early warning and positioning system using machine vision
By deeply coupling multimodal visual perception with source-tracing localization, the problem of deviation warning and localization in complex scenarios of machine vision technology is solved, realizing timely warning, accurate localization and effective source tracing of deviations, and improving the robustness and adaptability of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHUHAI BOMING VISION TECH CO LTD
- Filing Date
- 2026-01-10
- Publication Date
- 2026-04-24
AI Technical Summary
Existing machine vision technologies lack deviation warning capabilities in complex scenarios, cannot trace the propagation path and physical causes of deviations, and have not built a closed-loop optimization system, making it difficult to meet the needs for early warning, accurate positioning, traceability, and continuous optimization of deviations.
By employing a multimodal visual perception module, a semantic temporal integration module, a dynamic deviation early warning module, a source tracing and positioning module, and a feedback and optimization module, deviation quantification, trend prediction, cause tracing, and system optimization are achieved through the fusion of RGB images, infrared images, and temporal continuous frame data.
It achieves timely early warning of deviations, accurate positioning, and effective source tracing, improves the system's robustness and decision support capabilities in complex scenarios, reduces false alarm and missed alarm rates and rectification and investigation costs, and provides a more intelligent adaptive system operation mode.
Smart Images

Figure CN121921368A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine vision technology, and more specifically discloses a deviation warning and positioning system utilizing machine vision. Background Technology
[0002] With the rapid development of industrial automation and intelligent manufacturing technologies, production systems are increasingly demanding higher precision, efficiency, and reliability. Traditional manual inspection and single-point monitoring methods are no longer sufficient to meet the needs of real-time perception and precise control of minute deviations in complex scenarios. Against this backdrop, machine vision technology has emerged. It simulates human visual functions and uses image sensors and computer algorithms to identify, measure, and locate targets.
[0003] The prior art patent document with authorization announcement number CN107818587B discloses "a high-precision positioning method for machine vision based on ROS", which includes two high-precision industrial cameras and an image processing system connected to the high-precision industrial cameras. The image processing system includes a positioning chip, a development board running ROS, and a parallel computing unit. The positioning method involves driving the high-precision industrial cameras to align with the positioning platform to acquire images, calculating the coordinate positions 1 and 2 of the target signal respectively; integrating the final images to calculate the coordinate information 3 of the target signal and calculating the deviation of coordinate information 3 from coordinate positions 1 and 2; finally, the parallel computing unit comprehensively processes coordinate positions 1, 2, and 3 to precisely adjust the target information and obtain the final coordinate position.
[0004] The patent document with authorization announcement number CN118518005B discloses "a machine vision positioning method, device and system". The method first acquires an image of the product, then performs shape analysis on the product in the image, finds the geometric center of the inner and outer contours by recognizing and analyzing the inner and outer contours, and then determines the center point of the entire product based on the two geometric centers. In this application, the deviation between the center point and different positions of the inner and outer contours is reduced, the processing allowance is balanced, and the processing amount is more even on both the inner and outer sides of the product, thereby improving product quality.
[0005] While existing technologies can improve positioning accuracy through multi-coordinate comparison or contour analysis, and utilize dual cameras to acquire images and calculate multiple coordinates for comparison to eliminate errors, thereby improving the robot's positioning accuracy in areas with high feature similarity, and by identifying the geometric center of the product's inner and outer contours and combining it with area ratios to determine the center point, positioning deviations are reduced, processing allowances are balanced to improve product quality, and automatic positioning of the product and equipment center is achieved, reducing labor costs, these technologies generally rely on single-type image data or fixed positioning logic. They are weakly resistant to environmental interference such as changes in lighting and occlusion, do not integrate multimodal data and temporal dynamic information, lack deviation warning functions, can only achieve simple positioning output, cannot trace the propagation path and physical cause of deviations, and lack a mechanism for dynamically adapting to changes in the scene. They have not built a closed-loop optimization system and cannot adjust data acquisition, feature processing, or positioning logic based on feedback from actual applications, making it difficult to meet the needs of early deviation warning, accurate positioning, traceability, and continuous optimization in complex scenarios. Summary of the Invention
[0006] The main technical problem solved by this invention is to provide a deviation warning and positioning system using machine vision, which can solve the problems mentioned in the background art.
[0007] To address the aforementioned technical problems, according to one aspect of the present invention, more specifically, a deviation warning and localization system utilizing machine vision, comprising: a multimodal visual perception module, a semantic temporal integration module, a dynamic deviation warning module, a source tracing and localization module, and a feedback and optimization module. The multimodal visual perception module acquires RGB images, infrared images, and temporally continuous frame data, and extracts scene semantic features using a semantic segmentation model. The semantic temporal integration module, based on optical flow and semantic similarity calculation, achieves deep fusion of motion features and semantic features, generating a temporal semantic fusion feature map. The dynamic deviation warning module, using a dynamically updated reference frame, triggers tiered warnings and associates semantic information through deviation quantization, trend prediction, and dynamic threshold determination. The source tracing and localization module analyzes the deviation propagation path and physical causes, simultaneously outputting the three-dimensional coordinates of the deviation starting point and the cause determination result. The feedback and optimization module feeds back the source tracing and localization results, dynamically optimizing the reference frame, fusion weights, and deviation propagation model.
[0008] Furthermore, the multimodal visual perception module includes: a multi-source data acquisition module and a feature extraction module;
[0009] Multi-source data acquisition module: Simultaneously acquires scene RGB images, infrared images, and time-series continuous frame data through RGB cameras and infrared thermal imaging cameras;
[0010] Feature extraction module: Using the U-Net semantic segmentation model, key object categories, core component locations and component relationships are extracted from the collected multimodal data to generate a scene semantic feature set.
[0011] Furthermore, the semantic temporal integration module includes: a motion feature calculation module, a semantic association module, and a temporal semantic fusion module;
[0012] Motion feature calculation module: The Lucas-Kanade optical flow method is used to calculate the motion vectors between consecutive temporal frames, capture the dynamic motion trajectory of objects, and filter effective motion information;
[0013] Semantic association module: The cosine similarity algorithm is used to calculate the semantic feature matching degree between the current frame and the reference frame, and to quantify the consistency of semantic information;
[0014] Temporal semantic fusion module: Based on the consistency of motion vectors, abnormal motion points are removed, and weights are allocated by combining semantic matching degree. The motion features and semantic features are deeply fused to generate a temporal semantic fusion feature map.
[0015] Furthermore, the dynamic deviation early warning module includes: a deviation quantification module, a trend prediction module, a dynamic threshold generation module, and a graded early warning module;
[0016] Deviation quantization module: Using a dynamically updated reference frame as a reference, and based on the temporal semantic fusion feature map, calculates the feature difference between the current frame and the reference frame to achieve deviation quantization representation;
[0017] Trend prediction module: It uses the exponential smoothing method to analyze the time-series propagation law of deviation, predicts the evolution trend of deviation in the next frame, and avoids false early warnings triggered by instantaneous deviation;
[0018] Dynamic threshold generation module: Based on the statistical distribution of historical deviation data, the threshold parameters are adaptively adjusted according to the 3σ principle to generate dynamic early warning thresholds that match scene changes in real time;
[0019] Tiered early warning module: When the deviation quantification value exceeds the dynamic threshold and the trend prediction shows that the deviation is expanding, a three-level early warning is triggered: mild, moderate and severe. The early warning information is synchronously associated with the corresponding semantic feature set.
[0020] Furthermore, the source tracing and positioning module includes: a deviation tracing module, a semantic analysis module, and a three-dimensional coordinate positioning module;
[0021] Deviation tracing module: Based on temporal semantic fusion feature map, reverse analysis is performed to analyze the propagation path of deviation from the initial occurrence area to the diffusion area, and the starting point of deviation is located.
[0022] Semantic analysis module: Relates semantic categories and component relationships in the semantic feature set of the associated scene to determine the physical causes of deviations;
[0023] 3D coordinate positioning module: By using the depth complementarity of RGB and infrared images and employing multimodal feature matching technology, the 3D spatial coordinates of the deviation starting point are calculated.
[0024] Furthermore, the feedback and optimization module includes: a feedback acquisition module and a continuous optimization module;
[0025] Feedback Acquisition Module: Collects source tracing and location results, deviation cause determination conclusions, and subsequent maintenance verification data to form a standardized feedback dataset;
[0026] Continuous optimization module: Dynamically optimizes based on feedback datasets, including automatically updating the baseline frame when the scene changes normally, adjusting the multimodal feature fusion weights according to the causes of deviations, and optimizing the deviation propagation prediction model.
[0027] Furthermore, the RGB camera and infrared thermal imaging camera of the multimodal visual perception module are coaxially mounted to ensure consistent acquisition angles, and the acquired data is synchronized and aligned using timestamps.
[0028] The beneficial effects of this invention's deviation warning and positioning system utilizing machine vision are as follows: Through the deep coupling of multimodal temporal semantic fusion and source-tracing positioning, it can comprehensively cover the entire process of machine vision deviation management. It not only focuses on traditional deviation warning and coordinate positioning but also incorporates elements such as dynamic trend prediction, cause tracing, and closed-loop adaptive adjustment. Furthermore, by deeply integrating multimodal perception, temporal semantic information, and physical cause analysis, it improves the timeliness of deviation warning, the accuracy of positioning, and the effectiveness of source tracing compared to existing technologies, providing a more scientific technical solution for deviation management in scenarios such as industrial production and robot navigation. In addition, through dynamic deviation... The integrated design of early warning and source tracing in deviation propagation breaks through the existing technology's problem of separating early warning and location, and only "knowing the deviation" but "not knowing the cause." It enables the synergistic optimization of deviation-level early warning, precise coordinate output, and cause tracing, improving the efficiency and pertinence of deviation management, and significantly reducing false alarm and missed alarm rates and rectification and investigation costs. At the same time, through closed-loop feedback adjustment and multimodal feature matching, it provides a more intelligent and adaptive system operation mode, which can dynamically adapt to changes in scenarios and deviation evolution patterns. Furthermore, the three-dimensional coordinate output and cause analysis based on source tracing greatly enhance the system's robustness and decision support capabilities in complex scenarios. Attached Figure Description
[0029] The present invention will now be described in further detail with reference to the accompanying drawings and specific implementation methods.
[0030] Figure 1 This is a schematic diagram of the system module architecture. Detailed Implementation
[0031] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the present application can be combined with each other.
[0032] According to one aspect of the invention, such as Figure 1 As shown, a deviation warning and localization system utilizing machine vision is provided, including: a multimodal visual perception module that acquires RGB images, infrared images, and temporal continuous frame data, and extracts scene semantic features by combining a semantic segmentation model. This module includes:
[0033] Multi-source data acquisition module: Simultaneously acquires scene RGB images, infrared images, and time-series continuous frame data through RGB cameras and infrared thermal imaging cameras;
[0034] Specifically, the RGB camera and the infrared thermal imaging camera are fixed with a rigid bracket to achieve coaxial installation. The optical axes of the lenses of the two cameras are kept completely aligned, and the distance between the center of the lenses is controlled within a certain range (e.g., 5cm) to ensure that the acquisition angle of the same target scene is completely consistent and to avoid multimodal data misalignment caused by viewing angle deviation.
[0035] Simultaneously, the synchronous trigger signal generator outputs the same trigger pulse, which is connected to the trigger interfaces of the RGB camera and the infrared thermal imaging camera respectively. The trigger frequency is dynamically set according to the application scenario to realize the synchronous acquisition of the two types of image data, ensuring that the RGB texture information and infrared thermal radiation information of the target scene are captured at the same time.
[0036] In addition, after each frame is acquired, the camera’s built-in timestamp module adds a unified timestamp to the RGB and infrared images. The timestamp is based on the system’s UTC clock calibration. Subsequently, the two types of image data are aligned frame by frame using a timestamp matching algorithm to eliminate the data asynchrony problem caused by acquisition delay.
[0037] Finally, the time-series continuous frame data was acquired through the camera's continuous shooting mode. During the acquisition process, the trigger frequency remained stable, and the time interval between consecutive frames was uniform. The acquired RGB images, infrared images, and time-series continuous frame data were stored together and a lossless compression format was used to preserve the original data details, providing a complete and synchronous multimodal data source for subsequent feature extraction.
[0038] Feature extraction module: Using the U-Net semantic segmentation model, key object categories, core component locations, and component relationships are extracted from the collected multimodal data to generate a scene semantic feature set;
[0039] The synchronized RGB image and infrared image are used as the dual input channels of the U-Net semantic segmentation model. Before input, the two types of images are normalized and mapped to the [0,1] interval, while keeping the spatial resolution of the images unchanged, so that the model can simultaneously utilize the texture detail features of the RGB image and the environmental adaptation features of the infrared image.
[0040] The U-Net model extracts features from multimodal input data through an encoder-decoder structure. In the encoding stage, convolutional and pooling layers are used to gradually compress the spatial dimension of the image and improve the level of feature abstraction, focusing on capturing the global features of key objects. In the decoding stage, upsampling and skip connections are used to restore the spatial dimension and accurately locate the position of core components. During model training, the object categories and component positions actually labeled in the scene are used as supervision signals to optimize model parameters and improve feature extraction accuracy.
[0041] In addition, for the extracted key object category features, the semantic association rule base is used to analyze the component association relationship. Based on the spatial proximity and relative posture relationship of the core component location, the logical relationship such as assembly association and motion association between different components is determined. For example, in the industrial assembly scenario, the fixed association relationship between "conveyor belt-workpiece" and "fixture-workpiece" is identified, and this association relationship is used as an important part of the semantic features.
[0042] Finally, the extracted key object category labels, core component coordinate information, and component relationship data are structurally integrated to generate a standardized scene semantic feature set. The feature set is stored in JSON format and contains semantic description fields and numerical feature vectors corresponding to each frame of multimodal data, ensuring that the subsequent semantic temporal integration module can directly call the feature set to perform deep fusion of motion features and semantic features.
[0043] The semantic-temporal integration module, based on optical flow and semantic similarity calculation, achieves deep fusion of motion features and semantic features, generating a temporal-semantic fusion feature map. This module includes:
[0044] Motion feature calculation module: The Lucas-Kanade optical flow method is used to calculate the motion vectors between consecutive temporal frames, capture the dynamic motion trajectory of objects, and filter effective motion information;
[0045] Specifically, the sequential frames output by the multimodal visual perception module are preprocessed by first removing image noise through Gaussian filtering and then converting the images into grayscale images to ensure the stability and consistency of information between frames.
[0046] Meanwhile, by selecting a set of feature points in the grayscale image, the feature points are screened using the Shi-Tomasi corner detection algorithm, prioritizing edge points and corner points with drastic grayscale changes as calculation samples, and defining a neighborhood window of, for example, 3×3 around each feature point.
[0047] Furthermore, the motion vector is calculated based on the formula of the Lucas-Kanade optical flow method, as follows:
[0048]
[0049] In the formula, and This represents the gradient of the image grayscale in the x and y directions. To accurately calculate the horizontal and vertical motion components of each feature point between frames, the grayscale change rate is determined.
[0050] Finally, effective information is filtered by the amplitude and direction of the motion vectors. A threshold for the amplitude of the motion vectors is set (determined based on the statistical analysis of normal motion speed in the scene). Abnormal noise vectors with amplitudes exceeding the threshold are removed, while vectors with consistent motion directions for multiple consecutive frames are retained to form an effective motion vector set. The dynamic motion trajectory of the object is fitted based on this vector set to ensure the continuity and accuracy of the trajectory.
[0051] Semantic association module: The cosine similarity algorithm is used to calculate the semantic feature matching degree between the current frame and the reference frame, and to quantify the consistency of semantic information;
[0052] The numerical feature vectors corresponding to the current frame and the dynamically updated reference frame are extracted from the scene semantic feature set output by the multimodal visual perception module. The feature vectors contain quantitative data on key object categories, core component positions and component relationships to ensure that the data dimensions are completely consistent.
[0053] Then, the cosine similarity formula is used to calculate the matching degree between the two types of feature vectors, as follows:
[0054]
[0055] In the formula, This is the semantic feature vector of the current frame. The semantic feature vector of the reference frame is used to quantify the similarity of semantic information between the two frames by the ratio of the vector dot product to the magnitude.
[0056] At the same time, the validity of the calculated cosine similarity results is verified. Combined with the scene semantic rule library, the fluctuation of matching degree caused by local minor deformation is excluded. For example, when the position of the core component has not changed substantially, even if the matching degree decreases slightly, it is still judged as semantically consistent.
[0057] Finally, the verified cosine similarity value is mapped to a semantic consistency coefficient in the range [0,1]. The closer the coefficient is to 1, the more consistent the semantic information is, and the closer it is to 0, the greater the semantic difference is. This coefficient will serve as the core basis for subsequent fusion weight allocation.
[0058] Temporal semantic fusion module: Based on motion vector consistency, abnormal motion points are removed, and weights are allocated by combining semantic matching degree. Motion features and semantic features are deeply fused to generate temporal semantic fusion feature map;
[0059] Specifically, the effective motion vector set output by the motion feature calculation module is first screened a second time. Based on the principle of motion vector consistency, motion points that deviate significantly from the overall motion direction of the region are removed to ensure that the retained motion features can truly reflect the dynamic changes of the object.
[0060] It also includes fusion weight allocation based on the semantic consistency coefficient output by the semantic association module. The higher the semantic consistency coefficient, the more stable the semantic information in the region, and the higher the weight is assigned to the semantic features. The lower the semantic consistency coefficient, the more significant the dynamic changes, and the higher the weight is assigned to the motion features, thus realizing adaptive dynamic adjustment of weights.
[0061] Simultaneously, feature fusion is performed, which integrates the motion vector information of motion features with the category, location, and correlation information of semantic features. Each motion component and semantic attribute are weighted and superimposed to form a composite feature that includes dynamic motion and static semantics.
[0062] Finally, the fused composite features are subjected to dimensional regularization and normalization to eliminate fusion bias caused by feature scale differences, generating a temporal semantic fusion feature map. This feature map simultaneously carries dynamic motion trajectory, multimodal environment adaptation information, and semantic association information, providing a unified high-quality feature input for subsequent dynamic deviation warning and source tracing.
[0063] The dynamic deviation early warning module uses a dynamically updated baseline frame as a reference. Through deviation quantification, trend prediction, and dynamic threshold determination, it triggers tiered early warnings and associates semantic information. This module includes:
[0064] Deviation quantization module: Using a dynamically updated reference frame as a reference, and based on the temporal semantic fusion feature map, calculates the feature difference between the current frame and the reference frame to achieve deviation quantization representation;
[0065] The dynamic update of the reference frame comes from the closed-loop adjustment results of the feedback and optimization module. If there is no abnormal change in the scene, the average frame of the recent stable frame set is used. If the scene changes normally (e.g., the workpiece model changes), it is replaced with the initial stable frame of the new scene to ensure that the reference benchmark matches the actual scene in real time.
[0066] Based on the temporal semantic fusion feature map, a composite feature vector of the current frame and the reference frame is extracted. The composite feature vector covers dynamic motion trajectory, semantic relationship and multimodal environment adaptation information. The feature difference is quantified by calculating the Euclidean distance between the two vectors. The larger the distance value, the more significant the deviation.
[0067] In addition, the feature difference is normalized and mapped to the [0,1] interval to eliminate the quantization deviation caused by the difference in feature scale under different scenarios. At the same time, the feature difference corresponding to the core component is weighted and amplified by combining the importance weight of semantic features to ensure that the deviation of key areas is not masked by the overall mean.
[0068] Finally, the weighted normalized difference is used as the final deviation quantification value. This value reflects both the magnitude and the scope of the deviation, providing a precise quantitative basis for subsequent trend prediction and threshold determination.
[0069] Trend prediction module: It uses the exponential smoothing method to analyze the time-series propagation law of deviation, predicts the evolution trend of deviation in the next frame, and avoids false early warnings triggered by instantaneous deviation;
[0070] Specifically, the deviation quantization values of multiple consecutive historical frames are collected to construct a deviation time series. The sequence length is adaptively adjusted according to the scene dynamics to ensure that the complete change cycle of the deviation is covered.
[0071] Simultaneously, the biased time series is trend-fitted using exponential smoothing, with the following formula:
[0072]
[0073] In the formula, For the first The first exponential smoothing value of the frame, For the first Frame deviation quantization value, For smoothing coefficients;
[0074] Based on the fitted smooth sequence, the rate of change and acceleration of the deviation are calculated. The rate of change is used to determine whether the deviation is expanding, shrinking or stabilizing, and the acceleration is used to determine the degree of change of the deviation. For example, a positive rate and positive acceleration indicate that the deviation is expanding rapidly, while a positive rate but negative acceleration indicates that the deviation is expanding slowly.
[0075] Finally, the trend prediction conclusion is output through the change rate and acceleration results to clarify the evolution direction and magnitude of the deviation in the next frame. For example, if the prediction result is "recovery to stability after instantaneous fluctuation", it is marked as a non-risk trend. If it is "continuous expansion" or "rapid expansion", it is marked as a risk trend, thus avoiding false early warnings caused by instantaneous deviations from the root cause.
[0076] Dynamic threshold generation module: Based on the statistical distribution of historical deviation data, the threshold parameters are adaptively adjusted according to the 3σ principle to generate dynamic early warning thresholds that match scene changes in real time;
[0077] The process involves selecting valid samples from historical deviation data, removing extreme deviation values caused by abnormal interference (such as sudden occlusion or equipment failure), and retaining deviation data under normal scene conditions to ensure that the statistical distribution reflects the true range of deviation fluctuations.
[0078] In addition, statistical analysis was performed on the selected valid samples to calculate the mean and standard deviation of the samples, and the basic threshold range was determined based on the 3σ principle.
[0079] Simultaneously, the threshold parameters are adjusted based on the dynamic changes in the scenario. If the scenario is in a stable operating phase, the basic threshold remains unchanged. If the scenario changes (e.g., workpiece model change, workflow adjustment), the deviation data distribution under the new scenario is re-statistically analyzed, and the mean and standard deviation are updated in real time to generate a dynamic threshold adapted to the new scenario. If the deviation trend prediction shows that the risk level is rising, the threshold range can be temporarily narrowed to improve the early warning sensitivity. The adjusted threshold is stored, and the latest threshold is called after each frame is collected and processed to ensure that the threshold is synchronized with the scenario changes and deviation trends in real time.
[0080] The graded early warning module: When the deviation quantification value exceeds the dynamic threshold and the trend prediction shows that the deviation is expanding, it triggers three levels of early warning: mild, moderate and severe. The early warning information is synchronously associated with the corresponding semantic feature set.
[0081] The criteria for determining the three-level early warning system are based on a combination of the proportion of the deviation quantification value exceeding the dynamic threshold and the rate of trend expansion: for example, if the deviation quantification value exceeds the threshold by 0-30% and the trend expands slowly, a mild early warning is triggered; if it exceeds the threshold by 30%-60% and the trend continues to expand, a moderate early warning is triggered; and if it exceeds the threshold by more than 60% and the trend expands rapidly, a severe early warning is triggered, ensuring the accuracy and rationality of the classification.
[0082] The generation of early warning information needs to be synchronized with the scene semantic feature set output by the semantic and temporal integration module, clearly marking the key object categories, core component locations and component relationships associated with the deviation, such as "mild warning - workpiece positioning hole deviation - associated fixture component - deviation slowly expanding", so that staff can quickly locate the deviation-associated object;
[0083] Finally, the early warning information is simultaneously pushed to the feedback and optimization module as a reference for updating the baseline frame and adjusting the fusion weights, realizing the linkage between early warning and closed-loop optimization, and ensuring that the early warning can not only indicate risks, but also support the system's adaptive adjustment.
[0084] The source tracing and localization module analyzes the propagation path and physical causes of the deviation, and simultaneously outputs the three-dimensional coordinates of the deviation's starting point and the cause determination result. This module includes:
[0085] Deviation tracing module: Based on temporal semantic fusion feature map, reverse analysis is performed to analyze the propagation path of deviation from the initial occurrence area to the diffusion area, and the starting point of deviation is located.
[0086] Specifically, the current frame when the dynamic deviation warning module triggers the warning is taken as the starting node. The historical temporal semantic fusion feature map is retrieved in reverse order of the time axis. The feature difference distribution of the deviation-related region is compared frame by frame to track the complete evolution process of the difference from nothing to something, from local diffusion to widening range.
[0087] Simultaneously, by combining the historical motion vector set output by the motion feature calculation module, the motion trajectory of the deviation-related area is reversed and restored. Through the continuity verification of the trajectory, the interference of feature changes caused by normal operation motion is eliminated, and the propagation law specific to the deviation is focused.
[0088] In addition, through semantic feature continuity analysis, the semantic consistency coefficient of each frame in the backtracking process is compared one by one. When the semantic consistency coefficient changes and is accompanied by a significant increase in feature difference, the region corresponding to that frame is marked as the region where the deviation initially occurred.
[0089] Finally, based on the spatial coordinates, diffusion direction, and propagation rate of the initial appearance area, a path fitting algorithm is used to construct the complete propagation path of the deviation. The path is visualized with the starting point as the origin, the diffusion direction as the vector, and the propagation range as the radius, accurately locking the coordinates of the starting point of the deviation.
[0090] Semantic analysis module: Relates semantic categories and component relationships in the semantic feature set of the associated scene to determine the physical causes of deviations;
[0091] Specifically, the semantic feature information corresponding to the deviation starting point is extracted, and the key object category, core component name and related component list of the starting point are accurately matched from the scene semantic feature set. The physical attributes and functional positioning of the starting point are clarified, such as determining the specific functional components such as "camera mounting base", "workpiece clamping claw" and "conveyor belt guide rail" corresponding to the starting point.
[0092] In addition, by combining the design functions and working principles of the components, the interaction between the starting point and related components is analyzed to identify potential factors that may cause deviations. At the same time, the feature differences of multimodal data are linked to assist in the judgment. For example, if the RGB image features are normal but the infrared image features show significant fluctuations, it is determined to be a deviation caused by light interference or thermal deformation. If the trajectory of the starting point changes abruptly in a continuous temporal frame without corresponding semantic changes, it is determined to be a deviation caused by external occlusion or collision. Finally, a clear physical cause judgment conclusion is formed, including the deviation type, responsible component, and core cause.
[0093] 3D coordinate positioning module: By using the depth complementarity of RGB and infrared images and employing multimodal feature matching technology, the 3D spatial coordinates of the deviation starting point are calculated;
[0094] Using the coordinates of the starting point of the deviation as a reference, local images of the corresponding regions are cropped from the synchronously aligned RGB and infrared images, and the texture detail features (RGB image) and thermal radiation features (infrared image) of the local images are extracted respectively to construct a multimodal local feature set;
[0095] Simultaneously, by invoking the camera calibration parameters (such as intrinsic and extrinsic parameters) and the spatial constraints of coaxial mounting, the optimized coordinates are converted into initial values of three-dimensional coordinates in the camera coordinate system. The texture details of RGB images are used to improve the accuracy of the horizontal coordinates (X and Y axes), and the environmental adaptability of infrared images is used to improve the depth accuracy of the vertical coordinates (Z axis), achieving depth complementarity between the two types of data. Then, a dynamic compensation mechanism for time-series frames is introduced to extract the initial values of three-dimensional coordinates of the starting point of deviation in multiple consecutive frames. The sliding window averaging algorithm is used to eliminate the error caused by instantaneous jitter. Finally, the three-dimensional coordinates in the camera coordinate system are converted into standard coordinates in the coordinate system, providing accurate position references for subsequent rectification operations.
[0096] The feedback and optimization module feeds back the source tracing and localization results, dynamically optimizing the baseline frame, fusion weights, and bias propagation model. This module includes:
[0097] Feedback Acquisition Module: Collects source tracing and location results, deviation cause determination conclusions, and subsequent maintenance verification data to form a standardized feedback dataset;
[0098] Specifically, it connects to the output interface of the source tracing and positioning module in real time, and synchronously collects the three-dimensional coordinates of the deviation starting point, the deviation propagation path, the semantic information related to the physical cause determination and the early warning, to ensure the integrity and timeliness of data collection;
[0099] It also includes collecting subsequent maintenance operation data, covering information such as maintenance measure type, operation execution parameters, and maintenance completion time. At the same time, it collects system operation data after maintenance, including verification data such as deviation recurrence, early warning response speed, and changes in positioning accuracy, forming a complete data link of "tracing the source - maintenance - verification".
[0100] In addition, the collected multi-source data is standardized to unify the data format and units of measurement, eliminate invalid and redundant data, and ensure the logical consistency of traceability results, maintenance operations, and verification data through data correlation verification. Finally, a feedback dataset is built using structured storage to provide high-quality data support for continuous optimization.
[0101] Continuous optimization module: performs dynamic optimization based on feedback dataset, including automatically updating the baseline frame when the scene changes normally, adjusting the multimodal feature fusion weights according to the causes of deviation, and optimizing the deviation propagation prediction model;
[0102] For the optimization of the reference frame, by analyzing the types of deviation causes in the feedback dataset, if it is determined to be a normal change in the scene (such as workpiece model switching or workflow adjustment), the multimodal data frame of the stable operation phase in that scene is extracted, the original reference frame is replaced, and the historical reference frame version is saved. If it is determined to be abnormal interference (such as occlusion or temporary component failure), the original reference frame is retained, and the interference area and interference type are marked in the frame to avoid the interference affecting the subsequent deviation determination.
[0103] In addition, for the adjustment of multimodal feature fusion weights, a weight adjustment rule base is established by combining the causes of deviations and maintenance verification results: if the cause of deviations is environmental interference such as changes in illumination or severe weather, the weight of infrared image features in temporal semantic fusion is increased; if the deviations originate from changes in component posture or abnormal motion trajectories, the fusion weight of motion features is increased; if the semantic association deviation of core components leads to misjudgment, the weight ratio of semantic features is strengthened, thereby realizing adaptive dynamic adjustment of fusion weights.
[0104] Finally, for the optimization of the deviation propagation prediction model, based on the historical deviation time series data, trend prediction results and actual deviation evolution in the feedback dataset, the model prediction error is calculated, and the key parameters of the exponential smoothing method are optimized by iterative training. At the same time, combined with the propagation laws of different deviation causes, a classified deviation propagation prediction sub-model is constructed so that the model can accurately adapt to the deviation evolution characteristics under various scenarios and continuously reduce early warning delay and positioning error.
[0105] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention are also within the protection scope of the present invention.
Claims
1. A deviation warning and positioning system utilizing machine vision, characterized in that, include: The system comprises a multimodal visual perception module, a semantic temporal integration module, a dynamic deviation early warning module, a source tracing and localization module, and a feedback and optimization module. The multimodal visual perception module acquires RGB images, infrared images, and continuous temporal frame data, and extracts scene semantic features using a semantic segmentation model. The semantic temporal integration module is based on optical flow and semantic similarity calculation to achieve deep fusion of motion features and semantic features, and generate a temporal semantic fusion feature map. The dynamic deviation early warning module uses a dynamically updated reference frame as a reference, and triggers graded early warnings and associates semantic information through deviation quantification, trend prediction and dynamic threshold determination; the source tracing and positioning module analyzes the deviation propagation path and physical cause, and simultaneously outputs the three-dimensional coordinates of the deviation starting point and the cause determination result; the feedback and optimization module feeds back the source tracing and positioning results, and dynamically optimizes the reference frame, fusion weight and deviation propagation model.
2. The deviation warning and positioning system using machine vision according to claim 1, characterized in that: The multimodal visual perception module includes: a multi-source data acquisition module and a feature extraction module; Multi-source data acquisition module: Simultaneously acquires scene RGB images, infrared images, and time-series continuous frame data through RGB cameras and infrared thermal imaging cameras; Feature extraction module: Using the U-Net semantic segmentation model, key object categories, core component locations and component relationships are extracted from the collected multimodal data to generate a scene semantic feature set.
3. The deviation warning and positioning system using machine vision according to claim 1, characterized in that: The semantic temporal integration module includes: a motion feature calculation module, a semantic association module, and a temporal semantic fusion module; Motion feature calculation module: The Lucas-Kanade optical flow method is used to calculate the motion vectors between consecutive temporal frames, capture the dynamic motion trajectory of objects, and filter effective motion information; Semantic association module: The cosine similarity algorithm is used to calculate the semantic feature matching degree between the current frame and the reference frame, and to quantify the consistency of semantic information; Temporal semantic fusion module: Based on motion vector consistency, abnormal motion points are removed, and weights are allocated by combining semantic matching degree. Motion features and semantic features are deeply fused to generate temporal semantic fusion feature map.
4. The deviation warning and positioning system using machine vision according to claim 1, characterized in that: The dynamic deviation early warning module includes: a deviation quantification module, a trend prediction module, a dynamic threshold generation module, and a graded early warning module; Deviation quantization module: Using a dynamically updated reference frame as a reference, and based on the temporal semantic fusion feature map, calculates the feature difference between the current frame and the reference frame to achieve deviation quantization representation; Trend prediction module: It uses the exponential smoothing method to analyze the time-series propagation law of deviation, predicts the evolution trend of deviation in the next frame, and avoids false early warnings triggered by instantaneous deviation; Dynamic threshold generation module: Based on the statistical distribution of historical deviation data, the threshold parameters are adaptively adjusted according to the 3σ principle to generate dynamic early warning thresholds that match scene changes in real time; Tiered early warning module: When the deviation quantification value exceeds the dynamic threshold and the trend prediction shows that the deviation is expanding, a three-level early warning is triggered: mild, moderate and severe. The early warning information is synchronously associated with the corresponding semantic feature set.
5. A deviation warning and positioning system utilizing machine vision according to claim 1, characterized in that: The source tracing and localization module includes: a deviation tracing module, a semantic analysis module, and a three-dimensional coordinate localization module; Deviation tracing module: Based on temporal semantic fusion feature map, reverse analysis is performed to analyze the propagation path of deviation from the initial occurrence area to the diffusion area, and the starting point of deviation is located. Semantic analysis module: Relates semantic categories and component relationships in the semantic feature set of the associated scene to determine the physical causes of deviations; 3D coordinate positioning module: By using the depth complementarity of RGB and infrared images and employing multimodal feature matching technology, the 3D spatial coordinates of the deviation starting point are calculated.
6. A deviation warning and positioning system using machine vision according to claim 1, characterized in that: The feedback and optimization module includes: a feedback acquisition module and a continuous optimization module; Feedback Acquisition Module: Collects source tracing and location results, deviation cause determination conclusions, and subsequent maintenance verification data to form a standardized feedback dataset; Continuous optimization module: Dynamically optimizes based on feedback datasets, including automatically updating the baseline frame when the scene changes normally, adjusting the multimodal feature fusion weights according to the causes of deviation, and optimizing the deviation propagation prediction model.
7. A deviation warning and positioning system using machine vision according to claim 2, characterized in that: The RGB camera and infrared thermal imaging camera of the multimodal visual perception module are coaxially mounted to ensure consistent acquisition angles, and the acquired data is synchronized through timestamps.
Citation Information
Patent Citations
A high-precision positioning method for machine vision based on ROS
CN107818587B
A machine vision positioning method, device and system
CN118518005B