AR-based factory object identification system and method
By introducing multi-scale ViT and CNN feature extraction modules into the factory AR object recognition system and adopting adaptive fusion and multi-scale fusion technology, the problem that the system cannot effectively fuse features of different scales is solved, and high-precision object recognition is achieved.
Patent Information
- Application Number
- CN202510003587.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-13
AI Technical Summary
The existing factory AR object recognition system cannot effectively integrate feature information at different scales, resulting in insufficient detection accuracy.
The multi-scale ViT module, CNN feature extraction module, ViT-CNN adaptive fusion module and multi-scale feature fusion module are adopted to achieve a high degree of fusion of semantics and details through adaptive fusion and multi-scale fusion.
It improves the accuracy and reliability of object recognition, and can accurately detect and identify key production equipment, parts and other target objects from complex factory environments.
Smart Images

Figure CN119992395A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of object recognition technology, and in particular to an AR-based factory object recognition system and method. Background Art
[0002] With the continuous development of computer vision and augmented reality (AR) technology, it has become an important trend to apply AR technology to industrial production environments to improve production efficiency and safety. In factory workshops, workers need to quickly and accurately identify numerous equipment, parts and other objects. Traditional manual visual recognition methods are inefficient and prone to errors. Therefore, the development of an AR-based intelligent object recognition system can assist workers in real-time identification of key objects and display object information in an intuitive AR form, which has important application value.
[0003] Existing factory AR object recognition systems usually use a convolutional neural network (CNN)-based target detection model to identify and locate objects in images. However, this solution has a major flaw: it cannot effectively integrate feature information at different scales, resulting in insufficient detection accuracy. Summary of the invention
[0004] The present application provides an AR-based factory object recognition system and method, which can improve object recognition accuracy.
[0005] In a first aspect, the present application provides an AR-based factory object recognition system, the system comprising: an image acquisition module, a target detection module and a display module; An image acquisition module, used to acquire an image of the factory in front of the AR device through a camera of the AR device; The object detection module is used to identify objects in factory images and obtain object recognition results; The display module is used to perform AR rendering on the factory image according to the object recognition result to obtain an AR image, and present the AR image on a display screen of the AR device.
[0006] By adopting the above technical solutions, firstly, the system can obtain the image data of the factory site in front of the AR device in real time through the image acquisition module using the camera of the AR device. This not only ensures the timeliness and accuracy of image acquisition, but also lays a solid data foundation for subsequent target detection and AR rendering. Next, the core of the system is the innovative target detection module. This module performs high-precision object recognition on the acquired factory images and obtains accurate object recognition results. It is worth mentioning here that the target detection module adopts advanced image preprocessing, feature extraction, detection head and post-processing technologies, which can accurately detect and identify key production equipment, parts and other target objects from complex factory environments. Among them, the feature extraction module introduces innovative technologies such as multi-scale ViT module, CNN feature extraction module, ViT-CNN adaptive fusion module and multi-scale feature fusion module. The multi-scale ViT module uses the visual converter architecture to efficiently extract the long-range dependency and global semantic features of the image; the CNN feature extraction module is good at capturing local details and texture features. The features of the two different paradigms are adaptively fused to achieve a high degree of fusion of semantics and details. Finally, through multi-scale feature fusion, the system obtains rich multi-scale target feature representation, laying a solid foundation for high-precision object recognition.
[0007] After the object recognition results are further optimized by the post-processing module, the final object recognition results of the system can be output. It is worth mentioning that the post-processing module uses technologies such as non-maximum suppression, confidence threshold filtering and detection box decoding to ensure the accuracy and reliability of the recognition results.
[0008] Finally, the display module presents the above object recognition results clearly and intuitively on the display screen of the AR device through AR rendering, so that workers can view and refer to them in real time. Through the cooperation of the result rendering module and the 3D space projection module, the system can transform the 2D rendering results into highly immersive AR images, effectively improving the image and vividness of information transmission.
[0009] In a second aspect of the present application, a factory object recognition method based on AR is provided, comprising: Acquire the factory image in front of the AR device through the camera of the AR device; Perform object recognition on factory images to obtain object recognition results; The object recognition results are presented in the form of AR on the display screen of the AR device.
[0010] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. This application uses the camera of the AR device through the image acquisition module, and the system can obtain the image data of the factory site in front of the AR device in real time. This not only ensures the timeliness and accuracy of image acquisition, but also lays a solid data foundation for subsequent target detection and AR rendering. Next, the core of the system is the innovative target detection module. This module performs high-precision object recognition on the acquired factory images to obtain accurate object recognition results. It is worth mentioning here that the target detection module uses advanced image preprocessing, feature extraction, detection head and post-processing technologies, which can accurately detect and identify key production equipment, parts and other target objects from complex factory environments. Among them, the feature extraction module introduces innovative technologies such as multi-scale ViT module, CNN feature extraction module, ViT-CNN adaptive fusion module and multi-scale feature fusion module. The multi-scale ViT module uses the visual converter architecture to efficiently extract the long-range dependency and global semantic features of the image; the CNN feature extraction module is good at capturing local details and texture features. The features of the two different paradigms are adaptively fused to achieve a high degree of fusion of semantics and details. Finally, through multi-scale feature fusion, the system obtains a rich multi-scale target feature representation, laying a solid foundation for high-precision object recognition.
[0011] 2. After further optimization of the object recognition results of this application by the post-processing module, the final object recognition results of the system can be output. It is worth mentioning that the post-processing module uses technologies such as non-maximum suppression, confidence threshold filtering and detection frame decoding to ensure the accuracy and reliability of the recognition results.
[0012] 3. The display module of the present application presents the above object recognition results clearly and intuitively on the display screen of the AR device in the form of AR rendering, so that workers can view and refer to them in real time. Through the cooperation of the result rendering module and the 3D space projection module, the system can convert the 2D rendering results into highly immersive AR images, effectively improving the image and vividness of information transmission. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 An architecture diagram of an AR-based factory object recognition system provided in an embodiment of the present application; Figure 2 An architecture diagram of a target detection module provided in an embodiment of the present application; Figure 3 A flowchart of an AR-based factory object recognition method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0014] In order to enable technicians in this field to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0015] In the description of the embodiments of the present application, words such as "for example" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "for example" or "for example" is intended to present related concepts in a specific way.
[0016] In the description of the embodiments of the present application, the meaning of the term "multiple" refers to two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "include", "comprise", "have" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0017] In order to facilitate understanding of the method and system provided by the embodiments of the present application, before introducing the embodiments of the present application, the background of the embodiments of the present application is first introduced.
[0018] With the rapid development of computer vision and augmented reality (AR) technology, integrating AR technology into industrial production environments to improve production efficiency and operational safety has become an important trend and development direction in the industry. In the actual production process of factory workshops, workers need to quickly and accurately identify and distinguish numerous production equipment, parts, materials and other target objects, while traditional manual visual recognition methods are not only inefficient, but also prone to misjudgment and cannot fully meet the needs of real-time and efficient recognition. Therefore, it is of great theoretical significance and application value to develop an intelligent object recognition system based on AR technology to assist workers in accurately identifying key target objects in real time and present object-related information in an intuitive and vivid AR visualization form.
[0019] At present, most of the existing factory AR object recognition systems use a target detection model based on a convolutional neural network (CNN) to identify and locate objects in images. However, this solution based on a single CNN architecture has a key flaw - it cannot effectively fuse and mine feature information at multiple scales, resulting in less than ideal accuracy and robustness of target detection. Specifically, because the CNN model is based on local convolution operations, it has limitations in capturing global semantics and long-range dependencies. At the same time, it is also difficult for CNN to efficiently integrate features of multiple scales at different scale levels. In a factory environment, the size, shape, and scale changes of the target to be identified are often very complex. It is difficult for a single CNN architecture to fully mine multi-level semantics and scale features. Therefore, the detection accuracy and real-time performance of the system are affected and limited to a certain extent.
[0020] After the background introduction of the above content, those skilled in the art can understand the problems existing in the prior art. The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.
[0021] Based on the above background technology, please refer to Figure 1 , Figure 1 A schematic diagram of the architecture of an AR-based factory object recognition system provided in an embodiment of the present application, the system can be implemented by a computer program, or can be run as an independent tool application. Specifically, in an embodiment of the present application, the method can be applied on a server, but can also be applied to electronic devices such as servers. An AR-based factory object recognition system includes the following steps: an image acquisition module, a target detection module, and a display module; An image acquisition module 1 is used to acquire an image of the factory in front of the AR device through a camera of the AR device; Specifically, the image acquisition module uses the camera of the AR device to collect image data of the factory site in front of the AR device in real time. Using the camera of the AR device itself for image acquisition can ensure that the acquired image data is highly consistent with the worker's actual perspective and field of view, and truly reflects the on-site environment currently seen by the worker. At the same time, AR devices usually have functions such as motion tracking and perspective detection, and can adjust the image acquisition perspective in real time according to the worker's position and posture changes, so that the acquired image data always remains consistent with the worker's field of view, laying a solid foundation for subsequent image processing and AR rendering.
[0022] The collected factory image data is then input into the target detection module. The target detection module is the core and key of the entire system. It is responsible for object recognition and target detection of factory images to obtain accurate object recognition results. The module integrates advanced image preprocessing, feature extraction, target detection head and post-processing technologies, which can accurately identify and detect key production equipment, parts, tools and other target objects from complex factory environment images. Among them, the feature extraction stage introduces innovative technologies such as multi-scale visual transformer (ViT) module, convolutional neural network (CNN) feature extraction module, ViT-CNN adaptive fusion module and multi-scale feature fusion module. The multi-scale ViT module can efficiently learn the global semantics and long-range dependency features of the image through the self-attention mechanism; the CNN feature extraction module is good at extracting local details and texture features. The features of the two different paradigms are fused through the adaptive fusion module, so that the system obtains rich semantics, details and location information. Afterwards, the multi-scale feature fusion module fuses the above-mentioned features of different scales to generate high-quality target feature representation, laying a solid foundation for high-precision detection.
[0023] Finally, the object recognition results obtained after post-processing optimization are transmitted to the display module. The display module presents the object recognition results on the display screen of the AR device through AR visualization rendering. Workers only need to wear AR devices to view and obtain relevant information of the recognized targets in real time. Specifically, the result rendering module first renders the object recognition results into 2D image information, and then the 3D space projection module projects these 2D rendering contents into immersive AR images and projects them onto the display screen of the AR device, allowing workers to superimpose the object recognition results on the real scene to achieve AR enhanced visualization.
[0024] The object detection module 2 is used to perform object recognition on the factory image and obtain object recognition results; Specifically, the target detection module includes multiple sub-modules such as image preprocessing module, feature extraction module, target detection head module and post-processing module. The modules work closely together to realize an end-to-end target detection process.
[0025] First, the image preprocessing module performs a series of preprocessing operations on the original factory image data, including image cropping, image enhancement, image normalization, and image format conversion. These preprocessing steps can effectively eliminate noise and redundant information in the image, improve image quality, and lay a good data foundation for subsequent feature extraction and target detection.
[0026] After preprocessing, the optimized factory image data is passed to the feature extraction module. Feature extraction is the key link of target detection. The module integrates innovative technologies such as the multi-scale ViT module, CNN feature extraction module, ViT-CNN adaptive fusion module and multi-scale feature fusion module. The multi-scale ViT module can efficiently learn the global semantics and long-range dependency features of the image; the CNN module is good at extracting local details and texture features. The features of the two different paradigms are fused through the adaptive fusion module, and finally a high-quality target feature representation is generated in the multi-scale feature fusion module, which can fully characterize the semantics, details and location information of the target object at different scales. Next, the target detection head module inputs the above target feature representation into the target detection head and applies the target detection algorithm to obtain preliminary detection results. After that, it is further optimized through the post-processing module, including non-maximum suppression, confidence threshold filtering and detection frame decoding, so as to eliminate redundant detection frames, improve detection confidence, and finally output accurate and reliable object recognition results.
[0027] Based on the above embodiment, as an optional embodiment, please refer to Figure 2 , Figure 2 This is an architecture diagram of a target detection module provided in an embodiment of the present application. The target detection module includes: an image preprocessing module 21, a feature extraction module 22, a target detection head module 23 and a post-processing module 24; An image preprocessing module 21 is used to preprocess the factory image to obtain a preprocessed engineering image; Specifically, image preprocessing is the basic part of the target detection process. Its purpose is to eliminate noise and redundant information in the original image, improve image quality, and lay a good data foundation for subsequent feature extraction and target detection. The image preprocessing module usually performs the following main preprocessing steps: The first is image cropping, which is to crop the factory image according to the approximate position and scale of the target to be detected, and remove the background area in the image that is not related to the detection target, thereby reducing the amount of calculation and focusing on the target area. Next is image enhancement, including brightness adjustment, contrast enhancement, sharpening filtering and other enhancement technologies, which are used to improve the clarity and visualization of the image and provide better image data input for subsequent feature extraction. Then comes image normalization, which is to uniformly map the image data to a fixed numerical distribution range, such as [0,1] or [-1,1]. This step can reduce the difference between image data and improve the adaptability of subsequent models to images.
[0028] The last step is image format conversion, which converts the input image into a tensor format recognizable by the neural network model, such as the common NCHW (batch × channel × height × width), etc., to meet the input requirements of the deep learning framework.
[0029] Based on the above embodiment, as an optional embodiment, the image preprocessing module includes: an image cropping module, an image enhancement module, an image normalization module and an image format conversion module; An image cropping module, used for cropping the factory image according to preset cropping parameters to obtain a first factory image; Specifically, at the beginning of the entire image preprocessing process, the image cropping module first performs a cropping operation on the input original factory image to obtain a first factory image.
[0030] The purpose of image cropping is to remove redundant background areas in the input image that are irrelevant to the target detection task and retain only the target area of interest, thereby simplifying the subsequent feature extraction and model inference process and improving computational efficiency.
[0031] Specifically, the image cropping module will determine the range of the area to be cropped based on pre-set cropping parameters. These cropping parameters can be fixed coordinate values or adaptive region selection strategies, such as automatically determining the cropping area based on the content semantics of the image.
[0032] After determining the cropping area, the module will crop the pixel matrix of the original factory image according to the specified area range, retaining only the pixel data of the target area and filtering out the rest of the background area, thereby obtaining the first cropped factory image.
[0033] It is worth mentioning that the image cropping module also adopts a strategy to prevent the target object from being over-cropped during the cropping process, that is, a certain amount of pixel redundancy will be reserved near the edge of the target object to prevent the target from being cropped and affecting the subsequent detection accuracy.
[0034] After image cropping, the output first-plant image will only retain the valid areas closely related to the target detection task, and a large number of irrelevant background areas will be removed, thereby reducing the processing size of the image and reducing the computational complexity of subsequent feature extraction and target detection.
[0035] An image enhancement module, used for performing image enhancement processing on the first factory image to obtain a second factory image; Specifically, after the image cropping module outputs the first factory image, the image enhancement module immediately performs a series of image enhancement processes on it to obtain an enhanced second factory image.
[0036] The purpose of image enhancement is to generate more diverse training data by transforming and distorting the original image, enhancing the robustness of the model to target objects under different viewing angles, lighting, deformation and other conditions, thereby improving the generalization ability of the model and avoiding overfitting.
[0037] Specifically, the image enhancement module first determines which enhancement operations and their parameters need to be performed on the input image according to the preset enhancement strategy. Common image enhancement operations include rotation, flipping, scaling, cropping, translation, Gaussian noise addition, brightness adjustment, contrast adjustment, etc. Taking rotation enhancement as an example, the module will randomly select a rotation angle and rotate the first factory image according to the angle to generate a new rotated image as the enhanced output. Through rotation enhancement, the model can improve the ability to recognize different orientations of the target object.
[0038] In addition to rotation, the image enhancement module will also perform other enhancement operations in sequence, such as flipping, scaling, cropping, etc., and accumulate the results after each operation as new enhancement input, and finally generate rich and diverse second-plant images.
[0039] An image normalization module, used for normalizing the second factory image to obtain a third factory image; Specifically, the purpose of image normalization is to uniformly map the pixel values of the input image to the expected input range of the neural network model, eliminating the differences caused by different image shooting conditions (such as lighting, exposure, etc.), thereby improving the stability and generalization performance of feature extraction and target detection.
[0040] Specifically, the image normalization module first determines the expected input range of the neural network model, which is usually a floating point interval between [0, 1] or [-1, 1]. Then, the module linearly scales each pixel value of the input binary image to map it to the expected input range of the model.
[0041] Common normalization methods include minimum-maximum normalization and mean-variance normalization. Minimum-maximum normalization maps the minimum pixel value of the image to the minimum value of the target interval, and the maximum pixel value to the maximum value of the target interval, and other pixel values are linearly scaled proportionally. Mean-variance normalization first calculates the mean and standard deviation of all pixels in the image, then subtracts the mean from each pixel value and divides it by the standard deviation, thereby mapping the data to the expected standard interval with a mean of 0 and a variance of 1.
[0042] It is worth mentioning that the specific normalization method, mapping range and other parameters used in the image normalization module have undergone a lot of experiments and adjustments to ensure the best quality of the normalized image and provide high-quality input data for subsequent feature extraction and target detection.
[0043] After image normalization, the pixel value range of the output third-plant image will be uniformly within the expected input range of the neural network. At the same time, the influence of different shooting conditions is eliminated to the greatest extent, ensuring the consistency and stability of different input images in the feature extraction and target detection stages.
[0044] The image format conversion module is used to convert the format of the third factory image according to the preset format data to obtain a pre-processed engineering image.
[0045] Specifically, the purpose of image format conversion is to convert the preprocessed image data into the input format expected by the deep learning model, ensuring that the model can correctly read and parse the image data, laying the foundation for subsequent feature extraction and target detection.
[0046] Specifically, the image format conversion module first obtains the neural network model's requirements for the input image format, which usually includes data dimension, channel arrangement order, data type, etc. Different model frameworks may have different requirements for image formats.
[0047] Taking the PyTorch framework as an example, its expected input image format is a 4-dimensional tensor (batch_size, channels, height, width), the channel arrangement order is RGB, and the data type is a 32-bit floating point number.
[0048] Therefore, for the input third-factor image, the module will first check its current format, and if it does not meet the requirements, it needs to perform the corresponding format conversion. For example, if the input is a 3-dimensional numpy array of HWC (height, width, channel), the module will rearrange the data dimensions in the order of NCHW and convert the data type to a 32-bit floating point number to obtain a formatted tensor that meets the requirements of PyTorch.
[0049] In addition to adjusting data dimensions and types, the image format conversion module can also perform other necessary format conversion operations, such as channel order conversion (RGB to BGR or vice versa), encoding and decoding (such as PNG to JPG), etc., to adapt to the specific needs of different frameworks and models.
[0050] The feature extraction module 22 is used to extract features from the processed engineering image to obtain a target feature image; Specifically, after the image preprocessing module optimizes the original factory image, the preprocessed high-quality factory image data will be input into the feature extraction module. The feature extraction module is the core of the entire target detection process, responsible for extracting the key feature information of the target object from the input image and generating the target feature representation, laying the foundation for subsequent target detection.
[0051] The feature extraction module integrates innovative technologies such as the multi-scale visual converter module, the convolutional neural network feature extraction module, the visual converter-convolutional neural network adaptive fusion module and the multi-scale feature fusion module.
[0052] First, the multi-scale visual converter module uses the self-attention mechanism to efficiently learn the global semantics and long-range dependency features of images. Unlike the traditional CNN model based on local convolution operations, the visual converter directly models the relationship between image pixels through the self-attention mechanism, and is good at capturing the overall semantics and long-range contextual information of the image. At the same time, the module adopts the design of multi-head attention and hierarchical encoder to further enhance its ability to represent multi-scale features.
[0053] At the same time, the CNN feature extraction module is responsible for extracting local details and texture features of the image. CNN can efficiently obtain local low-level features such as edges, textures, and object parts of the image by means of local convolution and pooling operations. The visual converter-convolutional neural network adaptive fusion module adaptively fuses the features of the above two paradigms, allowing the system to simultaneously obtain rich semantics, details, and location information, thereby comprehensively characterizing the feature representation of the target object at different scales.
[0054] Finally, the multi-scale feature fusion module fuses the features from the visual converter and CNN modules at different scales to generate a high-quality target feature image. This feature image contains rich feature information of the target object at each scale and can provide strong feature support for the subsequent target detection head module.
[0055] Based on the above embodiment, as an optional embodiment, the feature extraction module includes: a multi-scale ViT module, a CNN feature extraction module, a ViT-CNN adaptive fusion module and a multi-scale feature fusion module; The multi-scale ViT module is used to extract features from the preprocessed engineering images and obtain ViT feature maps; Specifically, within the feature extraction module, the multi-scale ViT module is the first key module for image feature extraction. The module uses the self-attention mechanism to efficiently learn and extract the global semantics and long-range dependency features of the target object from the preprocessed factory image to obtain the ViT feature map.
[0056] Different from the traditional CNN model based on convolution operations, the ViT (Visual Transformer) model directly models the relationship between image pixels through the self-attention mechanism, and is good at capturing the overall semantics and long-range contextual information of the image.
[0057] Specifically, the multi-scale ViT module first divides the input pre-processed factory image into small patches, adds position encoding to each patch, and maps it to the embedding space. Then, through the multi-head self-attention mechanism, the module allows each patch to pay attention to the feature information of all other patches, thereby effectively modeling the long-range dependencies between patches.
[0058] Next, the encoder layer further processes and propagates the embedding features, and finally generates encoding vectors corresponding to each patch. These encoding vectors retain the global semantics and contextual information of the image and constitute the preliminary ViT feature map.
[0059] It is worth mentioning that the multi-scale ViT module adopts innovative designs such as multi-head attention and hierarchical encoders, which greatly improves its ability to represent multi-scale features. The module focuses on patches of different scales through the attention mechanism, learns and fuses multi-scale features at different encoder layers, and ensures that the generated ViT feature map can fully characterize the semantics and dependencies of the target object at all scales.
[0060] The generated ViT feature map not only contains rich global semantic features, but also retains the feature information of the target object at different scales, laying a solid foundation for subsequent feature fusion. The ViT feature map will then be adaptively fused with the CNN feature map to give full play to the feature representation advantages of the two different paradigms and obtain a richer and more comprehensive target feature description.
[0061] CNN feature extraction module, used to extract features from preprocessed engineering images to obtain CNN feature maps; Specifically, the function of the CNN feature extraction module is to extract local details and texture features of the image. With its unique local convolution and pooling mechanism, CNN is good at obtaining local low-level visual features such as edges, textures, and object parts of images.
[0062] Specifically, the CNN feature extraction module first inputs the preprocessed factory image into the convolution layer. The convolution layer extracts features from each local area of the image by sliding the convolution kernel on the image to generate a preliminary feature map. Next, the pooling layer downsamples the feature map to retain the main components of the features and resist some noise and deformation. The module then repeats the above process in multiple convolutional and pooling layers, each of which extracts features of different scales and levels of abstraction to form a hierarchical feature pyramid. High-level layers capture larger-scale semantic features, while low-level layers retain more local detail information.
[0063] Finally, the CNN feature extraction module integrates these hierarchical feature maps to construct a CNN feature map, which finely depicts the texture, edge, contour and other detailed features of the target object in the image, providing extremely valuable local visual information supplement for subsequent feature fusion.
[0064] It should be noted that the CNN feature extraction module adopts the currently recognized best and most advanced CNN network architecture, such as ResNet, InceptionNet, etc., and performs network modifications and parameter tuning according to the special needs of factory scenarios, thereby ensuring the accuracy and robustness of feature extraction.
[0065] The ViT-CNN adaptive fusion module is used to adaptively fuse the ViT feature map and the CNN feature map to obtain the ViT-CNN fused feature map; Specifically, after the multi-scale ViT module and the CNN feature extraction module generate the ViT feature map and the CNN feature map respectively, the feature information of these two different paradigms will be input into the ViT-CNN adaptive fusion module for fusion processing.
[0066] The function of the ViT-CNN adaptive fusion module is to adaptively fuse the ViT feature map and CNN feature map transmitted from upstream to generate a ViT-CNN fused feature map containing comprehensive feature information, laying the foundation for subsequent target detection.
[0067] The module first uses the channel attention mechanism to perform inter-channel feature weighting on the ViT feature map and the CNN feature map, and adaptively learns the importance weights of different channels for the target detection task. Through this adaptive channel weighting method, the model can automatically find the importance of each feature on different channels, and then assign reasonable feature weights. Next, the ViT-CNN adaptive fusion module uses the attention mechanism to allow the feature vector of each feature map to pay attention to the information of all other feature vectors in another feature map. Specifically, the module encodes the ViT feature map and the CNN feature map into a query vector and a key-value vector, respectively, and captures the correlation between the two feature maps through the attention mechanism, thereby realizing the adaptive fusion of the two feature representations.
[0068] During the adaptive fusion process, this module can flexibly adjust the fusion weights of ViT and CNN features, and automatically allocate a reasonable fusion ratio according to different scenarios, so that the fused feature map can contain both global semantics and local detail information, giving full play to the advantages of the two paradigm feature representations.
[0069] Finally, the ViT-CNN adaptive fusion module outputs a high-quality ViT-CNN fusion feature map. This feature map not only retains the long-range dependencies and semantic features in the ViT feature map, but also contains rich local texture and detail information in the CNN feature map, realizing the organic fusion of two different feature representations and ensuring the integrity and richness of feature information.
[0070] The multi-scale feature fusion module is used to perform multi-scale fusion of the ViT feature map, the CNN feature map and the ViT-CNN fusion feature map to obtain the target feature image.
[0071] Specifically, after obtaining three different types of feature representations, namely ViT feature map, CNN feature map and ViT-CNN fusion feature map, they need to be comprehensively fused through the multi-scale feature fusion module to generate the final target feature image to provide high-quality feature support for the subsequent target detection head.
[0072] The function of the multi-scale feature fusion module is to fuse the multi-scale features from the ViT, CNN and ViT-CNN fusion modules at different levels to ensure that the rich feature information of the target object at each scale can be fully utilized to obtain the final all-encompassing target feature image.
[0073] Specifically, the module first upsamples and downsamples the ViT feature map, CNN feature map, and ViT-CNN fusion feature map, respectively, and maps the three to the same feature scale space. In this way, feature information at different scales can be fused at the same feature level. Next, the multi-scale feature fusion module uses an advanced feature pyramid network structure to adaptively weighted fuse the same-scale features obtained in the previous step at different levels. Low-level fusion focuses on local details and texture information, mid-level fusion takes into account the structural information of the target, and high-level fusion focuses on semantic and contextual features.
[0074] While fusing, the module also introduces an attention mechanism, so that the features at each scale can adaptively learn feature information from other scales, thereby achieving cross-scale feature interaction and information flow, further enhancing the ability of feature expression.
[0075] Finally, after the above cross-module and cross-scale adaptive fusion, the multi-scale feature fusion module outputs a high-quality, multi-dimensional target feature image. This feature image comprehensively depicts the visual features of the target object at all scales, including rich information at multiple levels such as semantics, details, structure and context, ensuring a comprehensive description of the target and laying a solid foundation for subsequent high-precision target detection.
[0076] The target detection head module 23 is used to input the feature image into the target detection head to obtain the detection result; Specifically, the target detection head first divides the input target feature image into a large number of small areas, and performs target classification and bounding box regression on each small area. In the target classification stage, the classifier is used to determine whether there is a target object in each small area, and if so, the category confidence score of the object is given; in the bounding box regression stage, the regressor is used to predict the specific position and size of the target object in the small area to obtain a preliminary detection frame.
[0077] After the above two key steps, the module can generate a large number of preliminary detection results, each of which contains rich information such as the category of the predicted target, confidence score, detection box position and size.
[0078] It is worth mentioning that the specific algorithm implementation inside the target detection head module adopts the most cutting-edge technologies in the field of deep learning and machine learning, and a large number of model optimizations and parameter tunings are carried out for the application scenarios of this system to ensure the best balance between detection accuracy and speed. At the same time, the module also integrates some innovative improved algorithms, such as multi-scale feature pyramids and adaptive anchor box selection, which further improve the detection effect.
[0079] The post-processing module 24 is used to post-process the detection result to obtain an object recognition result.
[0080] Specifically, after being processed by the target detection head module, the system obtains preliminary detection results. However, due to the complexity of the target detection process and the influence of noise, these preliminary results often contain some redundant, overlapping and low-confidence detection frames, which require further post-processing optimization to output accurate and reliable final object recognition results. The post-processing module undertakes this key task, which includes multiple processing steps such as non-maximum suppression, confidence threshold filtering and detection frame decoding.
[0081] The first is non-maximum suppression. Since the target detection head will detect each small area, multiple overlapping detection frames may be generated for the same target object. The role of non-maximum suppression is to select the one with the highest confidence from these overlapping detection frames as the final result, and suppress the remaining redundant detection frames with lower confidence. The specific method is to calculate the degree of overlap between each pair of detection frames according to the preset non-maximum suppression threshold. If the overlapping area exceeds the threshold, the detection frame with higher confidence is retained and the frame with lower confidence is deleted. The next step is confidence threshold filtering. The purpose of this step is to further eliminate those detection frames with too low confidence scores and low credibility, and only retain detection results with higher confidence and stronger reliability. According to the pre-set confidence threshold, the detection frames with scores lower than the threshold are directly filtered out to ensure the reliability of the output results.
[0082] The last step is to decode the detection frame. Since the original detection frame coordinates output by the neural network are usually relative coordinates based on the feature image, they need to be decoded and restored to the absolute pixel coordinates on the original input image. At the same time, the frame coordinates are scaled accordingly according to the downsampling rate of the feature pyramid to match the original scale. The decoded detection frame is the final, accurately located target object position and size.
[0083] After the above three key steps, the post-processing module will output optimized, streamlined, high-quality object recognition results. Each result contains key information such as the category, confidence, and precise bounding box coordinates of the target object, and the accuracy and reliability of the information have been greatly improved.
[0084] Based on the above embodiment, as an optional embodiment, the display module includes: a result rendering module and a 3D space projection module; A result rendering module, used to generate 2D rendering content based on the object recognition results; Specifically, after the target detection and object recognition module outputs the final object recognition result, the result rendering module of the display module will generate 2D rendering content according to the result.
[0085] The purpose of generating 2D rendering content is to present the results of target detection and recognition on images or video screens in an intuitive and clear manner, so that human users can quickly understand and verify the detection results and improve the interpretability and interactivity of the system.
[0086] Specifically, the result rendering module first extracts key data such as the category information, detection box coordinates, and confidence of each detection target from the object recognition results. Then, the module will render these detection results on the original image or video frame according to the preset rendering strategy.
[0087] Common rendering strategies include: drawing a bounding box around the detection box and marking the target category and confidence level in the box; using different colors to identify targets of different categories; using dedicated icons for specific categories, etc. The rendering strategy can be flexibly configured according to actual needs.
[0088] In addition to rendering detection results, this module can also support other enhanced rendering effects, such as rendering the detection results of multiple consecutive frames as animation tracks, adding ruler auxiliary lines, displaying detection time, etc., to enhance the visualization effect and information content of the results.
[0089] It is worth mentioning that when the rendering module generates rendering content, it will automatically adjust the rendering strategy according to the format of the input data. For example, static rendering is used when the input is an image, and dynamic rendering is enabled when the input is a video stream, ensuring that high-quality rendering effects can be generated in any case.
[0090] After rendering, the output 2D rendering content will be superimposed on the original image or video frame in the form of an overlay, so that the original picture and the detection result are perfectly integrated, clearly and intuitively showing the entire process and effect of target detection and recognition.
[0091] The 3D space projection module is used to project the 2D rendering content into an AR image and project the AR image onto the display screen of the AR device.
[0092] Specifically, after the result rendering module generates 2D rendering content, the 3D space projection module of the display module will project and convert the 2D rendering content into an AR image, and project the AR image onto the display screen of the connected AR device.
[0093] The purpose of performing 3D space projection is to upgrade the traditional 2D plane rendering content to an augmented reality (AR) visual experience, so that the detection results are no longer limited to the screen image, but can be seamlessly integrated with the real three-dimensional environment, bringing users an immersive visual experience.
[0094] Specifically, the 3D space projection module first needs to obtain the position and posture information of the current AR device, including position coordinates, viewing direction, etc., which are usually provided in real time by various sensors built into the AR device (such as depth cameras, IMUs, etc.). Next, the module will calculate the projection position and angle of the 2D rendering content in the virtual 3D space based on the position and posture of the AR device and the pre-input or calibrated 3D model of the environment.
[0095] It is worth mentioning that the module also considers various geometric factors and optical properties when performing projection calculations to ensure that the generated AR effect is as close to reality as possible. At the same time, it also enhances the realism and immersion of AR rendering by introducing strategies such as lighting models.
[0096] Finally, the module transmits the transformed AR image to the display screen of the AR device in real time through wireless or wired connection for presentation. It should be noted that what is displayed is no longer a simple 2D image, but a three-dimensional augmented reality effect that is integrated with the real environment.
[0097] Based on the above embodiment, as an optional embodiment, the post-processing module includes: a non-maximum suppression module, a confidence threshold filtering module and a detection frame decoding module; A non-maximum suppression module is used to remove highly overlapping redundant detection frames from the detection results to obtain a first detection result; Specifically, after the target detection head module outputs the original detection result, the first step of the detection post-processing module is to remove highly overlapping redundant detection frames through a non-maximum suppression module to obtain a first detection result.
[0098] The purpose of performing non-maximum suppression is to solve the problem that the target detector generates multiple highly overlapping detection frames on the same target object, ensuring that the final output detection result contains only one optimal detection frame for each target object, thereby improving the accuracy and clarity of the detection.
[0099] Specifically, the non-maximum suppression module first sorts all detection results according to the confidence score of the detection frame, and the detection frame with higher confidence is ranked higher. Then, the module traverses the sorted detection frames one by one, and for the currently traversed detection frame, calculates its overlap with all other detection frames.
[0100] The commonly used method to calculate the degree of overlap is to calculate the intersection over union (IoU) between two detection boxes, that is, the ratio of the intersection area of the two boxes to the union area. If the IoU value of the current detection box and any other detection box exceeds a preset threshold (usually 0.5-0.7), the two detection boxes are considered to have a high degree of overlap, and the one with a lower confidence will be eliminated.
[0101] After the above traversal and screening, the first detection result finally output by the non-maximum suppression module only retains the detection box with the highest confidence for each target object, and removes all other highly overlapping redundant detection boxes, thereby greatly reducing the risk of false detection and improving the accuracy of detection.
[0102] A confidence threshold filtering module, used to remove the detection frame in the first detection result according to a preset confidence threshold to obtain a second detection result; Specifically, after the non-maximum suppression module outputs the first detection result, the confidence threshold filtering module immediately performs confidence threshold filtering on it to obtain the second detection result.
[0103] The purpose of confidence threshold filtering is to further improve the accuracy of the detection results. By setting the confidence threshold, detection boxes with low confidence and possible false positives are filtered out, thereby reducing the false alarm rate and improving the reliability of the final detection system.
[0104] Specifically, the confidence threshold filtering module will first obtain a preset confidence threshold, which is usually obtained through a large number of experimental tuning based on the actual application scenario and the tolerance for false alarm rate. In general, the higher the threshold, the lower the false alarm rate, but it may also increase the risk of missed detection. Next, the module will traverse each detection box in the first detection result and compare its confidence score with the preset threshold. For detection boxes with confidence lower than the threshold, the module will directly filter them out; for detection boxes with confidence higher than or equal to the threshold, they will be retained and output as part of the second detection result.
[0105] It is worth mentioning that when performing confidence threshold filtering, the module adopts an efficient memory management strategy to avoid unnecessary data copying and memory allocation, thereby maximizing processing efficiency.
[0106] After filtering by the confidence threshold, only the detection boxes with higher confidence and reliability are retained in the output second detection result, effectively removing the low-confidence detection results with a greater risk of false detection, thereby significantly reducing the false alarm rate of the entire target detection system.
[0107] The detection frame decoding module is used to decode the second detection result to obtain a third detection result.
[0108] Specifically, after the confidence threshold filtering module outputs the second detection result, the detection frame decoding module decodes it to obtain a final third detection result.
[0109] The purpose of performing detection frame decoding is to convert the encoded detection frame output by the neural network model into pixel-level coordinates in the actual image coordinate system to facilitate subsequent visualization and application processing.
[0110] Specifically, the detection frame decoding module first obtains the detection frame encoding method used by the neural network model during training, such as anchor frame encoding, default frame encoding, etc. Different encoding methods correspond to different decoding logics.
[0111] Taking anchor frame encoding as an example, this encoding method usually represents the detection frame as an offset relative to a set of pre-set anchor frames, including center coordinate offset, aspect ratio offset, etc. During decoding, the module will calculate the actual pixel coordinates of each detection frame in the original image coordinate system based on these input offsets, combined with the position and size of the corresponding anchor frame, according to the pre-defined decoding formula.
[0112] In addition to decoding the coordinates, the module also needs to decode other attributes of the detection box, such as the detection confidence score.
[0113] It is worth mentioning that the detection box decoding formula corresponds to the encoding formula and contains some hyperparameters, such as anchor box size, scaling factor, etc. These parameters need to be kept exactly the same as during training, otherwise the decoding results will be biased.
[0114] After decoding, each detection box in the output third detection result will be converted into absolute coordinates and other original attribute values at the image pixel level, so that the detection results can be directly visualized on the original image and provide directly usable input for subsequent high-level applications such as target recognition and tracking.
[0115] Detection box decoding is a key step in connecting the internal representation of the deep learning model with the actual application scenario. By converting the encoded detection box into intuitive and easy-to-understand image coordinates, it not only improves the interpretability of the detection results, but also lays the foundation for the smooth progress of the subsequent processing flow.
[0116] The display module 3 is used to perform AR rendering on the factory image according to the object recognition result to obtain an AR image, and present the AR image on the display screen of the AR device.
[0117] Specifically, after being processed by the target detection module, the system obtains accurate object recognition results. Next, these object recognition results will be input into the display module for AR rendering, and finally the rendered AR image will be presented on the display screen of the AR device.
[0118] Based on the above embodiment, as an optional embodiment, the system further includes: an eye tracking module; The eye tracking module is used to obtain eye state data and determine the attention area through the eye state data.
[0119] Specifically, the purpose of introducing the eye tracking function is to achieve accurate attention detection, so as to better grasp the user's line of sight, optimize system resource allocation, and enhance the user's interactive experience.
[0120] Specifically, the eye tracking module first needs to obtain the user's eye image or video stream data in real time through specialized eye tracker hardware or image analysis algorithms. The module will analyze and process these raw data to extract key state parameters such as the pupil position, opening degree, and line of sight direction.
[0121] After acquiring the eye state data, the eye tracking module will combine the pre-established eye movement model and mapping relationship to calculate the screen coordinates or three-dimensional spatial position of the user's current line of sight, that is, the user's gaze point. Furthermore, the module will determine the user's current focus area based on the gaze point and its movement trajectory. It is worth mentioning that when determining the attention area, the module will also consider other contextual information, such as the distance between the user and the screen, the duration of gaze, etc., to improve the sensitivity and accuracy of attention detection.
[0122] Based on the above embodiment, as an optional embodiment, the system further includes: an attention area rendering module; The attention area rendering module is used to enhance the rendering of the AR image according to the object recognition result and the attention area to obtain the attention AR image.
[0123] Specifically, based on the attention area information provided by the eye tracking module, the system further integrates an attention area rendering module to perform attention-enhanced rendering on the AR image to obtain an attention AR image.
[0124] The purpose of introducing attention area rendering is to fully tap the potential of eye tracking technology and further improve the sophistication of AR visual content and user experience quality by enhancing the detail-level rendering of the user's gaze area.
[0125] Specifically, the attention area rendering module first obtains the real-time attention area coordinates from the eye tracking module, which corresponds to the screen or three-dimensional space range where the user's current line of sight is focused. Next, the module extracts all target object information within the attention area from the object recognition results, including object category, location, size, etc.
[0126] Based on the above embodiment, as an optional embodiment, the attention area rendering module includes: a rendering strategy generation module and a region layered rendering module; A rendering strategy generation module is used to generate a rendering enhancement strategy based on the attention area, object recognition results and a preset rendering strategy library; Specifically, within the attention area rendering module, there are two sub-modules: the rendering strategy generation module and the area layered rendering module. The former is responsible for generating rendering enhancement strategies, and the latter performs specific layered rendering operations.
[0127] The rendering strategy generation module is responsible for comprehensively generating a rendering enhancement strategy that is suitable for the current scene based on the real-time attention area information from the eye tracking module, the target object attributes in the object recognition results, and the pre-configured rendering strategy library.
[0128] The purpose of this step is to dynamically formulate personalized rendering schemes to ensure tailored rendering enhancement effects for different gaze scenarios and fully meet the diverse needs of users.
[0129] Specifically, the rendering strategy generation module first analyzes the position, size, and shape of the current attention area, and extracts key attributes such as the type, quantity, and position of the target objects in the area from the object recognition results. Next, the module will generate the most optimized rendering enhancement strategy for the current scene based on the configuration rules in the preset rendering strategy library. The rendering strategy library pre-stores a large number of enhancement rules that have been optimized manually or by algorithms, covering the best rendering templates under different object types and attention area characteristics.
[0130] The regional layered rendering module is used to enhance the rendering of the AR image according to the rendering enhancement strategy to obtain the attention AR image.
[0131] Specifically, after the rendering strategy generation module outputs the rendering enhancement strategy, the regional layered rendering module will perform targeted layered enhancement rendering on the original AR image according to the strategy, and finally generate an attention AR image.
[0132] The purpose of introducing regional layered rendering is to give full play to the parallel processing advantages of computer graphics. By layering the elements inside and outside the attention area, a dynamic balance between the visual quality and computing power consumption of the attention object and the background environment can be achieved, bringing users an unparalleled AR visual experience.
[0133] Specifically, the regional layered rendering module will first divide the original AR image into three different rendering levels: attention object layer, attention object surround layer and background layer according to the rendering strategy.
[0134] For the attention object layer, this module will maximize its rendering quality, including adding multi-sampling anti-aliasing, detail level loading, physical lighting simulation and other enhancements to ensure that the texture, edges, details, etc. of the object can be presented in high resolution and realistically.
[0135] At the same time, in order to avoid the visual separation between the object and the environment, the module will also perform moderate rendering enhancement on the surround layer of the attention object to ensure a natural transition in quality between the object and the surrounding environment.
[0136] See also Figure 3 , Figure 3 A schematic flow chart of a factory object recognition method based on AR provided in an embodiment of the present application, the factory object recognition method based on AR includes: S101, obtaining a factory image in front of the AR device through a camera of the AR device; S102, performing object recognition on the factory image to obtain an object recognition result; S103: presenting the object recognition result in the form of AR on a display screen of the AR device.
[0137] The above are only exemplary embodiments of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification and the truth of practice, those skilled in the art will easily think of other embodiments of the present disclosure.
[0138] This application is intended to cover any variation, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the art not described in the present disclosure. The description and examples are to be regarded as exemplary only, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A factory object recognition system based on AR, characterized in that: The system comprises: an image acquisition module, a target detection module and a display module; The image acquisition module is used to acquire the factory image in front of the AR device through the camera of the AR device; The object detection module is used to perform object recognition on the factory image to obtain an object recognition result; The display module is used to perform AR rendering on the factory image according to the object recognition result to obtain an AR image, and present the AR image on a display screen of an AR device.
2. The system according to claim 1, characterized in that The target detection module includes: an image preprocessing module, a feature extraction module, a target detection head module and a post-processing module; The image preprocessing module is used to preprocess the factory image to obtain a preprocessed engineering image; The feature extraction module is used to extract features from the processed engineering image to obtain a target feature image; The target detection head module is used to input the feature image into the target detection head to obtain a detection result; The post-processing module is used to post-process the detection result to obtain an object recognition result.
3. The system according to claim 2, characterized in that The feature extraction module includes: a multi-scale ViT module, a CNN feature extraction module, a ViT-CNN adaptive fusion module and a multi-scale feature fusion module; The multi-scale ViT module is used to extract features from the preprocessed engineering image to obtain a ViT feature map; The CNN feature extraction module is used to extract features from the preprocessed engineering image to obtain a CNN feature map; The ViT-CNN adaptive fusion module is used to adaptively fuse the ViT feature map and the CNN feature map to obtain a ViT-CNN fused feature map; The multi-scale feature fusion module is used to perform multi-scale fusion on the ViT feature map, the CNN feature map and the ViT-CNN fusion feature map to obtain the target feature image.
4. The system according to claim 2, characterized in that The image preprocessing module includes: an image cropping module, an image enhancement module, an image normalization module and an image format conversion module; The image cropping module is used to crop the factory image according to preset cropping parameters to obtain a first factory image; The image enhancement module is used to perform image enhancement processing on the first factory image to obtain a second factory image; The image normalization module is used to perform normalization processing on the second factory image to obtain a third factory image; The image format conversion module is used to convert the format of the third factory image according to preset format data to obtain the pre-processed engineering image.
5. The system according to claim 2, characterized in that The post-processing module includes: a non-maximum suppression module, a confidence threshold filtering module and a detection frame decoding module; The non-maximum suppression module is used to remove highly overlapping redundant detection frames from the detection result to obtain a first detection result; The confidence threshold filtering module is used to remove the detection frame in the first detection result according to a preset confidence threshold to obtain a second detection result; The detection frame decoding module is used to decode the second detection result to obtain a third detection result.
6. The system according to claim 1, characterized in that The display module includes: a result rendering module and a 3D space projection module; The result rendering module is used to generate 2D rendering content according to the object recognition result; The 3D space projection module is used to project the 2D rendering content into an AR image and project the AR image onto a display screen of an AR device.
7. The system according to claim 1, characterized in that The system further includes: an eye tracking module; The eye tracking module is used to obtain eye state data and determine the attention area through the eye state data.
8. The system according to claim 1, characterized in that The system further includes: an attention area rendering module; The attention area rendering module is used to enhance the AR image according to the object recognition result and the attention area to obtain an attention AR image.
9. The system according to claim 8, characterized in that The attention area rendering module includes: a rendering strategy generation module and a regional layered rendering module; The rendering strategy generation module is used to generate a rendering enhancement strategy according to the attention area, the object recognition result and a preset rendering strategy library; The regional layered rendering module is used to perform enhanced rendering on the AR image according to a rendering enhancement strategy to obtain the attention AR image.
10. A factory object recognition method based on AR, characterized in that: The method comprises: Acquire the factory image in front of the AR device through the camera of the AR device; Performing object recognition on the factory image to obtain an object recognition result; The object recognition results are presented in the form of AR on the display screen of the AR device.
Citation Information
Cited By
Lightweight eye movement tracking method based on half-eye image reconstruction
CN120635497A