Intelligent operation inspection information identification method and system based on acquisition field
Through the intelligent operation and inspection information identification method based on the acquisition site, the image data processing is performed using a full convolutional structure and regional proposal network, and the precise positioning and classification of target items is achieved, the problem of inefficient traditional manual operation and inspection is solved, and the accuracy and automation of identification are improved.
Patent Information
- Application Number
- CN202510358253.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-22
AI Technical Summary
Traditional artificial operation inspection information identification methods are inefficient and susceptible to human factors, making it difficult to meet the needs of modern society for efficient and fast services.
Using an intelligent operation and inspection information recognition method based on the acquisition site, the image data is collected regularly, and the edge-complement and mean-reduction processing is performed, the feature map is extracted using a full convolutional structure, combined with the regional proposal network and bounding box regression, to achieve accurate positioning and classification of the goals, and finally the labeled image data is evaluated through confidence.
It realizes the full process of automated processing from image data acquisition to target item identification and labeling, improves the efficiency and accuracy of operation and inspection work, adapts to environmental changes, reduces manual intervention, and enhances the accuracy and reliability of detection.
Smart Images

Figure CN120355891A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic identification, and particularly to an intelligent operation and inspection information identification method and system based on a collection field. Background Art
[0002] Traditional operation and inspection information identification methods have long mainly relied on manual inspections. This method requires staff to carefully observe and judge each item to determine whether it meets the specified standards or whether it carries potential risks. However, this manual inspection method has obvious drawbacks.
[0003] Firstly, the efficiency of manual inspection is relatively low. In large-scale, high-traffic transportation, logistics, or security inspection scenarios, staff need to process a large number of items, and each item inspection takes a certain amount of time. This results in a long overall inspection process and cannot meet the modern society's demand for efficient and fast services.
[0004] Secondly, manual inspection is easily affected by human factors. The experience, skill level, concentration, and fatigue status of staff will all directly affect the inspection results. For example, inexperienced staff may not be able to accurately identify some complex or hidden target items, while over-fatigued staff may miss inspections or make misjudgments due to inattentiveness. The existence of these factors greatly reduces the accuracy and reliability of the identification results.
[0005] Therefore, some traditional operation and inspection information identification methods that rely on manual labor have become difficult to meet the needs of modern society. To improve operation and inspection efficiency and ensure the accuracy of identification results, it is urgent to explore and apply new automated operation and inspection information identification methods based on advanced technologies. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide an intelligent operation and inspection information identification method and system based on a collection field, which realizes the full-process automated processing from image data acquisition to target item identification and marking, and improves the efficiency and accuracy of operation and inspection work.
[0007] To solve the above technical problem, the technical solution of the present invention is as follows:
[0008] In the first aspect, an intelligent operation and inspection information identification method based on a collection field, the method includes:
[0009] Regularly collect image data of target items, and construct an intelligent operation and inspection information collection field to obtain image data of the security inspection channel;
[0010] Perform border padding and mean subtraction on the image data to obtain processed image data;
[0011] Pass the processed image data through a fully convolutional structure to extract the feature map of the image;
[0012] Pass the feature map through a region proposal network to generate the feature map of the region of interest containing the target, and generate a feature map of a unified size after pooling transformation of the feature map of the region of interest;
[0013] Locate the feature map of the unified size and calculate the offset of the bounding box at the same time to obtain the accurately located target;
[0014] Refine the region of interest of the feature map according to the accurately located target, and classify and evaluate the confidence of the refined region of interest to obtain the target detection result and confidence;
[0015] Mark the image of the image data containing the target item according to the target detection result and confidence to realize the recognition of operation and inspection information.
[0016] Preferably, perform border padding and mean subtraction on the image data to obtain the processed image data, including:
[0017] Determine the type and size of the image border padding;
[0018] Add additional pixel values around the boundary of the image according to the type and size of the image border padding to obtain the image after border padding processing;
[0019] Pass the image after border padding processing through Perform mean subtraction processing to obtain the processed image data, where P ijk represents the pixel value of the i-th row, j-th column, and k-th channel in the original image data, and P i ′ jk represents the pixel value of the i-th row, j-th column, and k-th channel in the image data after mean subtraction processing, M and N respectively represent the height and width of the image (i.e., the number of rows and columns), and i, j, and k represent the row number, column number, and channel index.
[0020] Preferably, pass the processed image data through a fully convolutional structure to extract the feature map of the image, including:
[0021] Construct a fully convolutional structure by combining multiple convolutional layers;
[0022] Perform forward propagation on the processed image data through the fully convolutional structure to extract a higher-level feature representation;
[0023] Obtain the feature map of the image according to the higher-level feature representation.
[0024] Preferably, the feature map is passed through a region proposal network to generate a feature map of regions of interest containing the target, and the feature map of the regions of interest is pooled to generate a feature map of a unified size, including:
[0025] In the region proposal network, a series of anchor points are generated on the feature map by means of a sliding window, and each anchor point corresponds to a region on the original image;
[0026] The scales and ratios of the regions of interest are preset, and according to the scales and ratios of the regions of interest, the region proposal network generates multiple candidate regions for each anchor point;
[0027] Two branches of classification and regression are used to predict whether the candidate region contains the target and the precise location of the target, so as to obtain the feature map of the candidate region;
[0028] The feature map of the candidate region is pooled to generate a feature map of a unified size.
[0029] Preferably, the feature map of the unified size is located, and at the same time, the offset of the bounding box is calculated to accurately locate the target, including:
[0030] According to the feature map of the unified size, each region of interest is obtained;
[0031] Four coordinate parameters of each region of interest are predicted by a bounding box regressor to obtain the predicted bounding box of each region of interest;
[0032] By Calculating the relative offset on the x-axis, by Calculating the relative offset on the y-axis, by Calculating the ratio of the width of the true bounding box to the width of the predicted bounding box to obtain the scaling factor of the width, by Obtaining the scaling factor of the height, where x gt and y gt are the x coordinate and y coordinate of the center point of the true bounding box, x pred and y pred are the x coordinate and y coordinate of the center point of the predicted bounding box, w pred and h pred are the width and height of the predicted bounding box respectively, α x and α y are the weight coefficients on the x-axis and y-axis, Δw and Δh are the scaling factors of the width and height respectively, β x and β y are the adjustment coefficients of the angle difference, Δx and Δy are the offsets of the center point on the x-axis and y-axis respectively, θ diff is the angle difference between the predicted box and the true box, Δw and Δh are the scaling factors of the width and height respectively, wgt and h gt are the width and height of the ground truth bounding box, γ w and γ h are adjustment coefficients based on the degree of overlap, and IoU is the intersection over union between the predicted bounding box and the ground truth bounding box;
[0033] Adjust the predicted bounding box according to the offset of the bounding box to obtain an accurate target location.
[0034] Preferably, according to the accurate target location, refine the region of interest of the feature map, and classify and evaluate the confidence of the refined region of interest to obtain the target detection result and confidence, including:
[0035] According to the accurate target location, through x ref = x roi +(Δx × μ x ) × w roi × η w,x Refine the x coordinate of the center point of the region of interest, through y ref = y roi +(Δy × μ y ) × η h,y × h roi Refine the y coordinate of the center point of the region of interest, through w ref = w roi × exp(Δw × ξ w ) × G w Refine the center width of the region of interest, through h ref = h roi × exp(Δh × ξ h ) × G h Refine the center height of the region of interest, where x ref and y ref are the coordinates of the center point of the refined region of interest, x roi and y roi are the coordinates of the center point of the initially detected region of interest, Δx and Δy are the relative offsets of the center point of the region of interest in the x-axis and y-axis directions, w roi and h roi are the width and height of the initially detected region of interest, μ x and μ y are learnable scaling coefficients, η w,x is an adjustment coefficient related to the width, η h,y is an adjustment coefficient related to the height, ξ w is a learnable adjustment coefficient, G w is an additional width adjustment coefficient, ξ h is a learnable adjustment coefficient, Gh is an additional height adjustment coefficient, Δw and Δh are the scaling factors of width and height respectively, and w ref and h ref are the width and height of the refined region of interest;
[0036] According to the refined center point coordinates, width and height, obtain the coordinates and size of the refined region of interest;
[0037] Classify the refined region of interest to obtain the category of the region of interest;
[0038] Convert the category of the region of interest into a confidence value through the Sigmoid function or Softmax function to obtain the object detection result and confidence. The object detection result includes the coordinates, size and corresponding category of the refined region of interest.
[0039] Preferably, according to the object detection result and confidence, mark the image data of the image containing the target item to realize the recognition of operation and inspection information, including:
[0040] Scan the image data according to the object detection result to obtain the detection result of the target item, and provide a confidence score for each detection result;
[0041] Set a confidence threshold, compare the confidence of the detection result with the set threshold, and screen out the results higher than the threshold;
[0042] For the high-confidence target items screened out, mark them with bounding boxes and labels on the original image to realize the recognition of operation and inspection information.
[0043] In a second aspect, an intelligent operation and inspection information recognition system based on an acquisition field includes a data acquisition module, a feature extraction module, a feature processing module, and a classification and recognition module, where:
[0044] The data acquisition module is used to regularly collect the image data of the target item, construct an intelligent operation and inspection information acquisition field, and obtain the image data of the security inspection channel; perform border padding and mean subtraction processing on the image data to obtain the processed image data;
[0045] The feature extraction module is used to pass the processed image data through a fully convolutional structure to extract the feature map of the image; pass the feature map through a region proposal network to generate the feature map of the region of interest containing the target, and generate a feature map of a unified size after pooling transformation of the feature map of the region of interest;
[0046] The feature processing module is used to locate the feature map of the unified size and calculate the offset of the bounding box at the same time to obtain an accurately located target;
[0047] A classification and recognition module, which is used to precisely refine the region of interest of the feature map according to the precisely located target, and classify and evaluate the confidence of the refined region of interest, so as to obtain the target detection result and confidence; according to the target detection result and confidence, mark the image data of the image containing the target item, so as to realize the recognition of operation and inspection information.
[0048] In a third aspect, a computing device includes:
[0049] One or more processors;
[0050] A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method.
[0051] In a fourth aspect, a computer-readable storage medium stores a program that implements the method when executed by a processor.
[0052] The above solution of the present invention has at least the following beneficial effects:
[0053] The present invention regularly collects data, which can ensure that the system always has the latest image data of the target item, so as to adapt to the changing environment and items. The intelligent operation and inspection information acquisition field provides a structured data collection environment, which is conducive to obtaining high-quality and standardized image data. The edge padding process can reduce the edge effect in the convolution operation, enabling the model to better learn image features. The mean subtraction process helps with data standardization. The fully convolutional structure can efficiently extract features in the image, avoiding the fully connected layer in the traditional convolutional neural network, thus retaining the spatial information of the image and being conducive to the precise positioning of the target. The region proposal network can pre-screen the regions that may contain the target, thus avoiding redundant calculations for the entire image. Generating candidate regions through preset scales and ratios can better adapt to targets of different sizes and shapes, improving the detection accuracy. The feature maps of unified size make subsequent processing more standardized and convenient. By calculating the offset of the bounding box, more precise positioning of the target can be achieved. The refinement step can further refine the position and size of the target, improving the detection accuracy. Through classification and confidence evaluation, the reliability of the detection result can be quantitatively evaluated. By marking the image, the detection result can be visually displayed, facilitating quick recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a schematic flowchart of a method for identifying intelligent operation and inspection information based on an acquisition field provided by an embodiment of the present invention.
[0055] Figure 2 is a schematic diagram of a system for identifying intelligent operation and inspection information based on an acquisition field provided by an embodiment of the present invention. Detailed implementation manners
[0056] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0057] As Figure 1 shown, an intelligent operation and maintenance information recognition method based on a collection field is proposed in an embodiment of the present invention. The method includes the following steps:
[0058] Step 11: Regularly collect image data of a target item, construct an intelligent operation and maintenance information collection field, and obtain image data of a security inspection channel;
[0059] Step 12: Perform border padding and mean subtraction on the image data to obtain processed image data;
[0060] Step 13: Pass the processed image data through a fully convolutional structure to extract a feature map of the image;
[0061] Step 14: Pass the feature map through a region proposal network to generate a feature map of an interested region containing the target, and generate a feature map of a unified size after pooling transformation of the feature map of the interested region;
[0062] Step 15: Locate the feature map of the unified size and calculate the offset of the bounding box at the same time to obtain an accurately located target;
[0063] Step 16: Refine the interested region of the feature map according to the accurately located target, and classify and evaluate the confidence of the refined interested region to obtain a target detection result and a confidence;
[0064] Step 17: Mark the image of the image data containing the target item according to the target detection result and the confidence to realize the recognition of operation and maintenance information.
[0065] In the embodiments of the present invention, the data set is continuously updated and improved to adapt to the changes and diversities of the target items; an intelligent operation and inspection information acquisition field is constructed, which can automatically and efficiently collect the image data of the security inspection passage, reduce manual intervention, and improve the efficiency and quality of data acquisition. Image data padding and mean subtraction processing can enhance the stability and consistency of the image data, reduce the influence of factors such as illumination and contrast on subsequent processing, and improve the reliability of image processing. The fully convolutional structure can efficiently extract rich feature information from the image, the region proposal network generates regions of interest that can quickly and accurately locate the regions that may contain the target, and the pooling transformation generates feature maps of a unified size, which can ensure the consistent representation of targets of different sizes and proportions on the feature map. Precise target localization and calculation of the bounding box offset can more accurately determine the position and range of the target, reduce the localization error, and improve the accuracy and reliability of target detection. Refining the regions of interest and performing classification and confidence evaluation can further improve the accuracy of target detection, and at the same time give the confidence of the target belonging to various categories. Marking the image according to the target detection results and confidence can intuitively display the position and recognition results of the target item in the image, facilitate the operation and inspection personnel to quickly and accurately identify the operation and inspection information, and improve the work efficiency and accuracy.
[0066] In a preferred embodiment of the present invention, in step 11 above, the image data of the target item is regularly collected, and an intelligent operation and inspection information acquisition field is constructed to obtain the image data of the security inspection passage, including:
[0067] Step 111, regularly collect the image data of the target item;
[0068] Step 112, construct an intelligent operation and inspection information acquisition field to obtain the image data of the security inspection passage.
[0069] In the embodiments of the present invention, regularly collecting the image data of the target item can continuously update the data set required for model training, enable the model to adapt to the changes of the target item, and improve the detection accuracy. By regularly collecting data, more images of the target item in different environments and at different angles can be captured, thereby enhancing the robustness and generalization ability of the detection system. By constructing an intelligent operation and inspection information acquisition field, the image data of the security inspection passage can be automatically obtained, reducing manual intervention, and thus improving the efficiency of security inspection. The intelligent operation and inspection information acquisition field can monitor the situation of the security inspection passage in real time. Once an abnormal or suspicious item is found, an alarm can be issued immediately to improve safety.
[0070] In another embodiment of the present invention, step 11 specifically includes the following implementation steps:
[0071] Step 111: Use a high-resolution camera or scanner to capture clear images of the target items. These devices should be installed at key positions in the security inspection passage to ensure that all items passing through the security inspection can be captured. Store the captured image data in a secure and reliable database.
[0072] Step 112: According to the characteristics of the operation and inspection process, reasonably layout the positions of the cameras and scanners, install high-resolution cameras and scanners to capture detailed images of the items passing through the security inspection passage, construct a stable network system to connect the cameras and scanners to the central processing unit, integrate the intelligent operation and inspection information collection field with the existing security inspection system, and develop a dedicated software platform to manage the intelligent operation and inspection information collection field. This platform should have functions such as real-time display, storage, retrieval, and analysis of image data. At the same time, an intuitive and easy-to-use user interface should also be designed. Use encryption technology and access control mechanisms to ensure the security of the image data.
[0073] In a preferred embodiment of the present invention, in the above step 12, perform edge padding and mean subtraction on the image data to obtain processed image data, including:
[0074] Step 121: Determine the type and size of the image edge padding.
[0075] Step 122: According to the type and size of the image edge padding, add additional pixel values around the boundary of the image to obtain the edge-padded image.
[0076] Step 123: Perform mean subtraction on the edge-padded image through to obtain the processed image data, where P ijk represents the pixel value of the i-th row, j-th column, and k-th channel in the original image data, P' ijk represents the pixel value of the i-th row, j-th column, and k-th channel in the image data after mean subtraction, M and N respectively represent the height and width of the image (i.e., the number of rows and columns), and i, j, and k represent the row number, column number, and channel index.
[0077] In the embodiments of the present invention, by determining appropriate padding types and sizes, the information at the image edges can be effectively processed, reducing the interference of edge effects on subsequent image processing and analysis, thereby improving the accuracy of image processing. Padding processing can effectively protect the information at the image edges. Through padding, the images can be made consistent in size, avoiding processing errors caused by size changes and enhancing the stability of the images. Mean subtraction processing can standardize the image data, making the data distribution more uniform. The standardized data can accelerate the training speed of machine learning models, improve the convergence efficiency and performance of the models. Through mean subtraction processing, the dependence of the model on a specific dataset can be reduced, enhancing the generalization ability of the model so that it can better adapt to different datasets and environments.
[0078] In another embodiment of the present invention, step 12, the specific implementation steps include:
[0079] Step 121, analyze the original image data to understand the size, resolution, and boundary conditions of the image. According to the results of the image analysis, select the padding type. Common padding types include zero padding (adding zero-valued pixels around the image boundary), edge replication (copying the pixel values at the image edge for padding), and reflection padding (reflecting the pixel values at the edge with the image edge as the axis for padding), etc. Which type to choose depends on the specific application scenario of the system and the automatic setting of the image characteristics. The padding size refers to the number of additional pixels added around the image boundary. This size should be determined according to the size of the image, the requirements of the target detection algorithm, and the desired results. The padding size can be set through experiments or experience.
[0080] Step 122, according to the padding type and size determined in step 121, add additional pixel values around the boundary of the original image. For example, if zero padding is selected, zero-valued pixels are filled in the boundary area of the new image; if edge replication is selected, the pixel values at the edge of the original image are copied for padding.
[0081] Step 123, initialize the mean-subtracted image, traverse each pixel in the padded image, and calculate the mean of the entire image according to its channel values. Specifically, for each channel, use the formula to calculate the global mean of that channel. Traverse each pixel in the padded image again, and subtract the value of each pixel from the global mean of the corresponding channel. Use the formula for calculation to obtain the mean-subtracted image.
[0082] In a preferred embodiment of the present invention, for the above step 13, passing the processed image data through a fully convolutional structure to extract the feature map of the image includes:
[0083] Step 131, combine through multiple convolutional layers to construct a fully convolutional structure;
[0084] Step 132, perform forward propagation on the processed image data through a fully convolutional structure to extract higher-level feature representations;
[0085] Step 133, obtain the feature map of the image based on the higher-level feature representations.
[0086] In the embodiment of the present invention, through the combination of multiple convolutional layers, a deep network structure can be constructed, so that deeper feature information in the image can be extracted. Through the forward propagation process, the fully convolutional structure can automatically learn the high-level feature representations in the image. These features are more representative than the original pixels and help improve the performance of subsequent tasks. Compared with the fully connected layer, the fully convolutional structure can greatly reduce the number of parameters of the model, thereby reducing the complexity of the model and improving the computational efficiency. The feature map contains rich feature information in the image, such as edges, textures, shapes, etc. These information are very useful for subsequent tasks such as image classification and localization. The feature map can provide an intuitive way to understand and explain how the model processes the image. By observing the feature map, it is possible to understand the information and features that the model focuses on in different convolutional layers.
[0087] In another embodiment of the present invention, Step 13 specifically includes the following implementation steps:
[0088] Step 131, (1) Design convolutional layers: Determine the parameters of each convolutional layer, including the size of the convolutional kernel, the stride, the padding method, and the number of output channels. The selection of parameters should be based on specific task requirements and image characteristics. For example, a smaller convolutional kernel can better capture the local details of the image, while a larger convolutional kernel helps capture more global information. (2) Stack convolutional layers: Stack multiple designed convolutional layers in a certain order to form a fully convolutional structure. In this process, activation functions (such as ReLU) can be introduced to increase the non-linear expression ability of the network, and batch normalization can be used to accelerate training and improve the generalization ability of the model. (3) Set connection layers: In the fully convolutional structure, set some connection layers as needed, such as pooling layers to reduce the dimension of the feature map, or upsampling layers to restore the resolution of the feature map.
[0089] Step 132, (1) Forward propagation: Automatically transfer the image data into the fully convolutional structure for forward propagation calculation. In this process, each convolutional layer will perform a convolution operation on its input to extract the local features in the image. As the data is transmitted in the network, these local features will gradually be combined into higher-level feature representations. (2) Activation and transmission: After each convolutional layer, an activation function is usually used to perform a non-linear transformation on the feature map. The activated feature map will be used as the input of the next layer and continue to be transmitted until the output layer of the fully convolutional structure is reached.
[0090] Step 133, Feature Map Extraction: At the output layer of the fully convolutional structure, obtain the high-level feature maps extracted after multiple convolutional operations.
[0091] In a preferred embodiment of the present invention, in the above step 14, passing the feature maps through a region proposal network to generate feature maps of regions of interest containing the target, and generating feature maps of a unified size after pooling transformation of the feature maps of regions of interest, includes:
[0092] Step 141, in the region proposal network, generate a series of anchor points on the feature maps by means of a sliding window, and each anchor point corresponds to a region on the original image;
[0093] Step 142, preset the scales and ratios of the regions of interest, and according to the scales and ratios of the regions of interest, the region proposal network will generate multiple candidate regions for each anchor point;
[0094] Step 143, predict whether the candidate regions contain the target and the exact position of the target through two branches of classification and regression to obtain the feature maps of the candidate regions;
[0095] Step 144, perform pooling transformation on the feature maps of the candidate regions to generate feature maps of a unified size.
[0096] In the embodiments of the present invention, generating anchor points by means of a sliding window can ensure that each position on the feature maps is taken into account, thereby comprehensively covering the regions that may contain the target. By presetting different scales and ratios, the region proposal network can generate candidate regions adapted to targets of various sizes and ratios, improving the detection ability for different targets. The classification branch can predict whether the candidate regions contain the target, thereby screening out the regions that truly contain the target, and the regression branch can further adjust the positions of the candidate regions to more accurately locate the target. Pooling transformation can convert the feature maps of candidate regions of different sizes and ratios into feature maps of a unified size, providing a standardized input for subsequent classification and recognition tasks and simplifying the processing flow.
[0097] In another embodiment of the present invention, step 14, the specific implementation steps include:
[0098] Step 141, on the feature maps, use a sliding window (usually a small network layer, such as a 3x3 convolutional layer) to traverse. This sliding window will stop at each position on the feature maps and generate an anchor point. Each anchor point corresponds to a reference region of a fixed size on the original image. As the sliding window moves, a dense set of anchor points will be generated on the entire feature maps, covering different positions and scales of the image.
[0099] Step 142, according to the task requirements, preset a set of scales and ratios of regions of interest. For example, different sizes (such as 128x128, 256x256, etc.) and aspect ratios (such as 1:1, 1:2, 2:1, etc.) can be set. For each anchor point, the region proposal network generates multiple candidate regions according to the preset scales and ratios. The region proposal network adjusts the boundaries of each candidate region through bounding box regression to make it closer to the actual position of the target.
[0100] Step 143, the region proposal network includes a classification branch for predicting whether each candidate region contains the target, which is usually implemented by a binary classifier to output the probability that each region is the target or the background. At the same time, the region proposal network also includes a regression branch for further finely adjusting the bounding box coordinates of each candidate region that contains the target. According to the classification and regression results, the candidate regions with high probability of containing the target are selected, and the regions with low probability or background are discarded. The filtered candidate regions constitute the feature map of the regions of interest.
[0101] Step 144, for each region of interest, intercept the corresponding region on the feature map according to its bounding box coordinates, and perform a pooling operation using Region of Interest Pooling or Region of Interest Align. After the pooling operation, the feature map has a unified size. After the pooling transformation, each region of interest is converted into a feature map with a fixed size.
[0102] In a preferred embodiment of the present invention, in the above step 15, the feature map with a unified size is positioned, and at the same time, the offset of the bounding box is calculated to obtain an accurately positioned target, including:
[0103] Step 151, according to the feature map with a unified size, obtain each region of interest;
[0104] Step 152, predict the four coordinate parameters of each region of interest through a bounding box regressor to obtain the predicted bounding box of each region of interest;
[0105] Step 153, through calculate the relative offset on the x-axis, through calculate the relative offset on the y-axis, through calculate the ratio of the true bounding box width to the predicted bounding box width to obtain the width scaling factor, through obtain the height scaling factor, where, x gt and y gt are the x coordinate and y coordinate of the center point of the true bounding box, x pred and y pred are the x coordinate and y coordinate of the center point of the predicted bounding box, w pred and h predThey are the width and height of the predicted bounding box, α x and α y are the weight coefficients on the x-axis and y-axis, Δw and Δh are the scaling factors of the width and height respectively, β x and β y are the adjustment coefficients for the angle difference, Δx and Δy are the offsets of the center point on the x-axis and y-axis respectively, θ diff is the angle difference between the predicted box and the ground truth box, Δw and Δh are the scaling factors of the width and height respectively, w gt and h gt are the width and height of the ground truth bounding box, γ w and γ h are the adjustment coefficients based on the overlap degree, IoU is the intersection over union between the predicted box and the ground truth box;
[0106] Step 154, adjust the predicted bounding box according to the offset of the bounding box to obtain an accurately positioned target.
[0107] In the embodiment of the present invention, obtaining the region of interest from the feature map of unified size can ensure that each region undergoes the same processing flow. Through the coordinate parameters predicted by the bounding box regressor, the position of the target in the image can be initially determined. By calculating the offset, the position and size of the predicted bounding box can be accurately adjusted to make it closer to the ground truth bounding box, thereby improving the accuracy of target positioning. The offsets in multiple dimensions such as the x-axis, y-axis, width, height, and angle are considered. By introducing the weight coefficients and adjustment coefficients, the calculation method of the offsets in each dimension can be flexibly adjusted according to the actual situation, further improving the accuracy and robustness of the positioning. By adjusting the predicted bounding box according to the offset, the accurate positioning of the target can be finally achieved, meeting the requirements for the accuracy of target positioning in practical applications.
[0108] In another embodiment of the present invention, step 15 specifically includes the following implementation steps:
[0109] Step 151, based on the unified-size feature map generated in step 144, obtain each region of interest. From the feature map, the coordinate information of each region of interest can be directly obtained.
[0110] Step 152, process each region of interest using the trained bounding box regressor, and predict more accurate bounding box coordinates according to the features of the region of interest. Four coordinate parameters will be obtained according to the regressor, which usually represent the center point coordinates, width, and height of the bounding box.
[0111] Step 153, calculate the relative offsets on the x-axis and y-axis using the given formula. These offsets represent the difference between the center point of the predicted bounding box and the center point of the ground truth bounding box. Parameters in the formula such as α x and αy , β x and β y and γ w and γ h are learned from the training historical data and are used to adjust the calculation of the offset to more accurately reflect the real situation. The scaling factors of the width and height are calculated using the given formula, and the sub is used to adjust the size of the predicted bounding box to make it closer to the real bounding box. The IoU in the formula is an index to measure the overlapping degree between the predicted box and the real box.
[0112] Step 154: According to the offset and the scaling factor calculated in Step 153, adjust the predicted bounding box in Step 152. By adjusting the center point position and size of the predicted bounding box, more accurate target positioning can be obtained.
[0113] In a preferred embodiment of the present invention, in the above Step 16, according to the accurately positioned target, refine the region of interest of the feature map, and classify and evaluate the confidence of the refined region of interest to obtain the target detection result and the confidence, including:
[0114] Step 161: According to the accurately positioned target, refine the x coordinate of the center point of the region of interest through x ref = x roi +(Δx × μ x ) × w roi × η w,x refine the y coordinate of the center point of the region of interest through y ref = y roi +(Δy × μ y ) × η h,y × h roi refine the width of the center of the region of interest through w ref = w roi × exp(Δw × ξ w ) × G w refine the height of the center of the region of interest through h ref = h roi × exp(Δh × ξ h ) × G h where x ref and y ref are the coordinates of the center point of the refined region of interest, x roi and y roi are the coordinates of the center point of the initially detected region of interest, Δx and Δy are the relative offsets of the center point of the region of interest in the x-axis and y-axis directions, w roi and h roi are the width and height of the initially detected region of interest, μ x and μ yis a learnable scaling factor, η w,x is a width-related adjustment factor, η h,y is a height-related adjustment factor, ξ w is a learnable adjustment factor, G w is an additional width adjustment factor, ξ h is a learnable adjustment factor, G h is an additional height adjustment factor, Δw and Δh are the scaling factors of width and height respectively, w ref and h ref are the width and height of the refined region of interest;
[0115] Step 162, according to the refined center point coordinates, width and height, obtain the coordinates and dimensions of the refined region of interest;
[0116] Step 163, classify the refined region of interest to obtain the category of the region of interest;
[0117] Step 164, convert the category of the region of interest into a confidence value through the Sigmoid function or Softmax function to obtain the object detection result and confidence. The object detection result includes the coordinates, dimensions of the refined region of interest and the corresponding category.
[0118] In the embodiment of the present invention, by refining the center point coordinates, width and height of the region of interest, the position and size of the target can be further refined, thereby improving the positioning accuracy of object detection. In the actual scenario, the target may be deformed due to factors such as pose change and occlusion. By refining the region of interest, these deformations can be better adapted, and the robustness of the detection can be improved. The coordinates and dimensions of the refined region of interest can more accurately describe the position and size of the target in the image. By classifying the refined region of interest, the category of the target can be accurately identified to meet the requirements of target classification in practical applications. By converting the category of the region of interest into a confidence value, the reliability of the object detection result can be quantitatively evaluated. Accurate confidence evaluation helps to improve the performance of the entire detection system, including reducing the false detection rate, improving the accuracy rate and other key indicators, thereby enhancing the practicability and reliability of the system.
[0119] In another embodiment of the present invention, step 16, the specific implementation steps include:
[0120] Step 161, use the formula to refine the x coordinate of the center point of the region of interest for the region of interest. Similarly, use the formula to refine the y coordinate of the center point of the region of interest, use the formula to refine the width of the region of interest, and use the formula to refine the height of the region of interest. In the formula, μ x and μ yis a learnable scaling factor, η w,x is a width-related adjustment factor, η h,y is a height-related adjustment factor, ξ w is a learnable adjustment factor, G w is an additional width adjustment factor, ξ h is a learnable adjustment factor, G h is an additional height adjustment factor. In deep learning, parameters such as μ x , μ y , η w,x , η h,y , ξ w , G w , ξ h , ξh, G h are usually automatically learned through the backpropagation algorithm during the training process. These parameters are part of the model and are usually initialized to a certain value (such as 0 or 1), and then updated according to the gradient of the loss function during the training process. Specifically, the calculation steps of these parameters are as follows: (1) Before the start of model training, assign initial values to these parameters. (2) During the model training process, each batch of data will go through a forward pass through the model. This includes using the current parameter values (including μ x , μ y etc.) to calculate the predicted values. (3) Compare the predicted values of the model with the true values and calculate the value of the loss function (such as mean squared error, cross entropy, etc.). (4) According to the gradient of the loss function with respect to the model parameters, use an optimization algorithm (such as gradient descent, Adam, etc.) to update the model parameters, including μ x , μ y , η w,x , μ h,y , ξ w , G w , ξ h , ξh, G h . (5) Iteratively update until the preset number of training epochs is reached.
[0121] Step 162: Integrate the refined center point coordinates and width and height calculated in Step 161. Based on the refined parameters, determine the upper left and lower right coordinates of the refined region of interest, thereby obtaining its complete coordinate and size information.
[0122] Step 163: According to the coordinates and dimensions of the refined region of interest, extract the features of the refined region of interest from the original feature map or the corresponding feature layer, and input the extracted features into a pre-trained classifier for classification operations, such as the last layer of a convolutional neural network (CNN); the output of the classifier is the score or probability of each category, and the category with the highest score or the largest probability is autonomously selected as the category of the refined region of interest.
[0123] Step 164: Convert the category score or probability obtained in Step 163 into a confidence score value through the Sigmoid function or the Softmax function, where the Sigmoid function is used for multi-label classification problems (i.e., a sample can belong to multiple categories at the same time), and the Softmax function is used for the probability distribution of each category (i.e., a sample can only belong to one category); integrate the coordinates, dimensions, corresponding category, and confidence value of the refined region of interest to form a complete object detection result, and the object detection result includes the coordinates, dimensions, category, and confidence score value of each detected object.
[0124] In a preferred embodiment of the present invention, the above-mentioned Step 17: Mark the image data of the image containing the target item according to the object detection result and the confidence level to achieve the recognition of operation and inspection information, including:
[0125] Step 171: Scan the image data according to the object detection result to obtain the detection result of the target item, and provide a confidence score for each detection result;
[0126] Step 172: Set a confidence threshold, compare the confidence level of the detection result with the set threshold, and filter out the results higher than the threshold;
[0127] Step 173: For the high-confidence target items selected, mark them with bounding boxes and labels on the original image to achieve the recognition of operation and inspection information.
[0128] In the embodiments of the present invention, by scanning the entire image data, it can be ensured that every target in the image is detected without missing any possible target items. Providing a confidence score for each detection result can quantitatively evaluate the reliability of each detection result, which helps with subsequent screening and labeling work. By setting a confidence threshold and screening the results higher than the threshold, low-confidence detection results can be filtered out, thereby improving the overall detection accuracy. The screening process helps to eliminate possible false detections, ensuring that only high-confidence targets are marked and recognized, reducing false alarms. By marking the target items with bounding boxes and labels on the original image, the detection results can be visually displayed, facilitating quick understanding and recognition by users. By marking the image data containing the target items, the identification of operation and inspection information can be effectively achieved, contributing to the efficient progress of operation and maintenance work.
[0129] In another embodiment of the present invention, step 17, the specific implementation steps include:
[0130] Step 171, obtain the image data for target detection, and use the convolutional neural network trained in the above steps to scan the image data to obtain a series of detection results. Each result includes the position of the target item in the image (usually the coordinates of a bounding box) and the confidence score that the item belongs to a certain category. This score represents the degree of certainty of the convolutional neural network that the detected target is a specific category.
[0131] Step 172, according to the application requirements and model performance, set a confidence threshold. This threshold is used to judge the reliability of the detection results. For example, the threshold can be set to 0.5, which means that only the detection results with a confidence score higher than or equal to 0.5 will be considered reliable. Traverse all the detection results and compare the confidence score of each result with the set threshold; if the confidence score of a certain detection result is higher than or equal to the set threshold, it is regarded as a high-confidence result and retained. The results below the threshold will be discarded because they are considered unreliable.
[0132] Step 173, obtain the position information and category of the target item from the high-confidence results screened in step 172, and use a graphics processing library to draw bounding boxes on the original image according to the position information of the target item. These bounding boxes are usually represented by rectangles to highlight the detected target items, and text labels indicating the category of the detected target items are added above or beside the bounding boxes. Save the marked image to the file system or directly display it on the user interface. The target items and their positions in the image can be quickly identified, thus effectively realizing the identification of operation and inspection information.
[0133] Such as Figure 2As shown in the figure, an intelligent operation and maintenance information recognition system 20 based on a collection field according to an embodiment of the present invention includes a data acquisition module, a feature extraction module, a feature processing module, and a classification and recognition module, where:
[0134] The data acquisition module 21 is configured to regularly collect image data of a target item, construct an intelligent operation and maintenance information collection field, and acquire image data of a security inspection channel; perform border padding and mean subtraction processing on the image data to obtain processed image data;
[0135] The feature extraction module 22 is configured to pass the processed image data through a fully convolutional structure to extract a feature map of the image; pass the feature map through a region proposal network to generate a feature map of an interested region containing the target, and generate a feature map of a unified size after pooling transformation of the feature map of the interested region;
[0136] The feature processing module 23 is configured to locate the feature map of the unified size and calculate the offset of the bounding box at the same time to obtain an accurately located target;
[0137] The classification and recognition module 24 is configured to refine the interested region of the feature map according to the accurately located target, classify and evaluate the confidence of the refined interested region to obtain a target detection result and a confidence; mark the image of the image data containing the target item according to the target detection result and the confidence to realize the recognition of operation and maintenance information.
[0138] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An intelligent operation and inspection information recognition method based on a collection field, characterized in that, The method includes: Regularly collecting image data of the target object, constructing an intelligent operation and inspection information acquisition field, and obtaining image data of the security inspection channel; Performing border padding and mean subtraction on the image data to obtain processed image data; Passing the processed image data through a fully convolutional structure to extract the feature map of the image; Passing the feature map through a region proposal network to generate a feature map of the region of interest containing the target, and generating a feature map of a unified size after pooling transformation of the feature map of the region of interest; Locating the feature map of the unified size and simultaneously calculating the offset of the bounding box to obtain the accurately located target; According to the accurately located target, refining the region of interest of the feature map, and classifying and performing confidence evaluation on the refined region of interest to obtain the target detection result and confidence; Marking the image of the image data containing the target object according to the target detection result and confidence to achieve the recognition of operation and inspection information.
2. The intelligent operation and inspection information recognition method based on the acquisition site according to claim 1, wherein Performing border padding and mean subtraction on the image data to obtain processed image data, including: Determining the type and size of image border padding; Adding additional pixel values around the boundary of the image according to the type and size of image border padding to obtain the border-padded processed image; The image processed with border filling is passed through to perform mean subtraction processing to obtain processed image data, where P ijk represents the pixel value of the i-th row, j-th column, and k-th channel in the original image data, and P i ′ jk represents the pixel value of the i-th row, j-th column, and k-th channel in the image data after mean subtraction processing. M and N respectively represent the height and width of the image, and i, j, and k represent the row number, column number, and channel index.
3. The intelligent operation and maintenance information recognition method based on the acquisition site according to claim 2, wherein Passing the processed image data through a fully convolutional structure to extract the feature map of the image, including: Constructing a fully convolutional structure by combining multiple convolutional layers; Performing forward propagation of the processed image data through the fully convolutional structure to extract a higher-level feature representation; Obtaining the feature map of the image according to the higher-level feature representation.
4. The intelligent operation and inspection information recognition method based on a collection site according to claim 3, wherein Passing the feature map through a region proposal network to generate a feature map of the region of interest containing the target, and generating a feature map of a unified size after pooling transformation of the feature map of the region of interest, including: In the region proposal network, generating a series of anchor points on the feature map by means of a sliding window, and each anchor point corresponds to a region on the original image; Presetting the scale and ratio of the region of interest, and according to the scale and ratio of the region of interest, the region proposal network generates multiple candidate regions for each anchor point; Predicting whether the candidate region contains the target and the accurate position of the target through two branches of classification and regression to obtain the feature map of the candidate region; Performing pooling transformation on the feature map of the candidate region to generate a feature map of a unified size.
5. The intelligent operation and inspection information recognition method based on a collection site according to claim 4, characterized in that, Locating the feature map of the unified size and simultaneously calculating the offset of the bounding box to obtain the accurately located target, including: Obtaining each region of interest according to the feature map of the unified size; Predicting the four coordinate parameters of each region of interest through a bounding box regressor to obtain the predicted bounding box of each region of interest; By calculating the relative offset on the x-axis, by calculating the relative offset on the y-axis, by calculating the ratio of the true bounding box width to the predicted bounding box width to obtain the width scaling factor, by obtaining the height scaling factor, where and are the x-coordinate and y-coordinate of the center point of the true bounding box, x pred and y pred are the x-coordinate and y-coordinate of the center point of the predicted bounding box, w pred and h pred are the width and height of the predicted bounding box respectively, α x and α y are the weight coefficients on the x-axis and y-axis, Δw and Δh are the width and height scaling factors respectively, β x and β y are the adjustment coefficients for the angle difference, Δx and Δy are the offsets of the center point on the x-axis and y-axis respectively, θ diff is the angle difference between the predicted box and the true box, Δw and Δh are the width and height scaling factors respectively, and are the width and height of the true bounding box, γ w and γ h are the adjustment coefficients based on the overlap degree, IoU is the intersection over union between the predicted box and the true box; Adjusting the predicted bounding box according to the offset of the bounding box to obtain the accurately located target.
6. The intelligent operation and inspection information recognition method based on the acquisition site according to claim 5, characterized in that According to the accurately located target, refining the region of interest of the feature map, and classifying and performing confidence evaluation on the refined region of interest to obtain the target detection result and confidence, including: According to the accurately positioned target, through x ref = x roi +(Δx × μ x ) × w roi × η w,x Refine the x - coordinate of the center point of the region of interest. Through y ref = y roi +(Δy × μ y ) × η h,y × h roi Refine the y - coordinate of the center point of the region of interest. Through w ref = w roi × exp(Δw × ξ w ) × G w Refine the center width of the region of interest. Through h ref = h roi × exp(Δh × ξ h ) × G h Refine the center height of the region of interest. Wherein, x ref and y ref are the coordinates of the center point of the refined region of interest, x roi and y roi are the coordinates of the center point of the initially detected region of interest, Δx and Δy are the relative offsets of the center point of the region of interest in the x - axis and y - axis directions, w roi and h roi are the width and height of the initially detected region of interest, μ x and μ y are the scaling factors, η w,x is the adjustment factor, η h,y is the adjustment factor, ξ w is the adjustment factor, G w is the width adjustment factor, ξ h is the adjustment factor, G h is the height adjustment factor, Δw and Δh are the scaling factors of the width and height respectively, w ref and h ref are the width and height of the refined region of interest; Obtaining the coordinates and size of the refined region of interest according to the refined center point coordinates, width and height; Classifying the refined region of interest to obtain the category of the region of interest; Convert the category of the region of interest into a confidence value through the Sigmoid function or Softmax function to obtain the object detection result and confidence. The object detection result includes the refined coordinates, size, and corresponding category of the region of interest.
7. The intelligent operation and inspection information recognition method based on a collection site according to claim 6, characterized in that Mark the image data of the image containing the target item according to the object detection result and confidence to achieve the recognition of operation and inspection information, including: Scan the image data according to the object detection result to obtain the detection result of the target item and provide a confidence score for each detection result; Set a confidence threshold, compare the confidence of the detection result with the set threshold, and filter out the results higher than the threshold; For the high-confidence target items selected, mark them with bounding boxes and labels on the original image to achieve the recognition of operation and inspection information.
8. An intelligent operation and maintenance information recognition system based on a collection site, characterized in that, It includes a data acquisition module, a feature extraction module, a feature processing module, and a classification and recognition module, where: The data acquisition module is used to regularly collect the image data of the target item, build an intelligent operation and inspection information collection field, and obtain the image data of the security inspection channel; perform border padding and mean subtraction on the image data to obtain the processed image data; The feature extraction module is used to pass the processed image data through a fully convolutional structure to extract the feature map of the image; pass the feature map through the region proposal network to generate the feature map of the region of interest containing the target, and generate a feature map of a unified size after pooling transformation of the feature map of the region of interest; The feature processing module is used to locate the feature map of the unified size and calculate the offset of the bounding box at the same time to obtain an accurately located target; The classification and recognition module is used to refine the region of interest of the feature map according to the accurately located target, and classify and evaluate the confidence of the refined region of interest to obtain the object detection result and confidence; mark the image data of the image containing the target item according to the object detection result and confidence to achieve the recognition of operation and inspection information.
9. A computing device, characterized in that, It includes: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A program is stored in the computer-readable storage medium, and when the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.