Workpiece operation identification method and device, medium and program product

By using multi-angle video acquisition and deep learning model annotation to identify workpiece operations, the problem of detection omissions caused by manual visual inspection is solved, and the accuracy of workpiece installation and the reliability and safety of equipment are improved.

CN121330583APending Publication Date: 2026-01-13中国电气装备集团科学技术研究院有限公司 +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511530512.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

In the intelligent manufacturing scenario of power transmission and transformation equipment, manual visual inspection of workpieces is prone to oversights, leading to incorrect assembly of workpieces and affecting the reliability and safety of the equipment.

Method used

Multi-angle video acquisition is used to obtain workpiece operation video stream information. Deep learning workpiece detection model and operation recognition model are used for annotation and recognition to determine the target workpiece and its operation category. The operation recognition result is output through confidence judgment.

Benefits of technology

This improved the accuracy and precision of operation identification during workpiece installation, and enhanced the manufacturing reliability and safety of power transmission and transformation equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121330583A_ABST
    Figure CN121330583A_ABST
Patent Text Reader

Abstract

The invention discloses a workpiece operation identification method and device, a medium and a program product. The method comprises the steps that multiple pieces of workpiece operation video stream information are acquired; inputting each workpiece operation video frame into a pre-trained workpiece detection model to obtain a marked workpiece image after the workpiece is marked and output by the workpiece detection model; according to multiple pieces of workpiece labeling information and multiple pieces of workpiece background labeling information in multiple pieces of labeling workpiece images of each piece of workpiece operation video stream information, determining a target workpiece in the multiple pieces of workpiece operation video stream information; and inputting the plurality of workpiece operation video frames corresponding to the target workpiece into a pre-trained operation identification model to obtain an operation identification result output by the operation identification model. According to the method, the accuracy of the operation identification result of workpiece installation and the precision of workpiece operation are improved, and then the reliability and safety of power transmission and transformation equipment manufacturing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of industrial intelligent technology, and in particular to a workpiece operation recognition method, device, medium and program product. Background Technology

[0002] In the intelligent manufacturing scenario of power transmission and transformation equipment, the precise positioning and accurate assembly of key components such as insulators and fittings in complex workshops are the core links to ensure the quality of power equipment.

[0003] Currently, workpiece operation inspection mainly relies on manual visual inspection. In the intelligent manufacturing scenario of power transmission and transformation equipment, the assembly process operation of the operators is checked to ensure the accuracy of workpiece installation.

[0004] However, manual visual inspection requires workers to spend a long time checking the operation of the workpieces, which may lead to oversights and result in incorrect assembly of the workpieces not being detected. Furthermore, the results of manual visual inspection rely on the experience of the personnel, which may lead to inaccurate identification of the workpiece installation process, thereby affecting the reliability and safety of power transmission and transformation equipment. Summary of the Invention

[0005] This application provides a workpiece operation recognition method, device, medium, and program product to solve the problem that in the prior art, manual visual inspection requires personnel to manually check the workpiece operation for a long time, which may lead to oversights and result in the workpiece being incorrectly assembled without being detected. Furthermore, the results of manual visual inspection rely on personnel experience, which may lead to inaccurate workpiece installation operation recognition results, thereby affecting the reliability and safety of power transmission and transformation equipment.

[0006] In a first aspect, this application provides a workpiece operation recognition method, the method comprising:

[0007] Acquire multiple workpiece operation video stream information; wherein, the multiple workpiece operation video stream information is video stream information from multiple shooting angles, and each workpiece operation video stream information includes multiple workpiece operation video frames;

[0008] Each workpiece operation video frame is input into a pre-trained workpiece detection model to obtain an annotated workpiece image output by the workpiece detection model; wherein, the annotated workpiece image includes at least one workpiece annotation information and workpiece background annotation information corresponding to the workpiece annotation information;

[0009] Based on multiple workpiece annotation information and multiple workpiece background annotation information in multiple annotated workpiece images of each workpiece operation video stream, the target workpiece in the multiple workpiece operation video streams is determined.

[0010] Multiple workpiece operation video frames corresponding to the target workpiece are input into a pre-trained operation recognition model to obtain the operation recognition result output by the operation recognition model; wherein, the operation recognition result includes the workpiece operation category and the corresponding confidence level.

[0011] Secondly, this application provides a workpiece operation recognition device, comprising:

[0012] The acquisition module is used to acquire multiple workpiece operation video stream information; wherein, the multiple workpiece operation video stream information is video stream information from multiple shooting angles, and each workpiece operation video stream information includes multiple workpiece operation video frames;

[0013] The first input module is used to input each of the workpiece operation video frames into a pre-trained workpiece detection model to obtain an annotated workpiece image output by the workpiece detection model after annotating the workpiece; wherein, the annotated workpiece image includes at least one workpiece annotation information and workpiece background annotation information corresponding to the workpiece annotation information;

[0014] The determination module is used to determine the target workpiece in the plurality of workpiece operation video stream information based on the plurality of workpiece annotation information and the plurality of workpiece background annotation information in the plurality of annotated workpiece images of each workpiece operation video stream information.

[0015] The second input module is used to input multiple workpiece operation video frames corresponding to the target workpiece into a pre-trained operation recognition model to obtain the operation recognition result output by the operation recognition model; wherein, the operation recognition result includes the workpiece operation category and the corresponding confidence level.

[0016] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the workpiece operation recognition method as described in the first aspect of this application.

[0017] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the workpiece operation recognition method as described in the first aspect of this application.

[0018] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the workpiece operation recognition method as described in the first aspect of this application.

[0019] The present application's solution involves acquiring multiple workpiece operation video streams; wherein the multiple workpiece operation video streams are video streams from multiple shooting angles, and each workpiece operation video stream includes multiple workpiece operation video frames; each workpiece operation video frame is input into a pre-trained workpiece detection model to obtain an annotated workpiece image output by the workpiece detection model; wherein the annotated workpiece image includes at least one workpiece annotation and corresponding workpiece background annotation; based on the multiple workpiece annotations and multiple workpiece background annotations in the multiple annotated workpiece images of each workpiece operation video stream, a target workpiece is determined in the multiple workpiece operation video streams; the multiple workpiece operation video frames corresponding to the target workpiece are input into a pre-trained operation recognition model to obtain the operation recognition result output by the operation recognition model; wherein the operation recognition result includes the workpiece operation category and the corresponding confidence level. The method of this application annotates the workpiece in the video stream information from multiple shooting angles, identifies the target workpiece, and identifies the operation of the target workpiece to obtain the operation type and confidence level. This avoids the possibility of detection omissions that may occur during manual visual inspection, improves the accuracy of the operation identification results of workpiece installation and the precision of workpiece operation, and thus enhances the reliability and safety of power transmission and transformation equipment manufacturing. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating the workpiece operation recognition method provided in this application;

[0022] Figure 2 This is another flowchart illustrating the workpiece operation recognition method provided in this application;

[0023] Figure 3 This is a schematic diagram of the workpiece operation recognition device provided in this application;

[0024] Figure 4 This is a schematic diagram of the electronic device provided in this application. Detailed Implementation

[0025] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0026] Figure 1 This is a flowchart illustrating a workpiece operation recognition method provided in this application. This method can be executed by a workpiece operation recognition device, which can be implemented using software and / or hardware. In a specific embodiment, the device can be applied in an electronic device, which can be a computer. The following embodiments will illustrate this using the application of the device in an electronic device as an example. Figure 1 The method may specifically include the following steps:

[0027] Step 101: Obtain video stream information of multiple workpiece operations.

[0028] Among them, the multiple workpiece operation video stream information consists of video stream information from multiple shooting angles, and each workpiece operation video stream information includes multiple workpiece operation video frames.

[0029] Specifically, in the manufacturing scenarios of power transmission and transformation equipment, to achieve precise workpiece positioning and accurate identification of process operations, multiple video acquisition devices need to be deployed to obtain workpiece operation video stream information from multiple shooting angles. This prevents workpieces from being unidentifiable due to equipment obstruction within the workshop. For example, by deploying multiple cameras within the workshop to acquire video stream information from multiple angles, these cameras are distributed in different positions and angles, thus covering all areas of workpiece operation. Each video acquisition device acquires a video stream containing the workpiece operation process in real time. This video stream information records the workpiece operation from different perspectives. Each workpiece operation video stream information consists of multiple workpiece operation video frames. These video frames are discrete sampling points of the video stream in time series, reflecting the workpiece's state and operation actions at different points in time. By acquiring video stream information from multiple shooting angles, rich multi-view data support can be provided for subsequent workpiece identification, positioning, and process operation identification, thereby effectively solving problems such as workpiece obstruction and limited viewing angles in complex workshop environments, and improving the robustness and accuracy of the entire system. For example, four 2-megapixel industrial cameras were deployed in the assembly workshop for ultra-high voltage switchgear basin insulators. These cameras were mounted at various angles: a top-down view directly above the workstation, 45-degree side views on both sides, and a global view at the conveyor belt entrance. All cameras were set to a resolution of 1920×1080 and a frame rate of 25fps, achieving millisecond-level synchronization via a network time protocol.

[0030] Step 102: Input each workpiece operation video frame into the pre-trained workpiece detection model to obtain the labeled workpiece image output by the workpiece detection model.

[0031] The labeled workpiece image includes at least one workpiece labeling information and the workpiece background labeling information corresponding to the workpiece labeling information.

[0032] Specifically, each workpiece operation video frame is input one by one into a pre-trained workpiece detection model, which is built based on a deep learning-based object detection algorithm, such as an improved version of the YOLOv8x architecture. The model, trained on a large amount of labeled data, is able to identify workpieces in the video frames. The model processing includes feature extraction, feature aggregation, and bounding box generation. First, a feature extraction network extracts multi-scale feature maps from the input video frames. These feature maps capture key information such as the shape and texture of the workpiece. The feature aggregation network further processes the extracted feature maps to generate high-resolution fused feature maps. This process enhances the model's ability to detect workpieces at different scales. The decoupled detection head uses the information from the fused feature maps for classification and bounding box regression tasks, respectively. The classification branch outputs the probability of the target workpiece's existence, while the regression branch outputs the coordinates and confidence score of the bounding box. To preserve background information, the bounding box is expanded to 1.5 times its original size horizontally and vertically. After processing, the workpiece detection model outputs an image of the labeled workpiece. This image contains both workpiece annotation information and background annotation information. For each detected workpiece, a bounding box is labeled. The box contains the workpiece's category information and confidence level, indicating the model's reliability in detecting that workpiece. Because the bounding box is expanded to 1.5 times its original size, the labeled workpiece image includes not only the workpiece itself but also some background annotation information. This background annotation provides the context of workpiece operations, such as surrounding tools, equipment, and other workpieces, and can be used for subsequent process operation recognition.

[0033] Optionally, the workpiece inspection model is trained according to steps 21 to 24.

[0034] Step 21: Obtain the first training workpiece operation video frame and the first training annotation image.

[0035] The first training workpiece operation video frame is an image from the training video frame library. The training video frame library includes multiple workpiece operation video frames to be trained and a corresponding labeled image to be trained for each workpiece operation video frame to be trained. The labeled image to be trained is the image after the workpiece label anchor box in the workpiece operation video frame to be trained.

[0036] Specifically, the training video frame library is a collection of multiple workpiece operation video frames for training. Each video frame corresponds to a labeled image containing the workpiece's annotation information. These workpiece operation video frames are all collected from actual production scenarios, encompassing various workpiece operation processes, thereby improving the accuracy of subsequent model recognition. Each workpiece operation video frame corresponds to a training labeled image, which is created by adding workpiece annotation information to the original video frame using manual or automatic annotation tools. The annotation information typically includes the workpiece category, bounding box, and confidence score. A certain number of workpiece operation video frames and their corresponding labeled images are randomly selected from the training video frame library. Random selection ensures the diversity and representativeness of the training data. The selected workpiece operation video frames are adjusted to a uniform image size. For example, all images are adjusted to 1920×1080 pixels to ensure consistency in model input. The bounding box coordinates in the labeled images are converted to the format required by the model. For example, the bounding box coordinates are converted from pixel coordinates to normalized coordinates so that the model can better handle images of different sizes. Data augmentation of images can be performed, such as random cropping to simulate different viewpoints and partial occlusion. Random rotation and flipping of images can increase data diversity. Illumination adjustments, such as changing brightness and contrast, can simulate different lighting conditions. Injecting noise, such as Gaussian noise, into the image can improve the model's robustness.

[0037] Step 22: Input the first training workpiece operation video frame into the detection model to be trained to obtain the training labeled workpiece image corresponding to the first training workpiece operation video frame output by the detection model to be trained.

[0038] Specifically, the first training video frame of the workpiece operation is used as input data and fed into the detection model to be trained, resulting in the training labeled workpiece image corresponding to the first training video frame output by the detection model. The detection model typically uses a convolutional neural network structure, extracting image features through multiple convolutional operations. For example, 32 3×3 convolutional kernels are used to extract edge and texture information. After the convolutional layers, activation functions are usually applied to introduce non-linearity. Max pooling or average pooling operations are used to reduce the spatial dimension of the feature map, reducing computation while retaining important features. The above convolution and pooling operations are repeated to progressively extract higher-level features. The detection model to be trained can use a detection head to predict the position and category of each workpiece. The detection head typically consists of multiple convolutional and fully connected layers, outputting the coordinates of the detection box and the workpiece detection confidence score. The model calculates the confidence score for each detection box, representing the model's confidence that a workpiece exists within that box. To reduce overlapping detection boxes, the model can use a non-maximum suppression algorithm to retain the detection box with the highest confidence score and remove other overlapping detection boxes. The training-annotated workpiece image output by the model is generated from the original input image by drawing bounding boxes, category labels, and workpiece detection confidence scores. The training-annotated workpiece image contains bounding boxes, where each detected workpiece is labeled with a bounding box, which can contain workpiece category information and workpiece detection confidence scores. For example, the bounding box coordinates are (x1, y1) and (x2, y2), the category is insulator, and the workpiece detection confidence score is 0.9.

[0039] Optionally, the detection model to be trained includes a feature extraction network, a feature aggregation network, and a detection head. The feature extraction network is used to extract multiple feature maps from the first training video frames of workpiece operation. The feature aggregation network is used to fuse the multiple feature maps to generate a fused feature map. The detection head is used to obtain the workpiece presence probability, detection box coordinates, and workpiece detection confidence based on the fused feature map, and to annotate the workpiece according to the workpiece presence probability, detection box coordinates, and workpiece detection confidence, outputting a trained annotated workpiece image.

[0040] Specifically, the detection model to be trained is a deep learning-based object detection model, consisting of a feature extraction network, a feature aggregation network, and a detection head. The feature extraction network extracts and outputs multiple feature maps from the input image, such as a first training video frame of the workpiece operation. For example, these multiple feature maps can be five-layer multi-scale feature maps with feature dimensions of H / 2×W / 2×64, H / 4×W / 4×128, H / 8×W / 8×256, H / 16×W / 16×512, and H / 32×W / 32×1024. The feature aggregation network fuses the multiple feature maps generated by the feature extraction network to generate a fused feature map. The feature aggregation network can be a dual-path structure of a fused feature pyramid and a path aggregation network. It transmits deep semantic information through a top-down upsampling path and enhances spatial details through a bottom-up downsampling path, fusing feature maps of different scales to generate a high-resolution fused feature map. For example, the fused feature map is a high-resolution fused feature map of H / 2×W / 2×256. The detection head generates the workpiece presence probability, bounding box coordinates, and workpiece detection confidence score based on the fused feature map, and then labels the workpieces, outputting a training-labeled workpiece image. The detection head can employ a decoupled design architecture, separating the classification task from the bounding box regression task into independent branches. The classification branch in the detection head outputs the probability of the target workpiece's presence. For example, it uses fully connected layers or convolutional layers to classify the fused feature map, generating the presence probability for each category. The regression branch in the detection head outputs the bounding box coordinates and workpiece detection confidence score. For example, it uses fully connected layers or convolutional layers to regress the fused feature map, generating the bounding box coordinates and workpiece detection confidence score. Based on the generated bounding box coordinates, category, and workpiece detection confidence score, the workpieces are labeled, generating a training-labeled workpiece image.

[0041] Step 23: Adjust the model parameters of the detection model to be trained based on the first training labeled image and the training labeled workpiece image to obtain a new detection model to be trained.

[0042] Specifically, the loss is calculated between the predicted class probabilities and the actual class labels. Common loss functions include cross-entropy loss. The loss is also calculated between the predicted bounding box coordinates and the actual bounding box coordinates. Common loss functions include mean squared error (MSE) or IoU loss. The classification and regression losses are then weighted and summed to obtain the total loss. The gradient of the total loss with respect to the model parameters is calculated using backpropagation. An optimizer (such as SGD, Adam, etc.) is used to update the model parameters based on the calculated gradient to minimize the total loss, thereby adjusting the model parameters of the training recognition model to obtain a new training recognition model.

[0043] Step 24: Use the second training workpiece operation video frame as the new first training workpiece operation video frame, and the second training annotation image corresponding to the second training workpiece operation video frame as the new first training annotation image. Return to step 22 until the preset iteration end condition is reached, and use the new detection model to be trained as the workpiece detection model.

[0044] The second training workpiece operation video frame is an untrained workpiece operation video frame from the training video frame library.

[0045] Specifically, an untrained workpiece operation video frame is selected from the training video frame library as the second training workpiece operation video frame. This second training workpiece operation video frame is then used as the new first training workpiece operation video frame, and the corresponding second training labeled image is used as the new first training labeled image. The process returns to step 22, which involves inputting the new first training workpiece operation video frame into the detection model to be trained. Based on the obtained training labeled workpiece images, the model parameters are adjusted again, updating the detection model. For example, adjusting the model parameters could include adjusting the learning rate, regularization parameters, etc. The preset iteration termination condition could be reaching a preset number of iterations, the loss value not significantly decreasing in multiple consecutive iterations, or the accuracy on the validation set reaching a preset threshold. When the preset iteration termination condition is met, the training process ends, and the workpiece detection model is obtained.

[0046] Step 103: Based on the multiple workpiece annotation information and multiple workpiece background annotation information in the multiple annotated workpiece images of each workpiece operation video stream information, determine the target workpiece in the multiple workpiece operation video stream information.

[0047] Specifically, in the manufacturing process of power transmission and transformation equipment, to achieve accurate identification and positioning of the target workpiece, it is necessary to determine the target workpiece from workpiece operation video streams acquired from multiple shooting angles. For each workpiece operation video stream, the labeled workpiece image is compared with the workpiece labeling information from different viewpoints. If the labeling information from multiple viewpoints points to the same workpiece—for example, if the detection boxes from multiple viewpoints are consistent in spatial position—then the workpiece is considered the target workpiece. Spatial mapping technology is used to convert the workpiece pixel coordinates from different viewpoints into actual spatial coordinates, and the actual spatial coordinates of the workpiece under different viewpoints are compared. If the actual spatial coordinates of the workpiece are consistent from multiple viewpoints, the workpiece is further confirmed as the target workpiece. Then, based on the workpiece background labeling information in each labeled workpiece image, it is checked whether the background information is consistent with the operation scenario of the target workpiece. For example, if the target workpiece is an insulator, the background information should include tools and equipment related to insulator assembly. For workpieces with high confidence and consistency across multiple viewpoints, the background information is further checked to see if it meets expectations. If the background information is inconsistent with the operation scenario of the target workpiece, the workpiece is excluded, thus confirming the target workpiece. The identified target workpieces are marked, and their position, category, and background information are recorded in each marked workpiece image to facilitate subsequent workpiece positioning and process operation identification. Through these steps, the target workpiece can be effectively identified from multiple workpiece operation video streams, providing accurate data support for subsequent workpiece positioning and process operation identification.

[0048] Optionally, step 103 can be implemented through steps 1031 to 1033.

[0049] Step 1031: Determine the image coordinates of each workpiece based on the multiple workpiece annotation information corresponding to each annotated workpiece image.

[0050] Specifically, each labeled workpiece image includes workpiece annotation information, which includes the coordinates of the detection boxes, typically represented by the pixel coordinates of the top-left and bottom-right corners. Based on the detection box coordinates from the multiple workpiece annotation information corresponding to each labeled workpiece image, the image coordinates of each workpiece are determined.

[0051] Step 1032: For each labeled workpiece image, input the labeled workpiece image and the image coordinates of each workpiece into a pre-trained workpiece coordinate mapping model to obtain the three-dimensional spatial coordinates of each workpiece output by the workpiece coordinate mapping model.

[0052] Specifically, in the workpiece operation video stream information acquired from multiple shooting angles, each video stream is decomposed into multiple workpiece operation video frames. These video frames are processed by a pre-trained workpiece detection model to generate labeled workpiece images. Each labeled workpiece image includes a detection box labeled for each detected workpiece. Each labeled workpiece image and its corresponding workpiece image coordinates are input into a pre-trained workpiece coordinate mapping model. This model is a deep learning-based spatial mapping model that can convert image coordinates into actual three-dimensional spatial coordinates. The workpiece coordinate mapping model can adopt a two-branch convolutional neural network structure, including an image branch and a coordinate branch. The output feature vectors of the image branch and the coordinate branch are concatenated together and fused through a fully connected layer. The output layer is a fully connected layer containing three neurons, each corresponding to the three-dimensional spatial coordinates of the workpiece, thus outputting the three-dimensional spatial coordinates of each workpiece.

[0053] Optionally, the workpiece coordinate mapping model includes an image feature extraction unit, a coordinate feature extraction unit, and an output unit. The image feature extraction unit is used to extract image feature information of the labeled workpiece image. The coordinate feature extraction unit is used to extract coordinate feature information of the image coordinates of each workpiece. The output unit is used to output the three-dimensional spatial coordinates of each workpiece based on the image feature information and the coordinate feature information.

[0054] Specifically, the workpiece coordinate mapping model is a deep learning-based neural network model whose main function is to convert the coordinates of the workpiece image in the labeled workpiece image into actual three-dimensional spatial coordinates. This model consists of an image feature extraction unit, a coordinate feature extraction unit, and an output unit. The image feature extraction unit is used to extract image feature information from the labeled workpiece image. Its specific processing involves using multi-layer convolution operations to extract local features of the image. For example, the first convolutional layer uses 32 3×3 filters with the ReLU activation function to extract edge and texture information of the image. Max pooling is used to reduce the spatial dimension of the feature map, reducing computation while retaining important features. Repeated convolution and pooling operations gradually increase the number of filters to extract higher-level image features. The extracted feature map is flattened into a one-dimensional vector and then compressed and further processed through one or more fully connected layers. For example, a fully connected layer with 256 neurons, combined with the ReLU activation function, is used to extract and output image feature information. The coordinate feature extraction unit is used to extract the coordinate feature information of the image coordinates of each workpiece. The specific implementation process can be as follows: The image coordinates of each workpiece are normalized. The normalized coordinates are then input into a fully connected layer for processing. Using the ReLU activation function, coordinate feature information is extracted and output. The output unit is used to output the three-dimensional spatial coordinates of each workpiece based on the image and coordinate feature information. Alternatively, the image feature vector and coordinate feature vector can be concatenated to form a comprehensive feature vector. This comprehensive feature vector is then further processed through one or more fully connected layers. A fully connected layer with three neurons, each corresponding to one of the workpiece's three-dimensional spatial coordinates, is used. A linear activation function is then used for regression prediction to output the three-dimensional spatial coordinates of each workpiece.

[0055] The model training process can be performed using labeled image data and corresponding 3D spatial coordinates. Labeled data includes labeled workpiece images, workpiece image coordinates, and pre-defined 3D physical coordinates. Mean squared error can be used as the loss function to optimize the model parameters, making the predicted 3D spatial coordinates as close as possible to the actual 3D spatial coordinates. After multiple rounds of training, the workpiece coordinate mapping model is obtained. The workpiece coordinate mapping model extracts image features and coordinate features through image feature extraction units and coordinate feature extraction units, respectively. Then, the output unit fuses these features and predicts the 3D spatial coordinates of the workpiece. This model structure can effectively convert image coordinates into actual 3D spatial coordinates, providing important basic data support for accurate workpiece positioning and subsequent process operation recognition.

[0056] Step 1033: Based on the three-dimensional spatial coordinates of multiple workpieces corresponding to each labeled workpiece image and the background annotation information of multiple workpieces, determine the target workpiece in the multiple workpiece operation video stream information.

[0057] Specifically, for each workpiece operation video stream, the annotated workpiece image is compared with its 3D spatial coordinates from different viewpoints. If the 3D spatial coordinates are consistent across multiple viewpoints, these workpieces are considered to be the same target workpiece. Furthermore, the background annotation information in each annotated workpiece image is analyzed to check if the background information matches the target workpiece's operation scenario. For example, a background information consistency threshold is set; typically, a background information matching degree greater than 80%. If the background information matching degree from different viewpoints is higher than this threshold, these workpieces are considered to be the same target workpiece. By comprehensively considering both 3D spatial coordinate consistency and background information consistency, the target workpiece in the multiple workpiece operation video streams is determined. For example, if the 3D spatial coordinates are consistent across different viewpoints and the background information matches the target workpiece's operation scenario, then the workpiece is identified as the target workpiece. The target workpiece is then marked with identification information to facilitate subsequent workpiece identification.

[0058] Step 104: Input multiple workpiece operation video frames corresponding to the target workpiece into the pre-trained operation recognition model to obtain the operation recognition result output by the operation recognition model.

[0059] The operation recognition results include the workpiece operation category and the corresponding confidence level.

[0060] Specifically, from the workpiece operation video stream information acquired from multiple cameras, the target workpiece and its corresponding multiple workpiece operation video frames have been filtered out using a workpiece detection model. These video frames contain information about the target workpiece and its background. The multiple workpiece operation video frames corresponding to the target workpiece are then input into a pre-trained operation recognition model. This model is a deep learning-based multi-view temporal image analysis model capable of identifying the type of process operation and its confidence level within the video frames. The video frames can be sequences of video frames corresponding to the target workpiece extracted from the video streams of each camera. For example, extracting 5 seconds of video frames before and after the target workpiece detection box and its background information forms a temporally continuous video segment. The video frame sequences from different perspectives are integrated to ensure that the video frame sequences from each perspective are synchronized in time. For instance, the operation recognition model uses a 3D convolutional neural network (3D CNN) to extract features from the video frame sequences of each perspective to capture the spatiotemporal features in the video frames, such as the workpiece's motion trajectory and changes in operational actions. The operation recognition model can employ a lightweight 3D convolutional neural network, such as one containing 5 layers of 3D convolutions (64 3×7×7 convolutional kernels in the first layer) and time-space pooling operations, with fully connected layers at the end to output operation class probabilities. Each viewpoint's 3D CNN branch outputs the probability of the workpiece operation class from that viewpoint. For example, the probability of a bolt tightening operation is 0.85, and the probability of a sealing installation operation is 0.15. A confidence threshold is set, for example, 0.7, to filter out operation classes with confidence levels below the threshold, reducing noise and false recognition. An adaptive fusion strategy is used to weight the operation class probabilities from different viewpoints. For example, the retained viewpoints are weighted according to their confidence levels to calculate the final operation class and its confidence level. For instance, if the operation class probabilities for two viewpoints are 0.85 and 0.9 respectively, the final operation class confidence level is (0.85 + 0.9) / 2 = 0.875. After processing, the operation recognition model outputs the operation recognition result, including the identified process operation class performed on the target workpiece and its corresponding confidence level.

[0061] Optionally, the operation recognition model is trained according to steps 41 to 44.

[0062] Step 41: Obtain the first training workpiece operation video frame group, the first training workpiece operation category corresponding to the first training workpiece operation video frame group, and the first training confidence level.

[0063] The first training workpiece operation video frame group is a continuous set of workpiece operation video frames in the training operation video library. Multiple video frames in the first training workpiece operation video frame group include the target training workpiece. The first training workpiece operation category is the operation category of the target training workpiece. The first training confidence is the confidence of the target training workpiece. The training operation video library includes multiple workpiece operation video frame groups to be trained, as well as the workpiece operation category and the confidence of the target training workpiece corresponding to each workpiece operation video frame group to be trained.

[0064] Specifically, the first training workpiece operation video frame group, the first training workpiece operation category, and the first training confidence level corresponding to the first training workpiece operation video frame group are obtained from the training operation video library.

[0065] Step 42: Input the first training workpiece operation video frame group into the recognition model to be trained, and obtain the training operation category and training confidence corresponding to the first training workpiece operation video frame group output by the recognition model to be trained.

[0066] Specifically, the first training workpiece operation video frame group is input into the recognition model to be trained to obtain the recognition model to be trained, and the training operation category and training confidence are output according to the first training workpiece operation video frame group.

[0067] Step 43: Based on the first training workpiece operation category and the training operation category, as well as the first training confidence and the training confidence, adjust the model parameters of the recognition model to be trained to obtain a new recognition model to be trained.

[0068] Specifically, the difference between the first training workpiece operation category and the training operation category, as well as the difference between the first training confidence and the training confidence, is calculated using a loss function. The model parameters of the recognition model to be trained are then adjusted to obtain a new recognition model to be trained.

[0069] Step 44: Take the second training workpiece operation video frame group as the new first training workpiece operation video frame group, take the second training workpiece operation category corresponding to the second training workpiece operation video frame group as the new first training workpiece operation category, take the second training confidence as the new first training confidence, return to step 42, until the preset iteration end condition is reached, and take the new recognition model to be trained as the operation recognition model.

[0070] The second training workpiece operation video frame group is a workpiece operation video frame group in the training operation video library that has not been input into the recognition model to be trained.

[0071] Specifically, an untrained set of workpiece operation video frames is selected from the training operation video library as the second training workpiece operation video frame set. The second training workpiece operation category corresponding to the second training workpiece operation video frame set is taken as the new first training workpiece operation category, and the second training confidence score is taken as the new first training confidence score. Return to step 42 and input the new first training workpiece operation video frame set into the new recognition model to be trained. Thus, the model parameters are adjusted again based on the obtained training operation category and training confidence score, and the recognition model to be trained is updated. For example, the model parameters to be adjusted can be the learning rate, regularization parameters, etc. The preset iteration termination condition can be reaching a preset number of iterations, the loss value not decreasing significantly in multiple consecutive iterations, or the accuracy on the validation set reaching a preset threshold. When the preset iteration termination condition is reached, the training process ends, and the operation recognition model is obtained.

[0072] The present application's solution involves acquiring multiple workpiece operation video streams; wherein the multiple workpiece operation video streams are video streams from multiple shooting angles, and each workpiece operation video stream includes multiple workpiece operation video frames; each workpiece operation video frame is input into a pre-trained workpiece detection model to obtain an annotated workpiece image output by the workpiece detection model; wherein the annotated workpiece image includes at least one workpiece annotation and corresponding workpiece background annotation; based on the multiple workpiece annotations and multiple workpiece background annotations in the multiple annotated workpiece images of each workpiece operation video stream, a target workpiece is determined in the multiple workpiece operation video streams; the multiple workpiece operation video frames corresponding to the target workpiece are input into a pre-trained operation recognition model to obtain the operation recognition result output by the operation recognition model; wherein the operation recognition result includes the workpiece operation category and the corresponding confidence level. The method of this application annotates the workpiece in the video stream information from multiple shooting angles, identifies the target workpiece, and identifies the operation of the target workpiece to obtain the operation type and confidence level. This avoids the possibility of detection omissions that may occur during manual visual inspection, improves the accuracy of the operation identification results of workpiece installation and the precision of workpiece operation, and thus enhances the reliability and safety of power transmission and transformation equipment manufacturing.

[0073] Figure 2 This is another flowchart illustrating the workpiece operation recognition method provided in this application. This embodiment... Figure 1 Based on the illustrated embodiments and various optional implementation schemes, the steps after obtaining the operation recognition result are described in detail. For example... Figure 2 As shown, the method may include the following steps:

[0074] Step 201: Obtain video stream information of multiple workpiece operations.

[0075] Step 202: Input each workpiece operation video frame into the pre-trained workpiece detection model to obtain the labeled workpiece image output by the workpiece detection model.

[0076] Step 203: Based on the multiple workpiece annotation information and multiple workpiece background annotation information in the multiple annotated workpiece images of each workpiece operation video stream information, determine the target workpiece in the multiple workpiece operation video stream information.

[0077] Step 204: Input multiple workpiece operation video frames corresponding to the target workpiece into the pre-trained operation recognition model to obtain the operation recognition result output by the operation recognition model.

[0078] Step 205: If the confidence level meets the preset confidence level alarm rule, then issue an alarm according to the alarm rule corresponding to the confidence level.

[0079] Specifically, pre-set confidence alarm rules are multiple rules that are pre-defined to trigger an alarm when the confidence level reaches a threshold, used to determine whether the confidence level is within an acceptable range. These pre-set confidence alarm rules can be configured according to the needs of the actual production scenario, and may include multiple confidence thresholds. For example, the lowest acceptable confidence threshold is the low confidence threshold, which can be 60%. If the confidence level is below this threshold, an alarm is triggered. A highest acceptable high confidence threshold, such as 95%, can also be defined. If the confidence level is higher than the high confidence threshold, other types of alarms may be triggered, such as abnormal operation alerts. Alternatively, by analyzing historical operation records, the number of consecutive low-confidence operations can be determined based on historical confidence levels over a period of time and the current confidence level. A threshold for the number of consecutive low-confidence operations can be defined in the pre-set confidence alarm rules, for example, three consecutive times below 60% confidence level. If the confidence level is repeatedly below the operation count threshold, an alarm is triggered. If the confidence level meets any one or more of the pre-set confidence alarm rules, the corresponding alarm is triggered. The alarm rules corresponding to different confidence levels are the alarm measures. Alarm measures can include alerting staff to current confidence level issues by illuminating indicator lights of different colors, such as yellow, orange, and red. For example, activating the orange indicator light when the confidence level falls below 60% three times consecutively. Alarms can also be sounded via buzzer or speaker to alert operators. Alarm information can also be sent to staff via SMS or email for appropriate action. Alarm rules can also include triggering alarms when other anomalies are detected. For example, activating the yellow indicator light when the effective viewing angle of the target workpiece is insufficient (i.e., when the target workpiece detection is abnormal). For example, activating the red indicator light and pausing the production line when workpiece operation recognition fails. All alarm events are automatically saved with 10-second video clips before and after the event to a log database for easy review by staff.

[0080] The solution proposed in this application can monitor the confidence level by pre-setting confidence alarm rules, and trigger an alarm in a timely manner when the confidence level does not meet the requirements. This helps to promptly identify and correct problems in the operation process, thereby further improving production efficiency and product quality.

[0081] Figure 3 This is a schematic diagram of a workpiece operation recognition device provided in this application, which is suitable for executing the countdown prediction method provided in this application. Figure 3 As shown, the device may specifically include:

[0082] The acquisition module 301 is used to acquire multiple workpiece operation video stream information; wherein, the multiple workpiece operation video stream information is video stream information from multiple shooting angles, and each workpiece operation video stream information includes multiple workpiece operation video frames.

[0083] The first input module 302 is used to input each of the workpiece operation video frames into a pre-trained workpiece detection model to obtain a labeled workpiece image output by the workpiece detection model; wherein, the labeled workpiece image includes at least one workpiece annotation information and workpiece background annotation information corresponding to the workpiece annotation information.

[0084] The determination module 303 is used to determine the target workpiece in the plurality of workpiece operation video stream information based on the plurality of workpiece annotation information and the plurality of workpiece background annotation information in the plurality of annotated workpiece images of each workpiece operation video stream information.

[0085] The second input module 304 is used to input multiple workpiece operation video frames corresponding to the target workpiece into a pre-trained operation recognition model to obtain the operation recognition result output by the operation recognition model; wherein, the operation recognition result includes the workpiece operation category and the corresponding confidence level.

[0086] In one embodiment, the determining module 303 is specifically configured to: determine the image coordinates of each workpiece based on the multiple workpiece annotation information corresponding to each labeled workpiece image; for each labeled workpiece image, input the labeled workpiece image and the image coordinates of each workpiece into a pre-trained workpiece coordinate mapping model to obtain the three-dimensional spatial coordinates of each workpiece output by the workpiece coordinate mapping model; and determine the target workpiece in the multiple workpiece operation video stream information based on the three-dimensional spatial coordinates of the multiple workpieces corresponding to each labeled workpiece image and the multiple workpiece background annotation information.

[0087] In one embodiment, the workpiece coordinate mapping model determined by module 303 includes an image feature extraction unit, a coordinate feature extraction unit, and an output unit; the image feature extraction unit is used to extract image feature information of the labeled workpiece image; the coordinate feature extraction unit is used to extract coordinate feature information of the image coordinates of each workpiece; and the output unit is used to output the three-dimensional spatial coordinates of each workpiece based on the image feature information and the coordinate feature information.

[0088] In one embodiment, the workpiece detection model of the first input module 302 is trained according to the following steps: acquiring a first training workpiece operation video frame and a first training labeled image; wherein, the first training workpiece operation video frame is an image in a training video frame library, the training video frame library includes multiple workpiece operation video frames to be trained and a corresponding labeled image to be trained for each workpiece operation video frame to be trained, the labeled image to be trained being the image after the workpiece label anchor box in the workpiece operation video frame to be trained; inputting the first training workpiece operation video frame into the detection model to be trained, obtaining the training labeled workpiece image corresponding to the first training workpiece operation video frame output by the detection model to be trained; and based on the first training labeled image and the... The training labeled workpiece images are used to adjust the model parameters of the detection model to be trained, resulting in a new detection model to be trained. The second training workpiece operation video frame is used as the new first training workpiece operation video frame, and the second training labeled image corresponding to the second training workpiece operation video frame is used as the new first training labeled image. The process of "inputting the first training workpiece operation video frame into the detection model to be trained, and obtaining the training labeled workpiece image corresponding to the first training workpiece operation video frame output by the detection model to be trained" is repeated until a preset iteration end condition is met. The new detection model to be trained is then used as the workpiece detection model. The second training workpiece operation video frame is an untrained workpiece operation video frame in the training video frame library.

[0089] In one embodiment, the detection model to be trained in the first input module 302 includes a feature extraction network, a feature aggregation network, and a detection head; the feature extraction network is used to extract multiple feature maps from the first training workpiece operation video frames; the feature aggregation network is used to fuse the multiple feature maps to generate a fused feature map; the detection head is used to obtain the workpiece existence probability, detection box coordinates, and workpiece detection confidence based on the fused feature map, and to annotate the workpiece based on the workpiece existence probability, detection box coordinates, and workpiece detection confidence, and output the trained annotated workpiece image.

[0090] In one embodiment, the operation recognition model of the second input module 304 is trained according to the following steps: acquiring a first training workpiece operation video frame group, and a first training workpiece operation category and a first training confidence level corresponding to the first training workpiece operation video frame group; wherein, the first training workpiece operation video frame group is a continuous set of workpiece operation video frames in the training operation video library, multiple video frames in the first training workpiece operation video frame group all include the target training workpiece, the first training workpiece operation category is the operation category of the target training workpiece, the first training confidence level is the confidence level of the target training workpiece, the training operation video library includes multiple workpiece operation video frame groups to be trained and the workpiece operation category and training confidence level corresponding to each workpiece operation video frame group to be trained; inputting the first training workpiece operation video frame group into the recognition model to be trained, and obtaining the training operation category and training confidence level corresponding to the first training workpiece operation video frame group output by the recognition model to be trained. The model parameters of the recognition model to be trained are adjusted according to the first training workpiece operation category and the training operation category, as well as the first training confidence and the training confidence, to obtain a new recognition model to be trained. The second training workpiece operation video frame group is used as the new first training workpiece operation video frame group, the second training workpiece operation category corresponding to the second training workpiece operation video frame group is used as the new first training workpiece operation category, and the second training confidence is used as the new first training confidence. The process of "inputting the first training workpiece operation video frame group into the recognition model to be trained, and obtaining the training operation category and training confidence corresponding to the first training workpiece operation video frame group output by the recognition model to be trained" is repeated until the preset iteration end condition is met. The new recognition model to be trained is then used as the operation recognition model. The second training workpiece operation video frame group is a training workpiece operation video frame group in the training operation video library that has not been input into the recognition model to be trained.

[0091] In one embodiment, the device further includes an alarm module, which, after the second input module 304 obtains the operation recognition result output by the operation recognition model, if the confidence level meets the preset confidence alarm rule, then issues an alarm according to the alarm rule corresponding to the confidence level.

[0092] The apparatus of this application acquires multiple workpiece operation video streams; wherein the multiple workpiece operation video streams are video streams from multiple shooting angles, and each workpiece operation video stream includes multiple workpiece operation video frames; each workpiece operation video frame is input into a pre-trained workpiece detection model to obtain an annotated workpiece image output by the workpiece detection model; wherein the annotated workpiece image includes at least one workpiece annotation information and workpiece background annotation information corresponding to the workpiece annotation information; based on the multiple workpiece annotation information and multiple workpiece background annotation information in the multiple annotated workpiece images of each workpiece operation video stream, a target workpiece in the multiple workpiece operation video streams is determined; the multiple workpiece operation video frames corresponding to the target workpiece are input into a pre-trained operation recognition model to obtain an operation recognition result output by the operation recognition model; wherein the operation recognition result includes the workpiece operation category and the corresponding confidence level. The method of this application annotates the workpiece in the video stream information from multiple shooting angles, identifies the target workpiece, and identifies the operation of the target workpiece to obtain the operation type and confidence level. This avoids the possibility of detection omissions that may occur during manual visual inspection, improves the accuracy of the operation identification results of workpiece installation and the precision of workpiece operation, and thus enhances the reliability and safety of power transmission and transformation equipment manufacturing.

[0093] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the workpiece operation recognition method provided in any of the above embodiments.

[0094] This application also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the workpiece operation recognition method provided in any of the above embodiments.

[0095] The following is for reference. Figure 4 It shows a schematic diagram of the structure of an electronic device 400 suitable for implementing the present application. Figure 4 The electronic device shown is merely an example and should not impose any limitations on the functionality and scope of this application.

[0096] like Figure 4 As shown, the electronic device 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage section 408 into a random access memory (RAM) 403. The RAM 403 also stores various programs and data required for the operation of the electronic device 400. The CPU 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0097] The following components are connected to I / O interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, modem, etc. The communication section 409 performs communication processing via a network such as the Internet. Drive 410 is also connected to I / O interface 405 as needed. Removable media 411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 410 as needed so that computer programs read from them can be installed into storage section 408 as needed.

[0098] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the computer program is executed by central processing unit (CPU) 401, it performs the functions defined above in the system of this application.

[0099] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0100] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0101] The modules and / or units described in this application can be implemented in software or hardware. The described modules and / or units can also be housed in a processor; for example, a processor may be described as including an acquisition module, a first input module, a determination module, and a second input module. The names of these modules do not necessarily limit the functionality of the module itself.

[0102] In another aspect, this application also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more programs that, when executed by the device, cause the device to perform the following operations:

[0103] Multiple workpiece operation video streams are acquired; these multiple workpiece operation video streams are video streams from multiple shooting angles, and each workpiece operation video stream includes multiple workpiece operation video frames; each workpiece operation video frame is input into a pre-trained workpiece detection model to obtain an annotated workpiece image output by the workpiece detection model; the annotated workpiece image includes at least one workpiece annotation and corresponding workpiece background annotation; based on the multiple workpiece annotations and multiple workpiece background annotations in the multiple annotated workpiece images of each workpiece operation video stream, the target workpiece in the multiple workpiece operation video streams is determined; the multiple workpiece operation video frames corresponding to the target workpiece are input into a pre-trained operation recognition model to obtain the operation recognition result output by the operation recognition model; the operation recognition result includes the workpiece operation category and the corresponding confidence level.

[0104] According to the technical solution of this application, multiple workpiece operation video streams are obtained; wherein, the multiple workpiece operation video streams are video streams from multiple shooting angles, and each workpiece operation video stream includes multiple workpiece operation video frames; each workpiece operation video frame is input into a pre-trained workpiece detection model to obtain an annotated workpiece image output by the workpiece detection model; wherein, the annotated workpiece image includes at least one workpiece annotation information and workpiece background annotation information corresponding to the workpiece annotation information; based on the multiple workpiece annotation information and multiple workpiece background annotation information in the multiple annotated workpiece images of each workpiece operation video stream, the target workpiece in the multiple workpiece operation video streams is determined; the multiple workpiece operation video frames corresponding to the target workpiece are input into a pre-trained operation recognition model to obtain the operation recognition result output by the operation recognition model; wherein, the operation recognition result includes the workpiece operation category and the corresponding confidence level. The method of this application annotates the workpiece in the video stream information from multiple shooting angles, identifies the target workpiece, and identifies the operation of the target workpiece to obtain the operation type and confidence level. This avoids the possibility of detection omissions that may occur during manual visual inspection, improves the accuracy of the operation identification results of workpiece installation and the precision of workpiece operation, and thus enhances the reliability and safety of power transmission and transformation equipment manufacturing.

[0105] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the workpiece operation recognition method provided in any embodiment of this application.

[0106] In the implementation of the computer program product, computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0107] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.

[0108] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A workpiece operation recognition method characterized by, The method comprises: obtaining a plurality of workpiece operation video stream information; wherein the plurality of workpiece operation video stream information is video stream information of a plurality of shooting angles, each of the workpiece operation video stream information comprises a plurality of workpiece operation video frames; inputting each of the workpiece operation video frames into a pre-trained workpiece detection model to obtain a labeled workpiece image output by the workpiece detection model after labeling the workpiece; wherein the labeled workpiece image comprises at least one workpiece labeling information and workpiece background labeling information corresponding to the workpiece labeling information; determining a target workpiece in the plurality of workpiece operation video stream information according to a plurality of workpiece labeling information and a plurality of workpiece background labeling information in a plurality of labeled workpiece images of each of the workpiece operation video stream information; inputting a plurality of workpiece operation video frames corresponding to the target workpiece into a pre-trained operation recognition model to obtain an operation recognition result output by the operation recognition model; wherein the operation recognition result comprises a workpiece operation category and a corresponding confidence.

2. The method of claim 1, wherein, The method comprises: determining image coordinates of each workpiece according to a plurality of workpiece labeling information corresponding to each of the labeled workpiece images; for each of the labeled workpiece images, inputting the labeled workpiece image and the image coordinates of each workpiece into a pre-trained workpiece coordinate mapping model to obtain three-dimensional space coordinates of each workpiece output by the workpiece coordinate mapping model; determining a target workpiece in the plurality of workpiece operation video stream information according to a plurality of workpiece labeling information and a plurality of workpiece background labeling information corresponding to a plurality of workpieces of each of the labeled workpiece images.

3. The method of claim 2, wherein, The workpiece coordinate mapping model comprises an image feature extraction unit, a coordinate feature extraction unit and an output unit; the image feature extraction unit is configured to extract image feature information of the labeled workpiece image; the coordinate feature extraction unit is configured to extract coordinate feature information of the image coordinates of each workpiece; the output unit is configured to output the three-dimensional space coordinates of each workpiece according to the image feature information and the coordinate feature information.

4. The method of claim 1, wherein, The workpiece detection model is trained according to the following steps: obtaining a first training workpiece operation video frame and a first training labeled image; wherein the first training workpiece operation video frame is an image in a training video frame library, the training video frame library comprises a plurality of to-be-trained workpiece operation video frames and a to-be-trained labeled image corresponding to each of the to-be-trained workpiece operation video frames, and the to-be-trained labeled image is an image after a workpiece labeling anchor box in the to-be-trained workpiece operation video frame is labeled; inputting the first training workpiece operation video frame into a to-be-trained detection model to obtain a training labeled workpiece image corresponding to the first training workpiece operation video frame output by the to-be-trained detection model; Adjust model parameters of the to-be-trained detection model according to the first training annotation image and the training annotation workpiece image, to obtain a new to-be-trained detection model; Take a second training workpiece operation video frame as a new first training workpiece operation video frame, and take a second training annotation image corresponding to the second training workpiece operation video frame as a new first training annotation image, and return to execute the step of "inputting the first training workpiece operation video frame into the to-be-trained detection model to obtain a training annotation workpiece image corresponding to the first training workpiece operation video frame output by the to-be-trained detection model", until a preset iteration end condition is reached, and take the new to-be-trained detection model as the workpiece detection model; wherein the second training workpiece operation video frame is an untrained workpiece operation video frame in the training video frame library.

5. The method of claim 4, wherein, The to-be-trained detection model comprises a feature extraction network, a feature aggregation network and a detection head; The feature extraction network is configured to extract a plurality of feature maps from the first training workpiece operation video frame; The feature aggregation network is configured to perform feature fusion on the plurality of feature maps to generate a fused feature map; The detection head is configured to obtain a workpiece existence probability, a detection box coordinate and a workpiece detection confidence according to the fused feature map, and to annotate a workpiece according to the workpiece existence probability, the detection box coordinate and the workpiece detection confidence, and to output the training annotation workpiece image.

6. The method of claim 1, wherein, The operation recognition model is trained according to the following steps: Obtain a first training workpiece operation video frame group, a first training workpiece operation category corresponding to the first training workpiece operation video frame group and a first training confidence; wherein the first training workpiece operation video frame group is a group of continuous workpiece operation video frames in a training operation video library, a plurality of video frames in the first training workpiece operation video frame group each comprise a target training workpiece, the first training workpiece operation category is an operation category of the target training workpiece, and the first training confidence is a confidence of the target training workpiece; the training operation video library comprises a plurality of to-be-trained workpiece operation video frame groups and a to-be-trained workpiece operation category and a to-be-trained confidence corresponding to each to-be-trained workpiece operation video frame group; Input the first training workpiece operation video frame group into a to-be-trained recognition model to obtain a training operation category and a training confidence corresponding to the first training workpiece operation video frame group output by the to-be-trained recognition model; Adjust model parameters of the to-be-trained recognition model according to the first training workpiece operation category and the training operation category, and the first training confidence and the training confidence, to obtain a new to-be-trained recognition model; The second training workpiece operation video frame set is taken as a new first training workpiece operation video frame set, the second training workpiece operation category corresponding to the second training workpiece operation video frame set is taken as a new first training workpiece operation category, the second training confidence is taken as a new first training confidence, and the step of inputting the first training workpiece operation video frame set into the to-be-trained identification model to obtain the training operation category and the training confidence corresponding to the first training workpiece operation video frame set output by the to-be-trained identification model is returned until a preset iteration end condition is reached, and the new to-be-trained identification model is taken as the operation identification model; wherein the second training workpiece operation video frame set is a to-be-trained workpiece operation video frame set in the training operation video library that has not been input into the to-be-trained identification model.

7. The method of claim 1, wherein, The method further comprises: If the confidence meets a preset confidence alarm rule, an alarm is performed according to the alarm rule corresponding to the confidence.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the workpiece operation identification method in any one of claims 1 to 7.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the workpiece operation identification method in any one of claims 1 to 7.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the workpiece operation identification method in any one of claims 1 to 7. The computer program is executed by the processor to implement the workpiece operation identification method in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Behavior identification method and device for multi-category engineering vehicle

    CN112800934A

  • Video action recognition method and system for smart factory

    CN114898466A

  • Identification method and device, equipment and computer readable storage medium

    CN114926766A

  • Visual inspection method, device and system for lean assembly process

    CN117423043A