Operation sequence detection method based on CV large model

By using a large CV model-based job sequence detection method, the error problem introduced by manual marking in the battery management system board insertion operation is solved, achieving high-accuracy job sequence identification and anomaly handling without human intervention, and improving the reliability of job supervision.

CN120877199AActive Publication Date: 2025-10-31INPAI BATTERY TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511394248.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-10-31
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

In the existing technology, the operation sequence supervision of battery management system board insertion operations relies on manual marking, which has problems such as low efficiency, large human error, and high recognition error rate, making it difficult to achieve reliable operation sequence recognition.

Method used

A job sequence detection method based on a large visual model is adopted. By dividing the video of the job to be identified into multiple sub-videos, the feature extraction and attention mechanism of the large visual model are used to identify the job positions of the operation, determine the job sequence, reduce human intervention, and improve the accuracy and reliability of recognition.

Benefits of technology

It achieves unmanned job sequence recognition, reduces human error, improves the accuracy and reliability of job sequence recognition, and can promptly detect and handle execution anomalies, reducing the impact of job errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877199A_ABST
    Figure CN120877199A_ABST
Patent Text Reader

Abstract

The invention provides an operation sequence detection method based on a CV large model, and the method comprises the steps: dividing a to-be-recognized operation video into a plurality of sub-videos, and enabling the number of data of the sub-videos to be not greater than the number of operation workstations; for each segment of the sub-video, identifying each segment of the sub-video, and inputting the identified segment of the sub-video into a visual large model; performing feature extraction through a feature extraction module of the visual large model to obtain video features; identifying a work position of the operation represented in the sub-video through an attention mechanism of the visual large model; and on the basis of the time sequence of each segment of the sub-video in the to-be-identified operation video and the operation working position corresponding to the sub-video, determining the current operation completion sequence of each operation working position. According to the method, the reliability of monitoring and identification of the standard operation program can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of detection technology, and more specifically, to a method for detecting job sequence based on a large CV model. Background Technology

[0002] In various operational environments, such as battery management system (BMS) board insertion operations, there are specific requirements for the order of operations in each step. Incorrect sequence may lead to malfunctions. Therefore, there is a need to monitor the operational sequence. Some methods involve dedicated supervisors; others rely on video surveillance. During monitoring, operators need to set a marker after each step to identify the current stage. The first method is inefficient and relatively resource-intensive; the second method depends on operator-mandated marking, which is susceptible to errors due to forgotten markings or inconsistent marker placement. Summary of the Invention

[0003] The purpose of this application is to provide a job sequence detection method based on a large CV model, which can improve the reliability of standard operating procedure monitoring and identification.

[0004] In a first aspect, the present invention provides a job sequence detection method based on a large-scale visual model, comprising: dividing a job video to be identified into multiple sub-videos, wherein the number of data in the sub-videos is no greater than the number of job positions; for each sub-video, identifying and inputting each sub-video into a large-scale visual model; extracting features through the feature extraction module of the large-scale visual model to obtain video features; identifying the job positions of the operations represented in the sub-videos through the attention mechanism of the large-scale visual model; and determining the current operation completion order of each job position based on the temporal order of each sub-video in the job video to be identified and the job positions corresponding to the sub-videos.

[0005] In the above implementation method, no human intervention is required during the overall identification of the job sequence. The job sorting and job sequence identification are achieved solely based on the obtained recognition to identify the corresponding actions performed in each video segment. This reduces the influence of human subjective factors and makes the identified sequence more reliable.

[0006] In an optional implementation, dividing the video of the operation to be identified into multiple sub-videos includes: starting from the first frame of the video of the operation to be identified, sequentially identifying the frame in each frame of the video of the operation to be identified where the starting operation state is first identified as the starting frame; wherein, the starting operation state includes the operation subject being located at the starting work position, and the starting work position is one of all operation work positions; dividing each frame of the video of the operation to be identified after the starting frame into multiple sub-videos.

[0007] In the above implementation method, recognition can be performed on the starting frame first, which can reduce the number of images that need to be processed in the actual recognition of sub-videos, reduce the influence of some interference images, improve the accuracy of action recognition, and further improve the accuracy of job sequence detection.

[0008] In an optional implementation, dividing the images after the starting frame in the video to be identified into multiple sub-videos includes: taking the starting frame as the starting point, sequentially comparing the subsequent images of the starting frame with the starting frame to determine multiple consecutive images with a similarity greater than a similarity threshold, and using them as the first sub-video; taking the image adjacent to the i-th sub-video and with a similarity not greater than the similarity threshold as the unit starting frame of the (i+1)-th sub-video, sequentially comparing the subsequent images of the unit starting frame with the unit starting frame to determine multiple consecutive images with a similarity greater than the similarity threshold, and using them as the (i+1)-th sub-video, where i is a positive integer not greater than N-1, and N is the number of workstations in the job.

[0009] In the above implementation, since the similarity of the images of operations at the same workstation is significantly higher than the similarity of the images of operations at different workstations, the determination of each sub-video segment can be based on the similarity of the images to effectively divide the videos of different work stages, thereby more accurately identifying the operations in each time period and the workstations involved in the operations.

[0010] In an optional implementation, the step of sequentially comparing the similarity of subsequent images of the starting frame with that of the starting frame includes: extracting the optical flow information of subsequent images of the starting frame and the starting frame; calculating the frequency domain histograms of subsequent images of the starting frame and the starting frame; and comparing the optical flow information and frequency domain histograms of subsequent images of the starting frame with the optical flow information and frequency domain histograms of the starting frame.

[0011] In an optional implementation, the workstation includes multiple board interfaces of the battery management system; the method further includes: identifying each sub-video segment to determine whether each board interface is in a plug-in operation state; determining the current operation completion order of each workstation based on the time sequence of each sub-video segment in the work video to be identified and the workstation corresponding to the sub-video includes: determining the current operation completion order of each board interface based on the time sequence of each sub-video segment in the work video to be identified, the board interfaces corresponding to each sub-video, and the plug-in operation state of each board interface.

[0012] The above implementation not only identifies the order of tasks but also identifies the results of tasks, thereby eliminating cases where the order of tasks is correct but the actual tasks are not completed, thus presenting the completion status of tasks more accurately and comprehensively.

[0013] In an optional implementation, the method further includes: comparing the current operation completion order with a preset execution order to determine whether there is an execution abnormality; and outputting alarm information if there is an execution abnormality.

[0014] In the above implementation method, the output prompt can also be combined with whether the sequence is abnormal, so that the abnormality can reach the relevant personnel better and reduce the overall impact of the operation error.

[0015] In an optional implementation, the alarm information includes the timestamp of the operation video used in the current operation completion sequence; the method further includes: in the event of an execution abnormality, retrieving the original video corresponding to the current operation completion sequence from the video database based on the timestamp; transmitting the original video to a designated terminal; and generating a processing work order for the execution abnormality after receiving an abnormality confirmation message from the designated terminal.

[0016] In the above implementation method, it is also possible to generate processing work orders based on abnormal situations, which can enable relevant personnel to respond to abnormalities more quickly.

[0017] In a second aspect, the present invention provides an electronic device, comprising: a processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the machine-readable instructions are executed by the processor to perform the steps of the method described in any of the foregoing embodiments.

[0018] Thirdly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the method described in any of the foregoing embodiments.

[0019] Fourthly, the present invention provides a computer program product, the computer program product comprising a computer program, which, when executed by a processor, implements the method described in any one of the foregoing embodiments. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 A block diagram illustrating an electronic device provided in an embodiment of this application; Figure 2 A flowchart of the job sequence detection method provided in the embodiments of this application; Figure 3 A schematic diagram of a job scenario for the job sequence detection method provided in this application embodiment; Figure 4 This is a partial flowchart of the job sequence detection method provided in the embodiments of this application; Figure 5 Another schematic flowchart of the job sequence detection method provided in the embodiments of this application. Detailed Implementation

[0022] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0023] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0024] To improve the accuracy of various work processes, existing technical solutions utilize video surveillance to monitor the sequence of battery management system (BMS) board insertion operations. To analyze the accuracy of the sequence in the video, a computer vision-based processing flow is typically employed. The core idea is to process the input video stream frame by frame, utilizing its efficient convolutional neural network structure to detect and accurately locate the BMS board target in the frame in real time, outputting its bounding box coordinates. After board location is completed, operator intervention is required. The operator manually marks the completion point or key location of specific insertion actions (such as connector insertion) in the video frame; this mark is typically represented by a coordinate point or a small area. By calculating the relative spatial relationship between the manually marked completion point and the BMS board target, the operator's insertion operation position is determined, ultimately inferring and calculating the specific sequence of multiple BMS insertion actions performed by the operator in the video stream. The key to this approach lies in the cooperation between video recognition and human intervention to verify the standardization of manual operations.

[0025] However, this method, which requires personnel cooperation, has some limitations, including the following shortcomings: 1) Relying on manually marked scribing lines as action completion markers, but in practical applications there are problems with uncontrollable scribing quality. For example, at the physical level: insufficient scribing clarity, poor reflective properties, too short length or local wear, making it difficult for the YOLOv8 algorithm to extract effective features. For example, at the environmental level: factors such as changes in workshop lighting and oil contamination further reduce the visual salience of scribing markings.

[0026] 2) Dynamic occlusions generated by operators during operation constitute systematic detection blind spots. For example, when hands / tools obscure the scribing area, static occlusions formed by dense wire bundles or equipment parts cause fragmentation of the visible scribing area, exceeding the minimum effective recognition threshold of the target detection algorithm.

[0027] 3) Operators rely on subjective judgment to mark, introducing a chain of human error. For example, spatial positioning deviation: mismarking non-related areas (such as adjacent boards or background areas) in dynamic video frames; timing judgment error: marking in the intermediate state before the insertion action is completed, or marking after the action is completed; there are also individual differences in the judgment criteria of different operators for "completed state".

[0028] 4) The above defects form a cascading failure path, which ultimately leads to an exponential increase in the error rate of action sequence analysis with the complexity of the process, severely restricting the practical application value of the system in the industrial field.

[0029] Based on the above research, the embodiments of this application can provide a job sequence detection method, electronic device, and computer-readable storage medium that can achieve reliable monitoring and identification of standard operating procedures without human intervention.

[0030] To facilitate understanding of this embodiment, the electronic device that performs the job sequence detection method disclosed in this application embodiment will first be described in detail.

[0031] like Figure 1 The diagram shown is a block illustration of an electronic device. The electronic device 100 may include a memory 111 and a processor 113. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the electronic device 100. For example, the electronic device 100 may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0032] The memory 111 and processor 113 described above are electrically connected to each other directly or indirectly to enable data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines. The processor 113 described above is used to execute executable modules stored in the memory.

[0033] The memory 111 can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory 111 stores programs, and the processor 113 executes these programs upon receiving execution instructions. The methods executed by the electronic device 100 as defined in any embodiment of this application can be applied to the processor 113, or implemented by the processor 113.

[0034] The aforementioned processor 113 may be an integrated circuit chip with signal processing capabilities. The processor 113 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a digital signal processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor.

[0035] To facilitate users in obtaining the detection results of the job sequence detection method, the electronic device may also include an input / output unit and a display unit.

[0036] Input / output units can be used to provide users with input data. Input / output units can be, but are not limited to, mice and keyboards.

[0037] The display unit provides an interactive interface (e.g., a user interface) between the electronic device 100 and the user, or displays image data for the user's reference. In this embodiment, the display unit can be a liquid crystal display (LCD) or a touch display. If it is a touch display, it can be a capacitive touchscreen or a resistive touchscreen that supports single-point and multi-point touch operations. Supporting single-point and multi-point touch operations means that the touch display can sense touch operations generated simultaneously from one or more locations on the touch display and pass the sensed touch operations to the processor for calculation and processing.

[0038] The electronic device 100 in this embodiment can be used to execute various steps in the various methods provided in the embodiments of this application. The implementation process of the job sequence detection method is described in detail below through several embodiments.

[0039] Please see Figure 2 This is a flowchart of a job sequence detection method provided in an embodiment of this application. The job sequence detection method provided in this application can be applied to an electronic device, which then executes the steps in the job sequence detection method. The following will describe... Figure 2 The specific process shown will be explained in detail.

[0040] Step 210: Divide the video to be identified into multiple sub-videos.

[0041] The number of sub-video data is no greater than the number of job workstations.

[0042] The specific workstations required for detecting the sequence of operations vary depending on the field, such as industry, medical, and service sectors. For example, if the detection needs to be performed on the sequence of battery management system (BMS) board insertion operations, the workstation could be the interface of the BMS board. Similarly, if the detection needs to be performed on the sequence of surgical procedures, the workstation could be different parts of the patient's body. Likewise, if the detection needs to be performed on the assembly sequence of automotive parts, the workstation could be different parts of the car. Finally, if the detection needs to be performed on the assembly of components based on molds, the workstation could be different positions on the molds.

[0043] Rules for dividing videos into multiple segments can include operations performed at the same workstation, where consecutive frames can be grouped into one sub-video segment. For example, if two frames represent operations performed at the same workstation, but are separated by other action frames (e.g., operations performed at other workstations), they will not be grouped into the same sub-video segment even if they are highly similar and at the same workstation. Similarly, if two frames are adjacent, but the operations in those frames are performed at different workstations, they will not be grouped into the same sub-video segment even if they are adjacent.

[0044] Step 220: For each sub-video segment, identify and input each sub-video segment into the large visual model.

[0045] This large-scale visual model can be a large-scale visual language model, such as the qwen2.5-vl model.

[0046] Step 230: Feature extraction is performed using the feature extraction module of the visual large model to obtain video features.

[0047] Step 240: Identify the job position of the operation represented in the sub-video through the attention mechanism of the visual big model.

[0048] For example, attention mechanisms can be used to focus on the plug-in contact points.

[0049] In this embodiment, the location of the operation is different for different workstations. Therefore, the specific workstation can be determined based on the location of the operation.

[0050] For example, in step 210, the video of the task to be identified is divided into M sub-video segments, and steps 220 to 240 can then identify M task workstations. M is a positive integer, and specifically, the value of M varies depending on the content to be detected. For example, if the sequence of battery management system board insertion operations needs to be detected, then the value of M can be 7.

[0051] Alternatively, a neural network model can be used to identify each video segment in order to determine the work position operated by each video segment.

[0052] In this embodiment, the large visual model can output the operation's workstation and the result status of that workstation. 0 indicates that the workstation in the current screen is in an incomplete operation state; 1 indicates that the workstation in the current screen is in a completed operation state. For example, taking the battery management system board as an example, 0 indicates that the connector in the current screen is in an unconnected operation state; 1 indicates that the connector in the current screen is in a connected operation state.

[0053] Step 250: Based on the temporal order of each sub-video segment in the video to be identified and the corresponding workstation of the sub-video, determine the current operation completion order of each workstation.

[0054] Optionally, steps 220 to 240 above can be identified sequentially according to the time order of each sub-video, and the order of the results can be used to determine the current operation completion order of each job workstation.

[0055] Optionally, each sub-video segment can be set with a time tag. The work positions identified in each sub-video segment are sorted based on their time tags to obtain the order in which the operation is completed.

[0056] The above implementation method allows for the pre-division of the video of the task to be identified into multiple segments. By recognizing the actions in each segment, the relevant workstations in the video can be determined. The chronological order is then used to determine the operational sequence for each workstation. This reduces operator intervention and allows for direct recognition based on the operated equipment, improving the objectivity and reliability of the recognition results.

[0057] To reduce the impact of interfering video frames on the recognition results, the influence of some interfering video frames can be removed during the sub-video segmentation stage. Step 210 above may include: Step 211: Starting from the first frame of the video to be identified, sequentially identify each frame of the video to be identified, and take the first frame of the starting operation state as the starting frame.

[0058] The initial operation status includes the operation subject being located at the starting work position, which is one of all operation work positions.

[0059] Alternatively, an object detection algorithm can be used to identify the starting frame. For example, the object detection algorithm could be the You Only Look Once version 8 (YOLOv8) object detection algorithm.

[0060] Optionally, to reduce processing load, in step 211, only a specified number of images per second can be extracted for identifying the starting frame. This specified number can be a pre-set value, for example, 5, 8, 10, etc.

[0061] Alternatively, OpenCV and FFmpeg can be used to implement frame extraction. For example, frame extraction can be achieved by extracting 5 frames per second.

[0062] In one use case, the starting frame can be the first frame in which the device to be operated appears. For example, the device to be operated could be a battery management system. The model can then be fine-tuned and trained based on the YOLOv8 algorithm, and the battery management system can be identified using the trained YOLOv8 algorithm. For example, the YOLOv8 algorithm can be configured with: channel = 3; number of target categories = 1 (identifying the battery management system board); training iterations EPOCHS = 1300; batch size = 16; and learning rate = 1e. -3 The learning rate decay momentum lrMomentum = 0.9 is applied to the Nesterovs updater.

[0063] In one use case, the starting frame can be the first frame in which an action is performed on the device to be operated. For example, the device to be operated could be a battery management system. The YOLOv8 algorithm's channel count is set to 3; the number of target categories to be detected is set to 1 (to identify the battery management system's plug-in preparatory action); the number of training iterations is EPOCHS = 1300; the batch size is 16; and the learning rate is 1e. -3 The learning rate decay momentum lrMomentum = 0.9 is applied to the Nesterovs updater.

[0064] In one example, detecting a battery management system (BMS) board insertion operation requires recognizing the preparatory actions for the BMS insertion; that is, the image must show the operator's hands and the BMS board's insertion interface. For example... Figure 3As shown in the figure, the operator's hands are in the position of operating the interface of the battery management system board.

[0065] The parameter values ​​used in training the YOLOv8 algorithm mentioned above are merely illustrative; some parameters can be adjusted adaptively. For example, parameters such as the number of training iterations (EPOCHS), batch size (batchSize), learning rate (learningRate), and learning rate decay momentum (lrMomentum) can be adjusted adaptively. The two use cases described above are simply examples.

[0066] In other scenarios, if the content of the starting frame to be identified is different, the number of target categories (classes) mentioned above can also take different values. For example, in a usage scenario, it is necessary to identify the device to be operated and also the tool that operates the device; in this case, the number of target categories (classes) can be 2.

[0067] Step 212: Divide the images of each frame after the starting frame in the video to be identified into multiple sub-video segments.

[0068] After the starting frame is identified, the images before the starting frame of the video to be identified can be deleted to divide the images after the starting frame in the video to be identified.

[0069] For example, based on the identification of each frame of the video of the job to be identified, if it is determined that the operation is for the same job workstation and the images are consecutive, they can be divided into the same sub-video.

[0070] In this embodiment, step 212 may include steps 2121 and 2122.

[0071] Step 2121: Starting from the starting frame, the subsequent images of the starting frame are compared with the starting frame in sequence to determine the multiple consecutive frames that have a similarity greater than the similarity threshold with the starting frame, and these frames are used as the first sub-video.

[0072] For example, the similarity can include optical flow information and a frequency domain histogram. If both the optical flow information and the frequency domain histogram have high similarity, then the two frames are considered as images with a similarity greater than a similarity threshold. For instance, the similarity threshold can include multiple thresholds, with one threshold set for each piece of information being compared. For example, if the compared information includes optical flow information and a frequency domain histogram, the similarity threshold can include both an optical flow threshold and a histogram threshold.

[0073] Optionally, step 2121 above may include: extracting subsequent images of the starting frame and optical flow information of the starting frame; calculating the frequency domain histograms of subsequent images of the starting frame and the starting frame; and comparing the optical flow information and frequency domain histograms of subsequent images of the starting frame with the optical flow information and frequency domain histograms of the starting frame.

[0074] For example, the optical flow information of subsequent images of the starting frame can be compared with the optical flow information of the starting frame. If the similarity is greater than the optical flow threshold, it can be determined that the similarity of the optical flow information of the two frames meets the requirements.

[0075] For example, the frequency domain histogram of subsequent images of the starting frame can be compared with the frequency domain histogram of the starting frame. If the similarity is greater than the histogram threshold, it can be determined that the frequency domain histogram similarity of the two frames meets the requirements.

[0076] If both optical flow information similarity and frequency domain histogram similarity meet the requirements, it can be determined that the similarity between the corresponding subsequent image and the starting frame is greater than the similarity threshold.

[0077] Step 2122: Take the image that is adjacent to the i-th sub-video and whose similarity is not greater than the similarity threshold as the unit starting frame of the (i+1)-th sub-video. Then, compare the similarity of the subsequent images of the unit starting frame with the unit starting frame in turn to determine the multiple consecutive frames that have a similarity greater than the similarity threshold with the unit starting frame and take them as the (i+1)-th sub-video.

[0078] Where i is a positive integer and not greater than N-1, and N is the number of workstations in the job.

[0079] Optionally, step 2122 above may include: extracting subsequent images of the unit start frame and optical flow information of the unit start frame; calculating the frequency domain histogram of the subsequent images of the unit start frame and the unit start frame; and comparing the optical flow information and frequency domain histogram of the subsequent images of the unit start frame with the optical flow information and frequency domain histogram of the unit start frame.

[0080] For example, the optical flow information of subsequent images of the unit's starting frame can be compared with the optical flow information of the unit's starting frame. If the similarity is greater than the optical flow threshold, it can be determined that the similarity of the optical flow information of the two frames meets the requirements.

[0081] For example, the frequency domain histogram of subsequent images of the unit's starting frame can be compared with the frequency domain histogram of the unit's starting frame. If the similarity is greater than the histogram threshold, it can be determined that the frequency domain histogram similarity of the two frames of images meets the requirements.

[0082] If both optical flow information similarity and frequency domain histogram similarity meet the requirements, it can be determined that the similarity between the corresponding subsequent image and the unit's starting frame is greater than the similarity threshold.

[0083] Optionally, the similarity threshold can be a single value, or the optical flow information similarity and frequency domain histogram similarity can be calculated to determine the total similarity, and the comparison result can be determined by comparing the total similarity with the similarity threshold.

[0084] Optionally, the similarity between optical flow information and frequency domain histogram can be weighted to obtain the total similarity. For example, the average of the optical flow information similarity and frequency domain histogram similarity can be taken as the total similarity.

[0085] Alternatively, optical flow information can be extracted using the open-source OpenCV library, for example, through the following calculation algorithm: flow=cv2.calcOpticalFlowFarneback(gray1,gray2,None,0.5,3,15,3,5,1.2,0); Here, `cv2.calcOpticalFlowFarneback()` represents the OpenCV library function used to calculate optical flow information in an image; `gray1` and `gray2` represent images at two adjacent time points; and the subsequent `None`, `0.5`, `3`, `15`, `3`, `5`, `1.2`, and `0` represent the parameters set for extracting optical flow information. These values ​​are merely examples and may vary depending on actual needs.

[0086] Optionally, the frequency domain histogram of each image can be calculated as follows: hist=cv2.calcHist([hsv_image],[0,1,2],None,bins,[0,180,0,256,0,256]); cv2.normalize(hist,hist); Here, `cv2.calcHist()` represents the OpenCV function used to calculate the frequency domain histogram of an image; `[hsv_image]` represents the image whose frequency domain histogram needs to be calculated; `[0, 1, 2]` represents the channel index list for calculating the frequency domain histogram, indicating that the H, S, and V channels of the HSV image need to be calculated simultaneously, i.e., calculating a three-dimensional histogram; `bins` represents the size (number of bins) of each dimension of the frequency domain histogram; and `[0,180,0,256,0,256]` represents the pixel value range of each channel. Of course, the specific values ​​mentioned above are just examples, and the values ​​may vary depending on the actual needs.

[0087] cv2.normalize() means normalizing the histogram histogram using the L2 norm, so that the L2 norm of the entire histogram is equal to 1, and then writing the normalization result back to histogram in place.

[0088] Alternatively, the similarity of frequency domain histograms can be achieved using the following algorithm: bh_distance=cv2.compareHist(hist1,hist2,cv2.HISTCMP_BHATTACHARYYA); Here, `bh_distance` represents the distance between frequency domain histograms, which can be used to represent the similarity of frequency domain histograms; `hist1` represents the frequency domain histogram of one frame of the image; `hist2` represents the frequency domain histogram of another frame of the image; `cv2.compareHist()` is a function in the OpenCV library used to compare the similarity of two frequency domain histograms; `cv2.HISTCMP_BHATTACHARYYA` indicates that the Bhattacharyya distance between two probability distributions is used to calculate the distance between them.

[0089] Optionally, optical flow similarity and frequency domain histogram similarity can be integrated: flow_similarity=np.dot(flow_features1,flow_features2) / (np.linalg.norm(flow_features1)*np.linalg.norm(flow_features2)); combined_similarity=(color_similarity+flow_similarity) / 2; Here, `flow_similarity` represents the optical flow information similarity, and `color_similarity` represents the frequency domain histogram similarity; `numpy.dot()` is a function used to calculate the dot product of two vectors, namely `flow_features1` and `flow_features2`; `numpy.linalg.norm()` is a function used to calculate the norm of the vectors; the result of `flow_similarity` is a cosine similarity value, used to represent the optical flow information similarity.

[0090] In this embodiment, a similarity threshold can be used to determine whether the work screens are from the same work station.

[0091] The overall similarity can be in the range [0,1], where 1 represents the most similar scene, indicating that they are the same scene, and 0 represents the least similar scene, indicating that they are completely different scenes.

[0092] Optionally, statistical verification and analysis can be performed based on the on-site operation scenario. The similarity threshold can be set to a value greater than 0.5, for example, 0.7. When the calculated total similarity is less than 0.7, it indicates that the operation actions are not for different workstations and can be divided into different sub-videos.

[0093] In this embodiment, OpenCV can be used to save the abruptly changing frames and steady-state frames of the retained calculation results without frame extraction. A compressed video retaining the core operational semantics is generated and determined to be multiple sub-video segments. A composite video is then created by retaining the abruptly changing frames and steady-state frames of the retained calculation results without frame extraction.

[0094] Using the method described above, images with higher similarity are grouped into the same sub-video segment.

[0095] To improve the effectiveness and accuracy of job identification, the sequence of jobs and the completion status of each stage are both identified. The sequence can be determined only after all actions in each step are completed, thus improving the reliability of the identified sequence. Based on this, the job sequence detection method can also include: identifying each sub-video segment to determine the result status of each job workstation.

[0096] Optionally, the result status of each job workstation can be determined based on the recognition of the last frame of each sub-video segment.

[0097] The result status can include whether the operation at the current job station has been completed or not.

[0098] For example, if the operation sequence of the battery management system board interface is being detected, the work station includes multiple board interfaces of the battery management system; the result status of the work station includes plug-in operation status and non-plug-in operation status.

[0099] The above-mentioned identification of each video segment to determine the result status of each work station includes: identifying each video segment and determining whether each work station is in the plug-in operation state by identifying each board interface.

[0100] The result status can include whether the battery management system board's interface is in a plug-in operation state, indicating that the operation of the current work position has been completed; or it can include whether the battery management system board's interface is in a non-plug-in operation state, indicating that the operation of the current work position has not been completed. The plug-in operation state of the battery management system board's interface indicates that the operation has been completed.

[0101] Step 250 above may include: determining the current operation completion order of each job workstation based on the temporal order of each sub-video in the job video to be identified, the job workstation corresponding to each sub-video, and the result status of each job workstation.

[0102] Optionally, the final output may include not only the order in which each job workstation was operated, but also whether the operation at each job workstation was completed.

[0103] In terms of detecting the operation sequence of the battery management system board interface, the above step 250 may include: determining the current operation completion sequence of each board interface based on the time sequence of each sub-video in the video to be identified, the board interface corresponding to each sub-video, and the insertion operation status of each board interface.

[0104] For example, the result status of a job workstation can be sorted as the operation of the current job workstation completed.

[0105] Optionally, all identified job workstations that have been operated can be sorted. The sorting result can include not only the order of each job workstation but also the completion status of each job workstation.

[0106] The sorting in step 250 above can be achieved using a Python program.

[0107] In one example, it is necessary to detect the order of battery management system board insertion operations. The output result can be represented as: S = "[{'right': '1234'},{'left': '567'}]". Here, represent the order as: the first insertion port on the right, the second insertion port on the right, the third insertion port on the right, the fourth insertion port on the right, the fifth insertion port on the left, the sixth insertion port on the left, and the seventh insertion port on the left.

[0108] Through the above implementation logic, it is possible not only to identify the order of each workstation, but also to identify and confirm the completion status of each workstation.

[0109] Based on the identified sequence, to ensure the results reach relevant personnel more quickly, further reminders can be output based on the correctness of the sequence after it has been determined, thus reducing the harm caused by incorrect operation sequences. Based on this, such as... Figure 4 As shown, the method in this application embodiment may further include the following steps.

[0110] Step 310: Compare the current operation completion order with the preset execution order to determine if there are any execution anomalies.

[0111] If an execution exception occurs, proceed to step 320.

[0112] For example, the preset execution order can be different depending on the actual use case and the task.

[0113] Step 320: Output alarm information.

[0114] For example, the alarm information can be an audible and visual alarm, such as a level three audible and visual alarm.

[0115] For example, the alarm message may be a text alarm message. The text alarm message may also include the time when the execution exception occurred.

[0116] Optionally, after an alarm is detected, the original video can be obtained based on the time when the execution anomaly occurred.

[0117] For example, the alarm information mentioned above includes the timestamp of the operation video used in the current operation completion sequence. In the event of an execution anomaly, the method of this embodiment may further include: Step 330: Retrieve the original video corresponding to the order in which the operation was completed in the video database based on the timestamp.

[0118] Step 340: Transmit the original video to the designated terminal.

[0119] After receiving the exception confirmation message from the designated terminal, proceed to step 350.

[0120] For example, the designated terminal may be a terminal used by the relevant person in charge, through which the original video can be displayed. The relevant person in charge may identify the time of the anomaly and the device being operated through manual identification.

[0121] The designated terminal may also include the ability to obtain recognition results input by the relevant responsible person. If the recognition result indicates an execution anomaly compared to the preset execution order, subsequent step 350 can be executed. If the recognition result indicates no execution anomaly compared to the preset execution order, it indicates an error in the current visual model recognition, and the visual model can be updated and optimized using video data generated during the operation. The video data generated during the operation can be the video used for the current recognition or video data generated within a recent specified time period, such as video data from within a week or a month.

[0122] Step 350: Generate a work order for handling execution exceptions.

[0123] For example, after the processing work order is generated, the relevant information of the processing work order can also be sent to the relevant person in charge so that the person in charge can process the processing work order.

[0124] Taking the testing of battery management system board insertion operation as an example, the processing work order can be to replace faulty battery management system board components and restart the battery management system board insertion operation process.

[0125] In this embodiment, a spatiotemporal feature analysis module (YOLOv8+ optical flow motion perception) can be used to replace traditional scribe line marker detection. By fusing multi-scale convolutional features from a large visual model and capturing frames of sudden action changes, the dependence on physical markers is eliminated, reducing the false negative rate caused by scribe line blurring / occlusion.

[0126] To facilitate understanding of the entire process of the job sequence detection method, such as Figure 5 As shown below, the following description uses an example of battery management system board insertion operation sequence detection: S1, the operator initiates the battery management system board insertion operation test.

[0127] S2 uses YOLOv8 to detect the start frame in order to locate the start frame.

[0128] S3, check if the starting frame positioning was successful.

[0129] If the initial frame is successfully located, proceed to step S4. If the initial frame is not located, terminate the process.

[0130] S4, adaptive frame extraction and reconstruction to generate simplified motion video as a sub-video.

[0131] S5 uses a large visual model to load each sub-video for action recognition.

[0132] S6, sort the identified actions.

[0133] S7 outputs the sequence recognition results.

[0134] S8, check if there is an abnormality in the order.

[0135] If the sequence is normal, proceed to step S9; if the sequence is abnormal, proceed to step S10.

[0136] S9 records compliance operation logs.

[0137] S10 triggers an abnormal alarm.

[0138] S11, retrieve the original video for manual review.

[0139] S12, Check if there are any abnormalities in the review order.

[0140] If the verification order is normal, proceed to step S13; if the verification order is abnormal, proceed to step S14.

[0141] S13, marked as a false positive in the large visual model, optimize and update the large visual model.

[0142] S14, output a notification message to notify the relevant responsible person to replace the battery management system board.

[0143] S15, execute the battery management system board replacement action.

[0144] S16, Re-execute the battery management system board insertion operation check.

[0145] Return to step S2 and continue with the sequential detection.

[0146] By employing a motion-aware frame extraction algorithm and analyzing inter-frame optical flow vectors, only action state transition frames (which constitute a small portion of the original video) are retained, generating a high-information-density compressed video summary. This significantly reduces the amount of video processing data, lowers GPU memory usage, and greatly improves the inference speed of large models. Furthermore, a three-level alarm system and a human-machine collaborative review mechanism can be implemented to form a "detection-verification-execution" closed loop.

[0147] Furthermore, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the job sequence detection method described in the above method embodiments.

[0148] The computer program product of the job sequence detection method provided in this application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the steps of the job sequence detection method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0149] In the several embodiments provided in this application, it should be understood that the disclosed methods can also be implemented in other ways. The method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0150] In addition, the method steps in the various embodiments of this application can be integrated together to form an independent part for execution, or each method step can be executed by a separate module, or two or more steps can be formed into an independent part for execution.

[0151] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. It should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0152] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for detecting job order based on a large CV model, characterized in that, include: The video of the task to be identified is divided into multiple sub-videos, wherein the amount of data in each sub-video is no greater than the number of workstations in the task. For each of the aforementioned sub-video segments, the sub-video segments are identified and input into the large visual model; Feature extraction is performed using the feature extraction module of the aforementioned large visual model to obtain video features; The attention mechanism of the large visual model is used to identify the job position of the operation represented in the sub-video; Based on the temporal sequence of each of the sub-video segments in the video to be identified, and the corresponding workstation of the sub-video, the current operation completion order of each workstation is determined.

2. The method according to claim 1, characterized in that, The process of dividing the video to be identified into multiple sub-videos includes: Starting from the first frame of the video of the task to be identified, the frame in which the initial task state is first identified is identified sequentially as the starting frame; wherein, the initial task state includes the task subject being located at the starting work position, and the starting work position is one of all task work positions; The images of each frame after the starting frame in the video to be identified are divided into multiple sub-video segments.

3. The method according to claim 2, characterized in that, The step of dividing the images in the video to be identified, after the starting frame, into multiple sub-video segments includes: Starting from the initial frame, the subsequent images of the initial frame are compared with the initial frame in sequence to determine the multiple consecutive frames that have a similarity greater than the similarity threshold with the initial frame, and these frames are used as the first sub-video. The image adjacent to the i-th sub-video and whose similarity is not greater than the similarity threshold is taken as the unit starting frame of the (i+1)-th sub-video. The subsequent images of the unit starting frame are compared with the unit starting frame in turn to determine the multiple consecutive frames with a similarity greater than the similarity threshold of the unit starting frame, and these are taken as the (i+1)-th sub-video, where i is a positive integer and not greater than N-1, and N is the number of workstations in the job.

4. The method according to claim 3, characterized in that, The step of sequentially comparing the similarity of subsequent images with the starting frame includes: Extract subsequent images from the starting frame and the optical flow information of the starting frame; Calculate the subsequent images of the starting frame and the frequency domain histogram of the starting frame; The optical flow information and frequency domain histogram of subsequent images of the starting frame are compared with the optical flow information and frequency domain histogram of the starting frame.

5. The method according to claim 1, characterized in that, The workstation includes multiple board interfaces for the battery management system; the method further includes: For each video segment, identification is performed to determine whether each of the aforementioned board interfaces is in a plug-in operation state; The step of determining the current operation completion order of each workstation based on the temporal order of each of the sub-video segments in the video to be identified and the workstation corresponding to each sub-video includes: Based on the temporal sequence of each sub-video segment in the video to be identified, the corresponding board interface for each sub-video, and the insertion status of each board interface, the current operation completion order of each board interface is determined.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: The order in which the current operation is completed is compared with the preset execution order to determine whether there is any execution abnormality; If an execution error occurs, an alarm message will be output.

7. The method according to claim 6, characterized in that, The alarm information includes the timestamp of the operation video used in the current operation completion sequence; The method further includes: In the event of an execution exception, the original video corresponding to the order in which the current operation was completed is retrieved from the video database based on the timestamp. Transmit the original video to the designated terminal; Upon receiving the exception confirmation message from the designated terminal, a processing work order is generated for the execution exception.

8. An electronic device, characterized in that, include: The processor and memory, wherein the memory stores machine-readable instructions executable by the processor, wherein when the electronic device is running, the machine-readable instructions are executed by the processor to perform the steps of the method as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Construction method and application of operation process normativity identification model

    CN111507277A

  • Operation monitoring method and device, electronic equipment and storage medium

    CN113395480A

  • Abnormal behavior recognition method and device

    CN118115935A

  • Production operation optimization strategy generation method and device and terminal equipment

    CN118963270A