A disassembly behavior detection method, an electronic device, and a storage medium
By identifying the key points of the human body and the disassembly tool areas in the video files and judging the changes in the disassembly personnel's posture, the problem of rapid and accurate review of disassembly behavior is solved, which reduces labor costs and improves detection efficiency.
Patent Information
- Application Number
- CN202310344376.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-27
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2043-03-27
AI Technical Summary
During the dismantling of used electrical appliances, existing technologies make it difficult to quickly and accurately review irregular dismantling behaviors, which results in the destruction of valuable electronic components. In addition, manual review is highly dependent, costly, and inefficient.
By identifying the key points of human hands, human postures, and disassembly tool areas in the images in the video files, it is determined whether there is a situation in which the human posture changes from the first disassembly posture to the second disassembly posture in multiple consecutive images, and a prompt message is issued to indicate improper disassembly behavior.
It enables rapid and accurate review of irregular disassembly behaviors, reduces labor costs, and improves detection accuracy and efficiency.
Smart Images

Figure CN116311529B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of waste electrical appliance treatment, and in particular to a disassembly behavior detection method, electronic equipment, and storage medium. Background Art
[0002] With economic development and technological progress, the electronics industry and information high-tech industry have developed rapidly, and the number of waste electrical appliances has increased. However, some electronic components in waste electrical appliances can still be recycled and put into use again.
[0003] During the dismantling process of used electrical appliances, if the dismantler uses dismantling tools to hit the dismantled objects in an irregular manner, some valuable electronic components will be destroyed and lose their recycling value.
[0004] Therefore, how to quickly and accurately review irregular dismantling behavior is a problem that needs to be solved. Summary of the Invention
[0005] The present application provides a disassembly behavior detection method, electronic device, and storage medium for quickly and accurately reviewing irregular disassembly behavior.
[0006] To achieve the above technical objectives, this application adopts the following technical solutions:
[0007] In a first aspect, an embodiment of the present application provides a disassembly behavior detection method, the method comprising:
[0008] Get images from video files;
[0009] Recognize the image and obtain the recognition results of the key points of the human hand, human posture and disassembly tool area in the image;
[0010] If multiple consecutive images in the video file all meet a first preset condition, determining whether a human body posture changes from a first disassembly posture to a second disassembly posture in the multiple images, the first preset condition being that a key point of the human hand overlaps with a disassembly tool area;
[0011] When a human body posture changes from a first disassembly posture to a second disassembly posture in the plurality of images, a first prompt message is issued, where the first prompt message is used to indicate that an irregular disassembly behavior exists.
[0012] The technical solution provided by the present application brings at least the following beneficial effects: by identifying the key points of the human hand, the human posture and the disassembly tool area in the image, when the key points of the human hand in the image coincide with the disassembly tool area, it can be considered that the disassembly personnel is holding the disassembly tool. When there are multiple continuous images that meet the condition that the disassembly personnel is holding the disassembly tool, and there is a situation where the human posture of the disassembly personnel changes from the first disassembly posture to the second disassembly posture, it indicates that the disassembly personnel has the behavior of knocking the disassembly object with the hand-held disassembly tool, that is, the disassembly personnel has an irregular disassembly behavior that does not comply with the relevant norms for disassembly behavior. The present application simplifies the complex disassembly behavior of knocking and hammering the disassembly object with hand-held disassembly tools into an identification process of the key points of the human hand, the human posture and the disassembly tool area, which reduces the difficulty of identification, can improve the accuracy of disassembly behavior detection, and saves labor costs.
[0013] In one possible implementation, image recognition is performed to obtain recognition results for key points of human hands, human posture, and the area of disassembly tools within the image. This includes: receiving a user's annotation of an area within an image from a video file and identifying the area marked by the annotation as the work area; and identifying the work area within the image to obtain recognition results for key points of human hands, human posture, and the area of disassembly tools within the work area. This allows the user to mark the work area themselves, while the electronic device only needs to identify the work area. This reduces workload and prevents irrelevant people or events in the image from affecting the accuracy of disassembly behavior detection results. This makes the detection process faster and meets the user's detection requirements.
[0014] In one possible implementation, the first disassembly posture is a hand-raising posture, and the second disassembly posture is a non-hand-raising posture; or the first disassembly posture is a non-hand-raising posture, and the second disassembly posture is a hand-raising posture. In this way, by setting the first disassembly posture and the second disassembly posture to a hand-raising posture or a non-hand-raising posture, the electronic device can more easily determine the change in the disassembly person's body posture, facilitating the detection of irregular disassembly behavior.
[0015] In one possible implementation, the hand-raising posture is defined as follows: on the side of the human hand key point that overlaps with the disassembly tool area, the angle formed by the hand key point, elbow key point, and shoulder key point is within a preset angle range, and the position of the hand key point is higher than that of the elbow key point. In this way, the angles formed by the hand, elbow, and shoulder can be accurately determined using the various key points of the human body. The relationship between the angles and the positions of the hand and elbow can be used to more accurately determine whether the disassembly operator is in the hand-raising posture, thereby determining the transition of the disassembly operator's body posture.
[0016] In one possible implementation, the method further includes recording a set of continuous image sequences as a single instance of irregular disassembly behavior if a continuous image sequence that satisfies a second preset condition exists among the multiple images, wherein the second preset condition is that the continuous image sequence includes, in sequence, a first preset number of consecutive images and a second preset number of consecutive images, the human body posture in the first preset number of images is a hand-raising posture, and the human body posture in the second preset number of images is a non-hand-raising posture, and the images in any two sets of continuous image sequences are different. In this way, by detecting continuous image sequences that satisfy the second preset condition, the number of instances of irregular disassembly behavior can be recorded more quickly, facilitating staff review of videos of disassembling used electrical appliances.
[0017] In one possible implementation, the method further includes counting the number of times non-standard disassembly behaviors occur in the video file and issuing a second prompt message, wherein the second prompt message is used to indicate the number of times non-standard disassembly behaviors occur in the video file. Thus, issuing the second prompt message can alert the user to the number of times non-standard disassembly behaviors occur in the video file, facilitating the user's review of the non-standard disassembly behaviors in the video file.
[0018] In a second aspect, the present application provides an electronic device, including: an acquisition module for acquiring images in a video file; a processing module for identifying the images to obtain identification results of key points of human hands, human postures, and disassembly tool areas in the images; the processing module is also used to, if multiple consecutive images in the video file all meet a first preset condition, determine whether there is a situation in which the human body posture changes from a first disassembly posture to a second disassembly posture in the multiple images, the first preset condition being that the key points of the human hand and the disassembly tool area have overlapping parts; a sending module for issuing a first prompt message if there is a situation in which the human body posture changes from a first disassembly posture to a second disassembly posture in the multiple images, the first prompt message being used to indicate that there is irregular disassembly behavior in the video file.
[0019] In one possible implementation, the electronic device also includes a receiving module, which is used to receive a user's annotation operation on an area in an image of a video file; the processing module is specifically used to use the area marked by the annotation operation as a working area; identify the working area in the image, and obtain recognition results of key points of human hands, human postures, and disassembly tool areas in the working area.
[0020] In a possible implementation, the first disassembly posture is a hand-raising posture, and the second disassembly posture is a non-hand-raising posture; or, the first disassembly posture is a non-hand-raising posture, and the second disassembly posture is a hand-raising posture.
[0021] In one possible implementation, the hand-raising posture is: on the side of the key point of the human hand that overlaps with the disassembly tool area, the angle formed by the hand key point, the elbow key point and the shoulder key point is within a preset angle range, and the position of the hand key point is higher than the position of the elbow key point.
[0022] In one possible implementation, the processing module is also used to record a group of continuous image sequences as an irregular disassembly behavior if there is a continuous image sequence that meets a second preset condition among multiple images. The second preset condition is that the continuous image sequence includes a continuous first preset number of images and a continuous second preset number of images in sequence, the human body posture in the first preset number of images is a hand-raising posture, and the human body posture in the second preset number of images is a non-hand-raising posture, and the images in any two groups of continuous image sequences are different.
[0023] In a possible implementation, the processing module is further used to count the number of times that irregular disassembly behaviors occur in the video file; the sending module is further used to send a second prompt message, and the second prompt message is used to indicate the number of times that irregular disassembly behaviors occur in the video file.
[0024] In a third aspect, the present application provides an electronic device comprising: one or more processors; one or more memories; wherein the one or more memories are used to store computer program code, the computer program code including computer instructions, and when the one or more processors execute the computer instructions, the electronic device executes any one of the disassembly behavior detection methods provided in the first aspect above.
[0025] In a fourth aspect, the present application provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are run on a computer, the computer executes any one of the disassembly behavior detection methods provided in the first aspect above.
[0026] In a fifth aspect, the present application provides a computer program product, which includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the disassembly behavior detection method provided in the first aspect and any possible design thereof.
[0027] For the specific descriptions of the second to fifth aspects and their various implementations in this application, reference can be made to the detailed descriptions in the first aspect and its various implementations; and for the beneficial effects of the second to fifth aspects and their various implementations, reference can be made to the analysis of the beneficial effects in the first aspect and its various implementations, which will not be repeated here.
[0028] These and other aspects of the present application will become more readily apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 A schematic diagram of the structure of a disassembly behavior detection system provided in an embodiment of the present application;
[0030] Figure 2 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;
[0031] Figure 3 A flowchart of a disassembly behavior detection method provided in an embodiment of the present application;
[0032] Figure 4 A schematic diagram of human body key point recognition provided in an embodiment of the present application;
[0033] Figure 5 A schematic diagram of the network structure of a human key point detection model provided in an embodiment of the present application;
[0034] Figure 6 A schematic diagram of the network structure of a disassembly tool detection model provided in an embodiment of the present application;
[0035] Figure 7 A schematic diagram of a model training process provided in an embodiment of the present application;
[0036] Figure 8 Schematic diagram of an application scenario of a disassembly behavior detection method provided in an embodiment of the present application Figure 1 ;
[0037] Figure 9 Schematic diagram of an application scenario of a disassembly behavior detection method provided in an embodiment of the present application Figure 2 ;
[0038] Figure 10 Schematic diagram of an application scenario of a disassembly behavior detection method provided in an embodiment of the present application Figure 3 ;
[0039] Figure 11 Schematic diagram of an application scenario of a disassembly behavior detection method provided in an embodiment of the present application Figure 4 ;
[0040] Figure 12 Schematic diagram of an application scenario of a disassembly behavior detection method provided in an embodiment of the present application Figure 5 ;
[0041] Figure 13 Schematic diagram of an application scenario of a disassembly behavior detection method provided in an embodiment of the present application Figure 6 ;
[0042] Figure 14 A logical diagram of a disassembly behavior detection method provided in an embodiment of the present application;
[0043] Figure 15 A schematic structural diagram of another electronic device provided in an embodiment of the present application;
[0044] Figure 16 A schematic structural diagram of another electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0046] It should be noted that, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design schemes. To be precise, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete way. The terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, "multiple" means two or more.
[0047] As mentioned in the background, during the dismantling of used electrical appliances, if dismantlers engage in irregular behavior, such as swinging dismantling tools and striking the objects, this can damage valuable electronic components and render them worthless. Currently, the primary method for determining whether irregular dismantling occurs is through manual review of dismantling videos. This relies heavily on manual experience, consumes significant labor costs, and has low audit efficiency. Furthermore, manual effort is required to obtain information about the time, location, and personnel involved in the irregular dismantling, resulting in limited accuracy and stability in the audit results. Therefore, how to quickly and accurately audit irregular dismantling behavior is a challenge that needs to be addressed.
[0048] In this regard, an embodiment of the present application provides a disassembly behavior detection method, which identifies the key points of human hands, human postures, and disassembly tool areas in images in a video file. If multiple consecutive images meet a first preset condition, and there is a situation in which the human posture changes from a first disassembly posture to a second disassembly posture, indicating that the disassembly personnel may have engaged in irregular disassembly behavior, the electronic device will issue a first prompt message, thereby reminding the user that there is irregular disassembly behavior in the video file.
[0049] Therefore, the disassembly behavior detection method provided by this application can not only quickly and accurately review irregular disassembly behaviors, but also save labor costs.
[0050] Please refer to Figure 1 , which shows a schematic diagram of the structure of the disassembly behavior detection system to which the disassembly behavior detection method provided in this application is applicable. Figure 1 As shown, the disassembly behavior detection system 1 may include: a video acquisition device 10 and an electronic device 20.
[0051] Among them, a communication connection is established between the video capture device 10 and the electronic device 20 and the prompt device 30. It should be understood that the connection method can be a wireless connection, such as a Bluetooth connection, a wireless fidelity (Wi-Fi) connection, etc.; or the connection method can also be a wired connection, such as an optical fiber connection, etc., without limitation. Exemplarily, the video capture device 10, the electronic device 20 or the prompt device 30 can be connected to the Internet through a router, thereby realizing a communication connection between the electronic device 20 and the video capture device 10 and the prompt device 30.
[0052] In some embodiments, the video capture device 10 is configured to output a video file to the electronic device 20. For example, when it is necessary to detect whether there is irregular disassembly behavior in an image of a video file, the video capture device 10 can send the video file to the electronic device 20, and the electronic device 20 then detects the image in the video file to determine whether there is irregular disassembly behavior.
[0053] In some embodiments, the electronic device 20 is used to identify images in a video file to obtain recognition results of key points of the human hand, human posture and disassembly tool area in the image, and then when there are overlapping parts of the key points of the human hand and the disassembly tool area in multiple consecutive images, and the human posture in multiple images changes from a first disassembly posture to a second disassembly posture, a first prompt information is issued to indicate the existence of irregular disassembly behavior.
[0054] In some embodiments, the electronic device 20 may include a processor. The processor is used to identify key points of the human body and disassembly tools in the image of the video file, and determine whether there is any irregular disassembly behavior in the video file based on the identification results of the key points of the human body and the disassembly tools. The processor can be a central processing unit (CPU), a general-purpose processor network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor can also be other devices with processing functions, such as circuits, devices, or software modules, and the embodiments of the present application do not impose any restrictions on this.
[0055] Optionally, the electronic device may further include a memory for storing images in the video file from the video acquisition device 10, and then the processor may search the image in the video file from the memory to identify key points of the human body and disassembly tools in the image in the video file.
[0056] In some embodiments, the disassembly behavior detection system 1 may further include a prompting device 30. The prompting device 30 is configured to display the first prompt information. For example, the prompting device 30 may be a voice prompting device, in which case the prompting device 30 presents the first prompt information to the user by reading the first prompt information aloud. Alternatively, the prompting device 30 may be a display device, in which case the prompting device 30 presents the first prompt information to the user by displaying the first prompt information on a display screen.
[0057] In some embodiments, the video capture device 10 and the electronic device 20 can be as follows Figure 1 As shown, they are two independent devices, or the video acquisition device 10 and the electronic device 20 can be integrated into the same device.
[0058] In some embodiments, the electronic device 20 and the prompting device 30 may be as follows: Figure 1 As shown, they are two independent devices, or the electronic device 20 and the prompting device 30 can be integrated into the same device.
[0059] In some embodiments, the disassembly behavior detection system 1 may include one or more video acquisition devices 10 .
[0060] In some embodiments, the video acquisition device 10 can be any device that can transmit video files to the electronic device 20, such as a camera, a terminal device with video transmission function (such as a mobile phone, a tablet computer, a laptop computer, etc.), a digital video disc (Digital Video Disc, DVD), a set-top box, a satellite receiver, etc. The embodiment of the present application does not limit the specific form of the video acquisition device 10.
[0061] In some embodiments, the electronic device 20 may be a single server or a server cluster, or the electronic device 20 may be a terminal device, such as a personal computer (PC), a notebook computer, a mobile device, a tablet computer, a laptop computer, etc. The embodiments of the present application do not limit the specific form of the electronic device 20.
[0062] The hardware structure of the electronic device 20 includes Figure 2 The computing device shown in FIG. Figure 2 Taking the computing device shown as an example, the hardware structure of the electronic device 20 is introduced.
[0063] like Figure 2 As shown, the computing device may include a processor 201 , a memory 202 , a communication interface 203 , and a bus 204 . The processor 201 , the memory 202 , and the communication interface 203 may be connected via a bus 204 .
[0064] Processor 201 is the control center of the computing device and can be a single processor or a collective term for multiple processing elements. For example, processor 201 can be a general-purpose central processing unit (CPU) or other general-purpose processor. A general-purpose processor can be a microprocessor or any conventional processor.
[0065] As an embodiment, the processor 201 may include one or more CPUs, such as Figure 2 CPU 0 and CPU 1 are shown in Figure 1.
[0066] The memory 202 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0067] In one possible implementation, memory 202 may exist independently of processor 201 and may be connected to processor 201 via bus 204 for storing instructions or program code. When processor 201 calls and executes the instructions or program code stored in memory 202, the model deployment method provided in the embodiments of the present application can be implemented.
[0068] In another possible implementation, the memory 202 may also be integrated with the processor 201 .
[0069] The communication interface 203 is used to connect the computing device to other devices via a communication network, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 203 may include a receiving unit for receiving data and a sending unit for sending data.
[0070] The bus 204 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 2 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0071] It should be pointed out that Figure 2 The structure shown in the figure does not constitute a limitation on the computing device, except Figure 2In addition to the components shown, the computing device may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0072] The implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0073] The disassembly behavior detection method provided in the embodiment of the present application can be executed by the electronic device 20 in the disassembly behavior detection system 1.
[0074] like Figure 3 As shown, the embodiment of the present application provides a disassembly behavior detection method, which includes the following steps:
[0075] S101: Acquire an image in a video file.
[0076] For example, the video file may be a video file of the disassembly process captured by a camera, or may be a video file of the disassembly process transmitted to the electronic device by a terminal device such as a mobile phone or a computer.
[0077] S102: Recognize the image to obtain recognition results of key points of the human hand, human posture, and disassembly tool area in the image.
[0078] Among them, the disassembly tool can be a tool such as a hammer that can be used to knock or hammer old electrical appliances, and the embodiments of this application do not specifically limit this.
[0079] In some embodiments, step S102 can be specifically implemented as follows: identifying human body key points in the image in the video file, determining the human hand key points from the identified human body key points, and then determining the human body posture based on the identified human body key points, and then identifying the disassembly tools in the image to determine the disassembly tool area in the image.
[0080] In this way, the complex disassembly behavior of knocking and hammering objects is simplified to the identification process of key points on the human body and the areas of disassembly tools. Compared with the identification process of complex and irregular disassembly behaviors, only key points on the human body and disassembly tools are identified, which reduces the difficulty of identification and makes the identification process simpler, thereby improving the speed of detection and the accuracy of the detection results.
[0081] In some embodiments, the electronic device recognizes key points of the human body in an image, which can be specifically implemented as follows: inputting the image in the video file into the first model to obtain the recognition result of the key points of the human body in the image, and the first model is used to recognize the key points of the human body in the image.
[0082] Exemplarily, the first model may be a YOLO-Pose human key point recognition model.
[0083] The following is an example of the training process of the YOLO-Pose human key point recognition model:
[0084] Before detecting human key points, you can use the public COCO keypoint track human key point dataset to train the YOLO-Pose human key point recognition model. The COCO keypoint track human key point dataset includes a large number of labeled sample images, such as Figure 4 As shown in Table 1, 17 key points are annotated on the human body in each sample image. For example, the left shoulder of the human body corresponds to key point 5, the left elbow of the human body corresponds to key point 7, the left hand of the human body corresponds to key point 9, the right shoulder of the human body corresponds to key point 6, the right elbow of the human body corresponds to key point 8, and the right hand of the human body corresponds to key point 10.
[0085] Table 1 Human body key point annotation table
[0086]
[0087]
[0088] Among them, for each human body, 57 elements are predicted, including: the position x, y and confidence conf of each key point in the 17 key points of the human body, a total of 51 elements, and the coordinates C of the center point of the prediction box (the smallest rectangle covering the human body) where the human body is located x and C y , width W, height H and confidence box conf There are 6 elements in total, so all the elements that need to be predicted can be expressed as a prediction vector, that is:
[0089]
[0090] in, Indicates the horizontal coordinate of the nth key point in training, Indicates the vertical coordinate of the nth key point in training, The visible flag representing the nth keypoint during training. If Indicates that this key point is not marked. In this case When , it means that the key point is marked but not visible (occluded). , which means that this key point is marked and visible at the same time. In the reasoning stage, the confidence is retained. Keypoints with a value greater than 0 are retained as labeled keypoints, while unlabeled keypoints are discarded. Since 17 keypoints are output for each human keypoint detection during model training, keypoints outside the field of view need to be filtered out. Otherwise, there will be dangling keypoints, causing deformation of the human skeleton. The COCO keypoint track human keypoint dataset is divided into training and test datasets in a 9:1 ratio. Finally, before model training, the COCO keypoint track human keypoint dataset is enhanced based on resolution and lighting factors. Histogram equalization is used to enhance image contrast, and the image resolution is appropriately adjusted based on the resolution requirements of the actual detection.
[0091] In practical applications, you can use Figure 5 The network structure of the YOLO-Pose human keypoint detection model shown in Figure 2 uses the CSP-darknet53 backbone feature extraction network as its backbone. This network extracts rich, informative features from the input image. This network addresses the gradient duplication issue in network optimization in other large convolutional neural network frameworks, integrating gradient changes throughout the feature map. This reduces the number of parameters and floating-point operations per second (FLOPS) of the YOLO-Pose human keypoint detection model, ensuring both inference speed and accuracy while also reducing model size. The Neck network of the YOLO-Pose human keypoint detection model uses the PANet architecture. The Neck network is primarily used to generate feature pyramids, which enhance the model's ability to detect objects at different scales, enabling it to recognize the same object at different sizes and scales. The PANet architecture enhances the FPN algorithm with a bottom-up approach, enabling the top-level feature map to leverage the rich location information from the bottom layers, thereby improving the detection of large objects. The PANet architecture fuses features of various scales from the backbone, followed by the head architecture for final detection. The architecture outputs four detection heads of varying sizes, along with predicted bounding boxes and keypoints for each head.
[0092] The YOLO-Pose human key point detection model uses complete-intersection over union (CIOU) as the regression positioning loss function of the prediction box (bounding box, bbox), where the intersection over union (IOU) means that for any prediction box A and B, the intersection and union are calculated respectively, and the ratio of the two is calculated. The expression of IOU is The regression positioning loss of the YOLO-Pose human keypoint detection model considers three geometric parameters: overlap area, center point distance, and aspect ratio. CIOU adds a penalty term (loss) for the predicted box scale based on the distance-intersection over union (DIOU), so that the predicted box is more consistent with the ground truth box. The expression of CIOU is as follows:
[0093]
[0094] Where α is the weight function, And υ is used to measure the similarity of aspect ratio, defined as Then the complete prediction box regression positioning loss function is defined as:
[0095] Target keypoint similarity (OKS) is a commonly used metric for evaluating keypoints. By directly defining the regressed keypoints as fixed centers, the concept of IOU loss can be extended from bbox to keypoints. In the presence of keypoints, OKS is treated as IOU, so the OKS loss is essentially scale-invariant and tilted towards specific keypoints. In other words, keypoints on a person's head, such as ears, nose, and eyes, receive more pixel-level error penalties than keypoints on the body, such as shoulders, knees, and hips. Unlike the standard IOU loss, the gradient of the IOU loss disappears when there is no overlap, while the OKS loss does not experience gradient vanishing. Therefore, the OKS loss is more similar to the DIOU loss.
[0096] Corresponding to each bbox, the entire pose information is stored. Therefore, OKS is calculated for each key point separately, and then added together to obtain the final OKS loss or key point IOU loss:
[0097]
[0098] Among them, d n is the Euclidean distance between the predicted nth key point and its true value, k nis the penalty weight factor of the key point itself, s is the scale of the target, δ(υ n ) is the visible mark of each key point, which is also mentioned in the above content
[0099] Corresponding to each key point, a confidence parameter is learned, which shows whether there is a key point on the human body. The loss function is:
[0100]
[0101] in, is the confidence of the predicted n-th key point.
[0102] Finally, the overall loss of the human key point prediction model is expressed as:
[0103]
[0104] Among them, λ cls =0.5,λ box =0.05,λ kpts =0.1,λ kpts_conf = 0.5 is a hyperparameter in the model, used to balance the losses between different scales.
[0105] In summary, the training process of the YOLO-Pose human keypoint detection model can be summarized as follows: Initialize the weights and parameters of the improved YOLO-Pose human keypoint detection model, including the convolutional layer parameter values, learning rate, number of iterations (epochs), and the number of data samples captured in one training session (batch_size). The COCOkeypoints human keypoint dataset is divided into training and test sets and placed in a predetermined directory. Then, run the program to perform training. Before training, the program selects 1 / 10 of the data from the training set as the validation set. Each iteration is validated against the validation set, and any difficult samples with poor performance are recorded. After the number of iterations is reached, the various metrics of the YOLO-Pose human keypoint detection model are output, including mean average precision (mAP), single-class average precision (AP), precision, and recall. When the number of iterations is reached, training ends, and the parameters and weights of the YOLO-Pose human keypoint detection model are saved.
[0106] In other embodiments, the electronic device identifies the disassembly tools in the image and determines the disassembly tool area in the image. This can be specifically implemented by inputting the image in the video file into the second model to obtain the identification result of the disassembly tool area in the image. The second model is used to identify the disassembly tools in the image.
[0107] Exemplarily, the second model may be a YOLO-v5 model.
[0108] The following exemplary introduces the process of training YOLO-v5 model to recognize disassembling tools:
[0109] First, a data set of scenes of non-standard disassembly behavior needs to be collected. The data set of scenes of non-standard disassembly behavior can be derived from video recordings provided by waste electrical appliance disassembly sites. Frame extraction is performed on the video recordings, and scene picture materials of non-standard disassembly behavior using disassembly tools are obtained. The disassembly tools in the scene picture materials are labeled, and the label format of the scene picture materials is converted into the TXT data set format of YOLO-V5 before training. The scene picture materials are divided into training data set and test data set in the ratio of 9:1. Finally, random scale ([0.5, 1.5]) data enhancement, random translation (-10, 10), random flip with a probability of 0.5, mosaic (Mosaic) enhancement with a probability of 1, image disturbance, changing brightness, contrast, saturation, hue, adding noise, random scaling, random cropping, flipping, rotating, random erasing, etc. are used.
[0110] Before detecting disassembly tools through the YOLO-v5 model, the network structure of the YOLO-v5 model needs to be constructed. The network structure of the YOLO-v5 model can be similar to the human key point network structure as shown in Figure 6 , and will not be described again. The loss function uses BCE-Logits to calculate the loss of the target score, uses cross-entropy loss (BCE-cals-loss) to calculate the loss of the category score, and uses generalized intersection over union (GIOU) to calculate the loss of the predicted frame. The GIOU formula is as follows: The meaning of this formula is that for two predicted frames A and B, we can find a minimum enclosing rectangle C that can contain A and B, then calculate the ratio of the area of C that does not cover A and B to the total area of C, and then subtract the ratio from the IOU of A and B. Similar to IOU, the loss function of GIOU can be represented as .
[0111] In summary, the YOLOv5 model training process can be summarized as follows: Initialize the weights and parameters of the improved YOLO-v5 model, including the convolutional layer parameter values, learning rate, number of iterations (epochs), and the number of data samples captured in one training run (batch_size). Place the disassembly scenario dataset in a predetermined directory and run the training program. Before training, the program selects 1 / 10 of the data from the training set as a validation set. Each iteration is validated against the validation set, and any difficult samples with poor performance are recorded. After reaching the required number of iterations, the program outputs various model metrics, including mean average precision (mAP), single-class average precision (AP), precision, and recall. When the required number of iterations is reached, training ends, and the model parameters and weights are saved. Average class accuracy, model parameters, and FLOPS are used as performance metrics to test the YOLOv5 model's ability to detect disassembly tools.
[0112] like Figure 7 As shown in the figure, through the above model training process, the YOLO-Pose human key point detection model and the YOLO-v5 model are finally combined into a comprehensive model. When the electronic device recognizes the image, it only needs to input the image in the video file into the comprehensive model to obtain the recognition results of the human key points and disassembly tool areas.
[0113] In this way, the complex disassembly behavior of knocking and hammering objects with handheld disassembly tools is simplified to the process of identifying key points on the human body and the areas of the disassembly tool. This is conducive to the training and tuning of the model, and makes it easier for the machine to learn the recognition process of the disassembly behavior. The recognition process is faster and the recognition results are more accurate.
[0114] In some embodiments, the electronic device receives a user's annotation operation on an area in an image of a video file, and uses the area marked by the annotation operation as a working area, and then identifies the working area in the image to obtain recognition results of key points of human hands, human postures, and disassembly tool areas in the working area.
[0115] For example, Figure 8 As shown in the figure, the user identifies the dismantling worker's workstation in an image based on the video file's shooting angle and marks an area in the image. The electronic device then identifies the marked area as the work area. The shaded area in the image represents the work area, and the electronic device only recognizes the key points of the human hand, the human posture, and the dismantling tool area in the shaded area, ignoring other areas.
[0116] In this way, the user marks the work area by himself, and the electronic device only needs to identify the work area, which reduces some of the workload and prevents irrelevant people or events in the image from affecting the accuracy of the disassembly behavior detection results. The detection process is faster and meets the user's detection requirements.
[0117] Optionally, before the electronic device begins the disassembly behavior detection process, the user can set the configuration information of the electronic device's detection task. The configuration information of the electronic device's detection task can include one or more of the following:
[0118] 1. Data channel, used to indicate the source of the video file.
[0119] For example, the acquisition source of the first video file may be a camera that shoots the first video file.
[0120] 2. Detection rules.
[0121] For example, the detection rule set by the user is to detect images of the working area in the video file of the dismantling personnel during working hours at full frame rate. If the dismantling personnel's working hours in a day are from 08:30 to 18:00, the electronic device detects images of all video files between 08:30 and 18:00.
[0122] 3. Alarm rules.
[0123] Among them, the alarm rules may include alarm interval, alarm event type, alarm event name, etc.
[0124] For example, the user sets an alarm rule to report detection results every 10 seconds, with the alarm event name being "Irregular disassembly behavior detected" and the alarm event type being irregular disassembly behavior. If, when the electronic device performs a target detection task, the next time it reports an alarm event is 08:30:30, irregular disassembly behavior is detected at 08:30:23, and irregular disassembly behavior is detected at 08:30:26, then at 08:30:30, the electronic device reports two first prompt messages: "Irregular disassembly behavior detected at 08:30:23" and "Irregular disassembly behavior detected at 08:30:26."
[0125] S103: If the plurality of consecutive images in the video file all meet the first preset condition, determine whether there is a situation in which the posture of the human body changes from the first disassembly posture to the second disassembly posture in the plurality of images.
[0126] Among them, the first preset condition is that the key points of the human hand and the disassembly tool area have overlapping parts.
[0127] In some embodiments, the first disassembly gesture is a hand-raising gesture, and the second disassembly gesture is a non-hand-raising gesture; or, the first disassembly gesture is a non-hand-raising gesture, and the second disassembly gesture is a hand-raising gesture.
[0128] In some embodiments, the hand-raising posture is: on the side of the key point of the human hand that overlaps with the disassembly tool area, the angle formed by the hand key point, the elbow key point and the shoulder key point is within the preset angle range, and the position of the hand key point is higher than the position of the elbow key point.
[0129] Optionally, the preset angle range is 0° to 145°. This preset angle range can be determined by the user by collecting relevant data on various hand-raising postures, and the angle range between the hand, elbow and shoulder that best fits the human hand-raising posture.
[0130] For example, Figure 9 As shown in the figure, the dismantler holds the dismantling tool in his right hand, and the angle formed by his right hand, elbow and shoulder is 90°, so the dismantler's right hand is in a raised hand posture. Figure 10 As shown, the dismantler holds the dismantling tool in his right hand, and the angle formed between his right hand, elbow and shoulder is 163°, so the dismantler's right hand is in a non-raised hand posture.
[0131] In this way, the angle formed between the hand, elbow and shoulder can be accurately determined through various key points of the human body. And through the relationship between the angle and the position of the hand and elbow, it is possible to accurately judge whether the dismantling personnel is in a hand-raising posture, thereby determining the change of the dismantling personnel's body posture.
[0132] S104: When a human body posture changes from a first disassembly posture to a second disassembly posture in the multiple images, issue a first prompt message.
[0133] The first prompt information is used to indicate that there is irregular disassembly behavior in the video file.
[0134] In some embodiments, the first prompt information may further include: the time when the non-standard disassembly behavior occurs, relevant information of the disassembly personnel who performed the non-standard disassembly behavior, and / or the source of the video file.
[0135] The relevant information of the dismantling personnel may include identification information of the dismantling personnel (such as work number, name, etc.), working hours of the dismantling personnel, the number of times the dismantling personnel have engaged in irregular dismantling behaviors, etc.
[0136] For example, Figure 11As shown, a plurality of first prompt information is displayed on the display screen of the electronic device, and the user can search for a specific disassembling personnel, a specific time, etc. in the interface to quickly find the corresponding first prompt information and determine the authenticity of the first prompt information.
[0137] In this way, the user can directly view the related information of the non-standard disassembling behavior from the first prompt information without the user searching by himself / herself, thereby saving the user's time and making the auditing process more convenient and fast.
[0138] Figure 3 The technical solution has at least the following beneficial effects: by identifying the human hand key points, the human posture, and the disassembling tool region in the image, when the human hand key points in the image coincide with the disassembling tool region, it can be considered that the disassembling personnel holds the disassembling tool. When there is a case that the human posture of the disassembling personnel changes from a first disassembling posture to a second disassembling posture in a plurality of images that satisfy the condition of the disassembling personnel holding the disassembling tool and are continuous, it indicates that the disassembling personnel has the behavior of holding the disassembling tool to knock the disassembling object, i.e., the disassembling personnel has the non-standard disassembling behavior, which does not conform to the relevant standard of the disassembling behavior. The complex disassembling behavior of holding the disassembling tool to knock and hammer the disassembling object is simplified to the identification process of the human hand key points, the human posture, and the disassembling tool region, which reduces the identification difficulty, improves the accuracy of the disassembling behavior detection, and saves the labor cost.
[0139] In some embodiments, if there is a continuous image sequence that satisfies a second preset condition in a plurality of continuous images that satisfy a first preset condition, a group of continuous image sequences is recorded as a non-standard disassembling behavior. The second preset condition is that the continuous image sequence includes a first preset number of continuous images and a second preset number of continuous images in sequence, the human posture in the first preset number of images is a hand-raising posture, the human posture in the second preset number of images is a non-hand-raising posture, and the images in any two groups of continuous image sequences are different.
[0140] Optionally, the human posture in the first preset number of images is a non-hand-raising posture, and the human posture in the second preset number of images is a hand-raising posture.
[0141] For example, the detection personnel set the first preset number to 4 and the second preset number to 2, and there is no interval between the first preset number and the second preset number, i.e., when 6 continuous images of holding the disassembling tool are detected, the human posture in the first 4 images is a hand-raising posture, and the human posture in the last 2 images is a non-hand-raising posture, then the 6 images are a group of continuous image sequences, and the group of continuous image sequences correspond to the non-standard disassembling behavior in sequence. Figure 12As shown, in the images at time T0, T1, T2, and T3, the dismantling worker's body posture is that his right hand is in a raised hand posture, and his right hand is holding a dismantling tool. In the images at time T4 and T5, the dismantling worker still holds the dismantling tool in his right hand, but his body posture changes to a non-raised hand posture, indicating that the dismantling worker held the dismantling tool and made a knocking action during this period. This behavior does not comply with the relevant regulations. The six images from time T0 to time T5 are a set of continuous image sequences, and this set of continuous image sequences is recorded as a corresponding non-standard dismantling behavior.
[0142] In this way, by detecting a continuous image sequence that meets the second preset condition, the number of times that irregular disassembly behaviors occur can be recorded more quickly, making it easier for staff to review videos of disassembling waste electrical appliances.
[0143] In some embodiments, the electronic device further counts the number of times that irregular disassembly behaviors appear in the video file and issues a second prompt message, wherein the second prompt message is used to indicate the number of times that irregular disassembly behaviors appear in the video file.
[0144] For example, each time an irregular disassembly behavior is detected in an electronic device, the total number of irregular disassembly behaviors occurring in the video file, including the irregular disassembly behavior at that time, is counted.
[0145] For another example, after the electronic device completes video file detection, the number of times that irregular disassembly behaviors appear in the video file is counted.
[0146] Optionally, the second prompt information may be different from the first prompt information, or the second prompt information may be included in the first prompt information. Figure 13 As shown, the first prompt information also includes the number of occurrences of non-standard disassembly behaviors by the disassembly personnel.
[0147] In this way, by issuing the second prompt information, the user can be informed of the number of times that irregular disassembly behaviors appear in the video file, making it easier for the user to review the irregular disassembly behaviors in the video file. For example, if a video file corresponds to the work process of a disassembly worker, the user can use the second prompt information to quickly determine the number of times the disassembly worker performed irregular disassembly behaviors.
[0148] The following is a complete description of the control flow of the disassembly behavior detection method provided by this application in conjunction with the above method embodiments:
[0149] like Figure 14As shown, a video file of the dismantling process of a used appliance is used to obtain multiple consecutive frames of images during the dismantling process. Using the trained YOLO-V5 model and YOLO-Pose human key point detection model, each frame of the consecutive multi-frame image is detected and analyzed, obtaining recognition results for human key points and the dismantling tool area in the image. Based on the identified human key points, the dismantling person's posture is analyzed. If, in the multiple consecutive frames of the image, the dismantling person holds the dismantling tool and the posture of the side holding the dismantling tool changes from a raised hand to a non-raised hand posture, or from a non-raised hand posture to a raised hand posture, it indicates that irregular dismantling behavior has occurred. The electronic device then issues a first prompt message to prompt the user to review the irregular dismantling behavior.
[0150] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. It should be easy to realize that the technical goals in this field are combined with the units and algorithm steps of each example described in the embodiments disclosed herein, and the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technical goals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0151] like Figure 15 As shown, the embodiment of the present application further provides an electronic device for executing the disassembly behavior detection method shown in the above method embodiment. The electronic device 300 includes: an acquisition module 301, a processing module 302, a sending module 303 and a receiving module 304.
[0152] Among them, the acquisition module 301 is used to acquire images in the video file; the processing module 302 is used to identify the image and obtain the identification results of the key points of the human hand, the human posture and the disassembly tool area in the image; the processing module 302 is also used to, if multiple consecutive images in the video file all meet the first preset condition, determine whether there is a situation in the multiple images where the human posture changes from the first disassembly posture to the second disassembly posture, and the first preset condition is that the key points of the human hand and the disassembly tool area have overlapping parts; the sending module 303 is used to send a first prompt message when there is a situation in the multiple images where the human posture changes from the first disassembly posture to the second disassembly posture, and the first prompt message is used to indicate that there is irregular disassembly behavior in the video file.
[0153] In one possible implementation, the receiving module 304 is used to receive the user's annotation operation on an area in the image of the video file; the processing module 302 is specifically used to: use the area marked by the annotation operation as the working area; identify the working area in the image, and obtain the recognition results of the key points of the human hand, the human posture and the disassembly tool area in the working area.
[0154] In another possible implementation, the processing module 302 is also used to record a group of continuous image sequences as an irregular disassembly behavior if there is a continuous image sequence that meets a second preset condition among the multiple images. The second preset condition is that the continuous image sequence includes a continuous first preset number of images and a continuous second preset number of images in sequence, the human body posture in the first preset number of images is a hand-raising posture, and the human body posture in the second preset number of images is a non-hand-raising posture, and the images in any two groups of continuous image sequences are different.
[0155] In another possible implementation, the processing module 302 is further configured to count the number of times irregular disassembly behaviors occur in the video file; the sending module 303 is further configured to issue a second prompt message, where the second prompt message is configured to indicate the number of times irregular disassembly behaviors occur in the video file.
[0156] It should be noted that Figure 15 The module division described is illustrative and represents only one logical functional division. Actual implementations may employ different divisions. For example, two or more functions may be integrated into a single processing module. These integrated modules may be implemented as either hardware or software functional modules.
[0157] Another embodiment of the present application further provides an electronic device, such as Figure 16 As shown, electronic device 400 includes memory 401 and processor 402; memory 401 and processor 402 are coupled; memory 401 is used to store computer program code, which includes computer instructions. When processor 402 executes the computer instructions, electronic device 400 performs each step performed by the electronic device in the method flow shown in the above method embodiment.
[0158] In actual implementation, the acquisition module 301, the processing module 302, the sending module 303 and the receiving module 304 can be Figure 16 The processor 402 is shown to implement the computer program code in the memory 401. The specific execution process can be referred to the description of the disassembly behavior detection method above, which will not be repeated here.
[0159] Another embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes each step performed by the electronic device in the method flow shown in the above method embodiment.
[0160] Another embodiment of the present application provides a chip system, which is applied to an electronic device. The chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via wiring. The interface circuits are configured to receive a signal from a memory of the electronic device and send the signal to the processor. The signal includes computer instructions stored in the memory. When the processor of the electronic device executes the computer instructions, the electronic device executes each step performed by the electronic device in the method flow shown in the above method embodiment.
[0161] In another embodiment of the present application, a computer program product is provided. The computer program product includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes each step executed by the electronic device in the method flow shown in the above method embodiment.
[0162] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using a software program, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer execution instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more servers that can be integrated with the medium. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), etc.
[0163] The above is only a specific embodiment of the present application. Those skilled in the art may conceive of changes or substitutions based on the specific embodiment provided in this application, and all such changes or substitutions shall fall within the scope of protection of this application.
Claims
1. A disassembly behavior detection method, characterized in that: Applied to electronic equipment; the method comprises: Get images from video files; Recognize the image to obtain recognition results of key points of a human hand, a human posture, and a disassembly tool area in the image; If the key points of the human hand in the image in the video file overlap with the disassembly tool area, it is determined that the disassembly person is holding the disassembly tool; When multiple consecutive images all show a dismantling person holding a dismantling tool, determining whether the multiple images show a situation in which the posture of the person changes from a first dismantling posture to a second dismantling posture; the first dismantling posture is a hand-raising posture, and the second dismantling posture is a non-hand-raising posture; or the first dismantling posture is the non-hand-raising posture, and the second dismantling posture is the hand-raising posture; In the case that the posture of the human body changes from the first disassembly posture to the second disassembly posture in the multiple images, a first prompt information is issued, where the first prompt information is used to indicate that an irregular disassembly behavior exists in the video file.
2. The method according to claim 1, characterized in that The step of recognizing the image to obtain recognition results of key points of a human hand, a human posture, and a disassembly tool area in the image includes: receiving a user's marking operation on an area in an image of the video file, and using the area marked by the marking operation as a working area; The working area in the image is identified to obtain identification results of key points of human hands, human postures and disassembly tool areas in the working area.
3. The method according to claim 1, characterized in that The hand-raising posture is: on the side of the key point of the human hand that overlaps with the disassembly tool area, the angle formed by the hand key point, the elbow key point and the shoulder key point is within the preset angle range, and the position of the hand key point is higher than the position of the elbow key point.
4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: If there is a continuous image sequence among the multiple images that meets the second preset condition, a group of the continuous image sequences will be recorded as an irregular disassembly behavior, and the second preset condition is that the continuous image sequence includes a continuous first preset number of images and a continuous second preset number of images in sequence, the human body posture in the first preset number of images is a hand-raising posture, and the human body posture in the second preset number of images is a non-hand-raising posture, and the images in any two groups of the continuous image sequences are different.
5. The method according to claim 1, wherein The method further comprises: The number of times the irregular disassembly behavior appears in the video file is counted, and a second prompt message is issued, where the second prompt message is used to indicate the number of times the irregular disassembly behavior appears in the video file.
6. An electronic device, characterized in that: include: An acquisition module is used to acquire images from video files; a processing module, configured to recognize the image and obtain recognition results of key points of a human hand, a human posture, and a disassembly tool area in the image; The processing module is further configured to determine that the dismantling person is holding a dismantling tool if the key points of the human hand in the image in the video file overlap with the dismantling tool area; When multiple consecutive images all show a dismantling person holding a dismantling tool, determining whether the multiple images show a situation in which the posture of the person changes from a first dismantling posture to a second dismantling posture; the first dismantling posture is a hand-raising posture, and the second dismantling posture is a non-hand-raising posture; or the first dismantling posture is the non-hand-raising posture, and the second dismantling posture is the hand-raising posture; The sending module is configured to send a first prompt message when the human body posture changes from the first disassembly posture to the second disassembly posture in the multiple images, wherein the first prompt message is used to indicate that an irregular disassembly behavior exists in the video file.
7. The electronic device according to claim 6, wherein: The electronic device further includes a receiving module configured to receive a user's annotation operation on an area in an image of the video file; the processing module specifically configured to: use the area marked by the annotation operation as a working area; and identify the working area in the image to obtain recognition results of key points of a human hand, a human posture, and a disassembly tool area in the working area; The hand-raising posture is: on the side of the key point of the human hand that overlaps with the disassembly tool area, the angle formed by the key point of the hand, the key point of the elbow and the key point of the shoulder is within a preset angle range, and the position of the key point of the hand is higher than the position of the key point of the elbow; The processing module is further configured to record a group of continuous image sequences as an irregular disassembly behavior if a continuous image sequence that satisfies a second preset condition exists among the plurality of images, wherein the second preset condition is that the continuous image sequence includes, in sequence, a first preset number of continuous images and a second preset number of continuous images, the human body posture in the first preset number of images is a hand-raising posture, and the human body posture in the second preset number of images is a non-hand-raising posture, and the images in any two groups of the continuous image sequences are different; The processing module is further configured to count the number of times the irregular disassembly behavior occurs in the video file; the sending module is further configured to issue a second prompt message, wherein the second prompt message is configured to indicate the number of times the irregular disassembly behavior occurs in the video file.
8. An electronic device, characterized in that: include: one or more processors; one or more memories; The one or more memories are used to store computer program codes, and the computer program codes include computer instructions. When the one or more processors execute the computer instructions, the electronic device executes the disassembly behavior detection method according to any one of claims 1 to 5.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed on a computer, the computer is caused to execute the disassembly behavior detection method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and device for determining disassembly information of electronic product, electronic equipment, and system
CN114005071A
Method for judging continuous action process through visual identification
CN114533038A