Target behavior recognition method, device and equipment based on adaptive perception domain
Patent Information
- Application Number
- CN202610879636.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-04
AI Technical Summary
员工在工作时间频繁使用手机或接打电话,存在极大安全隐患,员工在分心状态下易忽视设备警示信号、操作失误或误触危险区域,导致伤害、物料泄漏等事故
[0021] The aforementioned target behavior recognition method, device, and equipment based on adaptive perceptual domain, after acquiring the image to be detected, determines the human body detection box and human body key point set for each target object through target detection. For each target object, an adaptive perceptual domain is constructed based on the corresponding human body detection box and human body key point set, thereby accurately narrowing the potential occurrence area of mobile phone usage behavior for each target object and effectively reducing the detection range. Then, the positional relationship between the hand position and the mobile phone detection box is combined to complete the initial judgment. Based on the initial judgment result, it is determined whether further secondary judgment based on ear key points is needed. Through hierarchical and progressive multi-dimensional positional association, various similar mobile phone usage behaviors such as making calls and playing with mobile phones are distinguished. In the above recognition process, by integrating human posture, key points of multiple parts, and relative positional relationships for multi-dimensional comprehensive judgment, the probability of misjudgment and missed judgment is greatly reduced, and the overall recognition accuracy and robustness of mobile phone usage behavior are significantly improved.
Smart Images

Figure CN122695686A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of target recognition technology, and in particular to a target behavior recognition method, apparatus and device based on adaptive perception domain. Background Technology
[0002] Against the backdrop of intelligent transformation in manufacturing, factories, as the core carriers of industrial production, directly impact a company's competitiveness and industry reputation through their operational efficiency, safe production, and product quality. However, with the widespread use of mobile devices, employees' unauthorized use of mobile phones during work hours (such as watching videos or chatting) or making phone calls poses significant safety hazards and is gradually becoming a "hidden pain point" that interferes with the refined management of factories.
[0003] Because factory workshops are typically complex environments with dense workforces, high-speed equipment operation, and tightly integrated processes, they demand extremely high levels of focus and operational standardization from employees. Frequent use of mobile phones or making calls during work hours poses significant safety hazards. Distracted employees are prone to ignoring equipment warning signals, making operational errors, or accidentally touching dangerous areas, leading to injuries, material leaks, and other accidents. According to the "China Industrial Safety Annual Report," over 30% of minor injuries in the manufacturing industry in the past three years are related to employee distraction, with mobile phone interference listed as the second leading cause. Furthermore, even brief mobile phone use can disrupt standardized work rhythms (such as assembly, quality inspection, and logistics transfer), causing delays in process connections, reducing overall operational efficiency, and impacting the overall equipment efficiency of the production line. Moreover, in critical processes (such as precision assembly and parameter calibration), employee errors due to mobile phone interference can lead to missed or incorrect inspections, resulting in batches of defective products, increased rework costs, and even customer complaints.
[0004] In related technologies, images are typically used to identify employee postures, but this method suffers from low accuracy. Summary of the Invention
[0005] Therefore, it is necessary to provide a target behavior recognition method, device, and equipment based on an adaptive perceptual domain that can improve the accuracy of recognizing mobile phone usage behavior, addressing the aforementioned technical problems.
[0006] Firstly, this application provides a target behavior recognition method based on an adaptive perceptual domain, including:
[0007] Acquire the image to be detected;
[0008] Target detection is performed based on the image to be detected to determine the human body detection box and the set of human body key points for each target object in the image to be detected.
[0009] For each target object in the image to be detected, the perceptual domain of the target object in the image to be detected is determined based on the human body detection box and the human body key point set; the perceptual domain of the target object is used to describe the potential area where the target behavior of the target object may occur.
[0010] When the hand key point in the human body key point set is located within the perceptual domain and a mobile phone object is identified from the image region corresponding to the human body detection box, the initial behavior recognition result is determined based on the first positional relationship between the mobile phone object and the hand of the target object.
[0011] If the initial behavior recognition result indicates that the first positional relationship satisfies the preset rule, the target behavior recognition result of the target object is determined based on the second positional relationship between the mobile phone detection frame of the mobile phone object and the ear key point in the human body key point set.
[0012] Secondly, this application also provides a target behavior recognition device based on an adaptive perceptual domain, comprising:
[0013] The acquisition module is used to acquire the image to be detected;
[0014] The detection module is used to perform target detection based on the image to be detected, and to determine the human body detection box and the set of human body key points for each target object in the image to be detected;
[0015] The perceptual domain determination module is used to determine the perceptual domain of each target object in the image to be detected, based on the human body detection box and the human body key point set; the perceptual domain of the target object is used to describe the potential occurrence area of the target object's target behavior;
[0016] The first recognition module is used to determine an initial behavior recognition result based on a first positional relationship between the mobile phone object and the hand of the target object when the hand key point in the human body key point set is located within the perception domain and a mobile phone object is recognized from the image area corresponding to the human body detection box.
[0017] The second recognition module is used to determine the target behavior recognition result of the target object based on the second positional relationship between the mobile phone detection frame of the mobile phone object and the ear key point in the human body key point set, when the initial behavior recognition result indicates that the first positional relationship meets the preset rules.
[0018] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.
[0019] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0020] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0021] The aforementioned target behavior recognition method, device, and equipment based on adaptive perceptual domain, after acquiring the image to be detected, determines the human body detection box and human body key point set for each target object through target detection. For each target object, an adaptive perceptual domain is constructed based on the corresponding human body detection box and human body key point set, thereby accurately narrowing the potential occurrence area of mobile phone usage behavior for each target object and effectively reducing the detection range. Then, the positional relationship between the hand position and the mobile phone detection box is combined to complete the initial judgment. Based on the initial judgment result, it is determined whether further secondary judgment based on ear key points is needed. Through hierarchical and progressive multi-dimensional positional association, various similar mobile phone usage behaviors such as making calls and playing with mobile phones are distinguished. In the above recognition process, by integrating human posture, key points of multiple parts, and relative positional relationships for multi-dimensional comprehensive judgment, the probability of misjudgment and missed judgment is greatly reduced, and the overall recognition accuracy and robustness of mobile phone usage behavior are significantly improved. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is an application environment diagram of a target behavior recognition method based on an adaptive perceptual domain in one embodiment.
[0024] Figure 2 This is a flowchart illustrating a target behavior recognition method based on an adaptive perceptual domain in one embodiment;
[0025] Figure 3 This is a schematic diagram of the confirmation process for the initial behavior recognition result in one embodiment;
[0026] Figure 4 This is a structural block diagram of a target behavior recognition device based on an adaptive perceptual domain in one embodiment;
[0027] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0029] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0030] The target behavior recognition method based on adaptive perceptual domain provided in this application can be applied to, for example... Figure 1 The application environment shown is illustrated. Terminal 101 communicates with server 102 via a network. A data storage system can store the data that server 102 needs to process. The data storage system can be integrated onto server 102, or it can be located in the cloud or on another network server.
[0031] Administrators can initiate a target behavior recognition task through terminal 101. Server 102 responds to the target behavior recognition task by acquiring the image to be detected. Based on the image to be detected, target detection is performed to determine the human body detection box and human body key point set for each target object in the image to be detected. For each target object in the image to be detected, the perceptual domain of the target object in the image to be detected is determined based on the human body detection box and human body key point set. The perceptual domain of the target object is used to describe the potential occurrence area of the target object's target behavior. If the hand key point in the human body key point set is located within the perceptual domain and the mobile phone object is identified from the image area corresponding to the human body detection box, the initial behavior recognition result is determined based on the first positional relationship between the mobile phone object and the hand of the target object. If the initial behavior recognition result indicates that the first positional relationship satisfies a preset rule, the target behavior recognition result of the target object is determined based on the second positional relationship between the mobile phone object's mobile phone detection box and the ear key point in the human body key point set.
[0032] Terminal 101 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, and projection equipment. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0033] In one exemplary embodiment, such as Figure 2 As shown, a target behavior recognition method based on adaptive perceptual domain is provided, which is then applied to... Figure 1 Taking the server in the example, the explanation includes the following steps 201 to 205. Wherein:
[0034] Step 201: Obtain the image to be detected.
[0035] The image to be detected is the image collected for the target object; the target object can be a worker in a workshop, factory area or other work environment.
[0036] In some embodiments, the original image acquired by the image acquisition device can be used as the image to be detected.
[0037] In some embodiments, preprocessing operations such as cropping, denoising, brightness correction, and image scaling can be performed sequentially on the acquired original image, and the processed image can be used as the final image to be detected, thereby reducing the interference of factors such as image noise and light and shadow changes on the subsequent detection and recognition process.
[0038] In some embodiments, the image to be detected may be acquired in response to a target signal, wherein the target signal refers to a signal that initiates target behavior recognition.
[0039] In some embodiments, the target signal may be automatically triggered based on production status information; for example, when the production status information indicates that there is a production anomaly alarm within a preset time range from the current time, the target signal is triggered; or, for example, when the production status information indicates that there is a declining trend in production efficiency within a preset time range from the current time and the decline exceeds a preset threshold, the target signal is triggered.
[0040] In other embodiments, the target signal can also be manually triggered. For example, managers can manually issue control commands through devices such as a back-end management terminal or a field control panel. Alternatively, the target signal can be triggered periodically according to a preset frequency to achieve all-time detection.
[0041] Step 202: Perform target detection based on the image to be detected, and determine the human body detection box and human body key point set of each target object in the image to be detected.
[0042] In some embodiments, a target recognition algorithm can be used to perform global traversal recognition on the entire image to be detected, thereby achieving the classification and localization of human targets and obtaining the human detection box and human key point set of each target object in the image to be detected.
[0043] The human detection bounding box of the target object is used to define the effective area of a single human target in the image, thereby isolating background interference areas through the human detection bounding box.
[0044] The set of key points of the target object contains various types of key points, such as the pixel coordinates of key parts such as ear key points, head key points, neck key points, shoulder key points, left and right wrist key points, and left and right hand key points. The spatial distribution of each key point can be used to infer the posture of the human body.
[0045] Step 203: For each target object in the image to be detected, the perceptual domain of the target object in the image to be detected is determined based on the human body detection box and the set of human body key points; the perceptual domain of the target object is used to describe the potential area where the target behavior of the target object may occur.
[0046] It is understandable that the human body detection box encompasses the entire target object, thus defining the maximum area where the target object's target behavior can occur. However, due to the characteristics of the target behavior itself, there will be areas within the human body detection box where the target behavior is impossible. Therefore, this application uses the human body detection box and the set of human body key points to determine the receptive domain of the target object in the image to be detected. By using the receptive domain of the target object, the potential area where the target object's target behavior can occur can be further defined, thereby effectively simplifying the detection range and improving detection accuracy.
[0047] For example, when the target behavior is mobile phone usage, such as making a call or playing on the phone, this behavior requires interaction with the phone through the hand. Simultaneously, it requires raising the hand to bring the phone closer to the eyes for viewing. Therefore, the lower boundary of the perceptual domain can be determined based on the elbow keypoint, the upper boundary based on the eye or nose keypoint, and the left and right boundaries based on the human body detection box. Thus, the perceptual domain is constructed based on the lower, upper, and left / right boundaries. Clearly, if mobile phone usage exists, it will necessarily be transmitted within the perceptual domain. It should be noted that up, down, left, and right are defined with reference to the target object; for example, the direction of the human head is considered up, and the direction of the human feet is considered down.
[0048] Step 204: If the hand key points in the human body key point set are located within the perceptual domain and the mobile phone object is identified from the image region corresponding to the human body detection box, the initial behavior recognition result is determined based on the first positional relationship between the mobile phone object and the hand of the target object.
[0049] In the embodiments of this application, the target behavior is the interaction between the target object and the mobile phone object. As mentioned above, the completion of this behavior requires raising the hand to bring the mobile phone close to the eyes for browsing. The key points of the hand are located within the perceptual domain, which means that the target object has raised its hand, that is, the target behavior may occur.
[0050] Based on this, mobile phone objects can be identified in the image area corresponding to the human body detection box, and then the initial behavior recognition result can be determined according to the first positional relationship between the mobile phone object and the hand of the target object.
[0051] Understandably, the perceptual domain is a small area defined by key points on the human body that may potentially lead to a behavior. In real-world scenarios, the outline of a mobile phone may extend beyond the boundaries of the perceptual domain due to factors such as the angle at which the target is held or limb obstruction. If detection is only performed within the perceptual domain, it is easy to miss detections. However, the human detection box completely surrounds the entire human target and personal belongings. Therefore, recognizing mobile phone objects through the image area corresponding to the human detection box can ensure that the mobile phone target is completely captured, avoiding recognition failures caused by target truncation.
[0052] In some embodiments, the first positional relationship between the mobile phone object and the hand of the target object can be the distance between the center point of the mobile phone object and the center point of the hand, the degree of intersection between the detection box of the mobile phone object and the detection box of the hand, or the inclusion relationship between the detection box of the mobile phone object and the center point of the hand.
[0053] In other embodiments, it may involve simultaneously detecting whether the hand key points in the human body key point set are located within the perceptual domain, and detecting whether the image region corresponding to the human body detection box recognizes the mobile phone object.
[0054] In other embodiments, it may be possible to first detect whether the hand key points in the human body key point set are located within the perceptual domain. If they are located within the perceptual domain, then mobile phone object detection is performed on the image area corresponding to the human body detection box; if they are not located within the perceptual domain, then it is determined that there is no mobile phone usage behavior.
[0055] It should be noted that the target object has left-hand key points and right-hand key points. If any key point is located within the perceptual field, it is considered that the hand key point is located within the perceptual field.
[0056] In some embodiments, the initial behavior recognition result may be determined based on the first positional relationship between the mobile phone object and the hand of the target object when the hand key points in the human body key point set are located within the perceptual domain and the mobile phone object is identified from the image region corresponding to the perceptual domain.
[0057] Step 205: If the initial behavior recognition result indicates that the first positional relationship meets the preset rules, the target behavior recognition result of the target object is determined based on the second positional relationship between the mobile phone detection box of the mobile phone object and the ear key point in the human body key point set.
[0058] In some embodiments, the first positional relationship satisfies a preset rule, indicating that the initial behavior recognition result is that there is mobile phone usage behavior. Based on this, the second positional relationship between the mobile phone detection box of the mobile phone object and the ear key point in the human body key point set can be used to specifically determine whether the mobile phone usage behavior is a phone call behavior or a mobile phone entertainment behavior (i.e., playing with the mobile phone), thereby obtaining the target behavior recognition result of the target object.
[0059] The second positional relationship can be represented in various ways. For example, it can be the distance between the ear key point and the center point of the mobile phone detection frame, or it can be the inclusion relationship between the ear key point and the mobile phone detection frame.
[0060] For example, if the ear key point in the human body key point set is located within the mobile phone detection frame, it means that the mobile phone object is close to the ear, and thus the target behavior recognition result of the target object is the act of making a phone call; otherwise, the target behavior recognition result of the target object is the act of mobile phone entertainment.
[0061] In the aforementioned target behavior recognition method based on adaptive perceptual domain, after acquiring the image to be detected, the human body detection box and human body key point set of each target object are determined through target detection. For each target object, an adaptive perceptual domain is constructed based on the corresponding human body detection box and human body key point set, thereby accurately narrowing the potential occurrence area of mobile phone usage behavior for each target object and effectively reducing the detection range. Then, the positional relationship between the hand position and the mobile phone detection box is combined to complete the initial judgment. Based on the initial judgment result, it is determined whether further secondary judgment based on ear key points is needed. Through hierarchical and progressive multi-dimensional positional association, various similar mobile phone usage behaviors such as making calls and playing with mobile phones are distinguished. In the above recognition process, by integrating human posture, key points of multiple parts, and relative positional relationships for multi-dimensional comprehensive judgment, the probability of misjudgment and missed judgment is greatly reduced, and the overall recognition accuracy and robustness of mobile phone usage behavior are significantly improved.
[0062] In one exemplary embodiment, such as Figure 3 As shown, Figure 3 A schematic diagram of the determination process for the initial behavior recognition result in an embodiment of this application is shown. Step 204 includes steps 301 to 304. Wherein:
[0063] Step 301: If the hand key points in the human body key point set are located within the perceptual domain, crop the image to be detected based on the human body detection box to obtain the cropped image.
[0064] It is understandable that if the hand key points in the human body key point set are located within the perceptual domain, it indicates that mobile phone use behavior may occur. Therefore, by cropping the image to be detected using the human body detection bounding box, a cropped image can be obtained. On the one hand, this can reduce the range of subsequent hand recognition and mobile phone object recognition, and on the other hand, it can avoid interference from other image areas for subsequent recognition.
[0065] Image cropping can be achieved through image coordinate extraction algorithms. For example, based on the pixel coordinates and width and height parameters of the human body detection box, the image data of the corresponding area is extracted from the original image to be detected to generate a cropped image.
[0066] Step 302: Perform hand recognition on the cropped image to obtain a hand detection box.
[0067] In some embodiments, image detection algorithms can be used to identify hands in cropped images to obtain hand detection boxes.
[0068] In other embodiments, the coordinates of the hand key points located in the perceptual domain can be mapped to the coordinate system of the cropped image to obtain the cropping coordinates of the hand key points in the cropped image. Then, the image detection model performs hand recognition based on the cropped image and the cropping coordinates of the hand key points to obtain a hand detection box. Thus, the cropping coordinates of the hand key points can be used to guide the image detection model to perform detection, further improving the hand positioning accuracy.
[0069] Step 303: Perform mobile phone object recognition on the cropped image to obtain the mobile phone detection box.
[0070] In some embodiments, image detection algorithms can be used to identify mobile phone objects in the cropped image to obtain a mobile phone detection box.
[0071] In other embodiments, the area can be defined in the cropped image based on the hand detection box to obtain a defined region, and the mobile phone object can be recognized within the defined region to obtain a mobile phone detection box; for example, the region can be extended outward based on the hand detection box, for example, the length and width of the hand detection box can be extended by a preset multiple, and the extended region can be determined as the defined region.
[0072] Step 304: Determine the initial behavior recognition result of the target object based on the area intersection ratio between the hand detection box and the mobile phone detection box.
[0073] The initial behavior recognition result of the target object is used to indicate whether the target object has mobile phone usage behavior.
[0074] For example, if the area overlap ratio between the hand detection box and the mobile phone detection box is greater than a preset ratio threshold, the initial behavior recognition result of the target object is determined to be that there is mobile phone usage behavior; otherwise, the initial behavior recognition result of the target object is determined to be that there is no mobile phone usage behavior.
[0075] In this embodiment, the detection area is compressed by image cropping, reducing the amount of computational data and accelerating the detection rate. Simultaneously, image cropping removes irrelevant image interference, improving the detection effect of the hand and phone. The positional relationship between the hand and phone is quantified using the intersection ratio of the detection box areas, providing an objective and clear judgment standard, effectively reducing the false judgment rate of behavior recognition, and further improving the overall accuracy of mobile phone usage behavior detection.
[0076] In an exemplary embodiment, the initial behavior recognition result of the target object is determined based on the area intersection ratio between the hand detection bounding box and the mobile phone detection bounding box, including at least one of the following:
[0077] If the area overlap ratio between the hand detection box and the mobile phone detection box is greater than a preset ratio threshold, the initial behavior recognition result is determined to be that mobile phone usage behavior exists.
[0078] It is understandable that if the area overlap ratio between the hand detection box and the mobile phone detection box is greater than the preset ratio threshold, it means that the target object is holding a mobile phone. Combined with the hand-raising action determined in the previous steps, the initial behavior recognition result can be determined as the presence of mobile phone use behavior.
[0079] If the area overlap ratio between the hand detection box and the mobile phone detection box is not greater than a preset ratio threshold, the initial behavior recognition result is determined to be that there is no mobile phone usage behavior.
[0080] It is understandable that if the area overlap ratio between the hand detection box and the phone detection box is not greater than the preset ratio threshold, it means that the phone and the target object's hand may only be spatially adjacent, such as a spatial visual misalignment, and the target object does not form a grip on the phone. Therefore, it can be determined that the initial behavior recognition result is that there is no phone usage behavior.
[0081] In this embodiment, the spatial relationship between the hand and the phone is quantified by calculating the area intersection ratio of the hand detection frame and combining it with a preset threshold to determine the behavior. Compared to the traditional method that relies solely on coordinate position, this determination rule is objective and standardized, effectively distinguishing between different states such as holding, approaching, and separating, significantly reducing the probability of misjudgment. Furthermore, the determination logic is simple and clear, with low computational load, ensuring both recognition accuracy and algorithm efficiency, thus meeting the application requirements of real-time detection.
[0082] In an exemplary embodiment, for each target object in the image to be detected, the receptive domain of the target object in the image to be detected is determined based on the human body detection bounding box and the set of human body key points, including:
[0083] For each target object in the image to be detected, the first boundary is determined based on the maximum ordinate of the left elbow keypoint and the right elbow keypoint in the human body keypoint set.
[0084] It is understandable that when there is a hand raising motion, the hand will be above the elbow. Therefore, the first boundary is determined by the maximum value of the ordinate of the key points of the left and right elbows.
[0085] The first boundary is the lower boundary of the perceptual field. As mentioned above, the upper, lower, left, and right boundaries of the perceptual field are determined with reference to the target object. The direction of the target object's head is the upper boundary, the direction of its feet is the lower boundary, and the direction of its left and right hands is the left and right boundary.
[0086] The second boundary of the perceptual domain is determined based on the ordinate of the nose key point in the human body key point set.
[0087] The second boundary is also the upper boundary of the perceptual domain.
[0088] It should be noted that the determination of the key point of the nose is based on the actual mobile phone usage behavior. Usually, when using a mobile phone, you need to look down, and the mobile phone is usually used at a position lower than the human face.
[0089] In other embodiments, the ordinate of the eye key point or the ordinate of the top of the head key point in the human body key point set may be used as the second boundary of the perceptual domain.
[0090] Based on the boundaries of the human detection box, the first boundary, and the second boundary, the receptive domain of the target object in the image to be detected is determined.
[0091] In some embodiments, the left and right boundaries of the human detection box can be reused to construct the perceptual domain of the target object in the image to be detected together with the first boundary and the second boundary.
[0092] In other embodiments, the left and right boundaries of the human detection box can be extended, such as by widening or scaling by a preset factor, to obtain the left and right boundaries of the perceptual field. Then, the perceptual field of the target object in the image to be detected can be constructed together with the first boundary, the second boundary and the left and right sides of the perceptual field.
[0093] It should be noted that both the human detection box and the receptive field are rectangular boxes. The first and second boundaries, which serve as the upper and lower boundaries of the receptive field, are parallel to each other. Similarly, the left and right boundaries of the human detection box are also parallel to each other and perpendicular to either the first or second boundary.
[0094] In this embodiment, the perceptual domain is defined by combining the coordinate information of key points on the human elbow and nose with a human detection bounding box, accurately locating high-probability areas where mobile phone usage behavior occurs. By defining the upper and lower boundaries using key points on the human body and determining the left and right boundaries using the human detection bounding box, invalid detection areas within the human body region can be effectively eliminated, significantly reducing the computational area for subsequent target recognition and lowering computational overhead. Simultaneously, defining the region based on human motion features reduces interference from background and irrelevant human body parts, improving the overall accuracy and algorithm efficiency of hand and mobile phone detection and behavior recognition from the source.
[0095] In an exemplary embodiment, if the initial behavior recognition result indicates that the first positional relationship satisfies a preset rule, the target behavior recognition result of the target object is determined based on the second positional relationship between the mobile phone object's detection frame and the ear keypoint in the human body keypoint set, including:
[0096] If the initial behavior recognition result indicates that the first positional relationship meets the preset rule, and the ear key point is located in the mobile phone detection box of the mobile phone object, the target behavior recognition result of the target object is determined to be making a phone call; the preset rule is that the positional intersection between the mobile phone object and the hand of the target object is greater than the preset ratio threshold.
[0097] Understandably, the initial behavior recognition result indicates that the first positional relationship meets the preset rules, that is, the initial behavior recognition result is that there is mobile phone usage behavior. Based on this, the ear key point is located in the mobile phone detection box of the mobile phone object, which means that the mobile phone is close to the ear, and it is likely to be making a call. Therefore, the target behavior recognition result of the target object is making a call.
[0098] If the initial behavior recognition result indicates that the positional relationship between the target object's hand and the mobile phone object meets the preset rules, and the ear key point is not located in the mobile phone detection box of the mobile phone object, then the target object's target behavior recognition result is determined to be mobile phone entertainment behavior.
[0099] Understandably, the initial behavior recognition result indicates that the first positional relationship meets the preset rules, that is, the initial behavior recognition result is that there is mobile phone use behavior. Based on this, the ear key point is not located in the mobile phone detection box of the mobile phone object, which means that the mobile phone is not close to the ear and the probability of using the mobile phone to make a call is low. Therefore, the target behavior recognition result of the target object is mobile phone entertainment behavior.
[0100] In this embodiment, based on the initial behavior recognition results, it is preliminarily determined that there is mobile phone usage behavior. Then, it is further distinguished by combining the positional relationship between the mobile phone detection frame and the key points on the ear, so as to accurately distinguish between making phone calls and mobile phone entertainment, thereby improving the recognition accuracy of mobile phone usage behavior.
[0101] In some embodiments, target detection is performed based on the image to be detected, determining the human body detection bounding box and the set of human body key points for each target object in the image to be detected, including:
[0102] Human body recognition is performed based on the image to be detected, and human body detection boxes of each target object in the image to be detected are obtained.
[0103] The human body detection box is a rectangular box, which is determined based on the maximum outer contour of the target object. It covers the entire body area of the target object and can include the human body and personal items such as mobile phones.
[0104] In some embodiments, human recognition can be achieved through human detection models; for example, using mainstream object detection networks such as YOLO (You Only Look Once, an object detection algorithm), RetinaNet, and Faster R-CNN (Faster Region-based Convolutional Neural Network), the image to be detected is taken as input, and through forward inference of the network, the position coordinates of all human targets in the image are located, and finally the corresponding human detection boxes are generated.
[0105] Human keypoints are detected in the image to be detected based on preset keypoint types, resulting in an initial set of human keypoints for each target object in the image to be detected.
[0106] The preset key point types can include eyes, hands, ears, elbows, nose, etc.
[0107] In some embodiments, human keypoint detection can be achieved through a human pose estimation network; for example, the image to be detected is input into a trained pose estimation model, the model extracts local human features and global related features, locates and outputs the pixel coordinates of each preset type of human keypoint, forming an initial set of human keypoints.
[0108] For each target object, if there are missing key points in the initial human key point set, key point prediction is performed based on the initial human key point set and the image to be detected to obtain the human key point set of the target object.
[0109] Understandably, due to factors such as shooting angle, limb occlusion, image blur and posture distortion, key point detection models are prone to missing or missing some key points. Therefore, when the initial key points are not fully detected, the coordinate information of the detected key points can be combined with the human features of the image to be detected to complete the reasoning and prediction of the missing key points, fill in the missing points and update the key point data to obtain a complete set of human key points.
[0110] In some embodiments, if any keypoint of any preset keypoint type is empty in the initial human body keypoint set, then it is determined that there is a keypoint missing in the initial human body keypoint set.
[0111] In the above embodiments, a human body detection box is generated by human body recognition, an initial set of human body key points is obtained by human body key point detection, and key points are filled in if there are missing points in the initial set of human body key points, so as to form a complete human body detection and key point output result, providing reliable data support for subsequent processing.
[0112] In one exemplary embodiment, the target behavior recognition method based on adaptive perceptual domain further includes:
[0113] Obtain the historical target behavior recognition results corresponding to each of the multiple historical images that have temporal correlation.
[0114] Among them, multi-frame historical images with temporal correlation refer to multi-frame historical images that are correlated in time; for example, multi-frame historical images can be images acquired continuously, or they can be obtained by random sampling from a set of continuously acquired images.
[0115] It is understandable that a single image can only reflect instantaneous behavior at a single moment, which is sporadic. Generating anomaly reports based solely on the target behavior recognition results of a single image can easily lead to misjudgments. In contrast, the historical target behavior recognition results of multiple historical images with temporal correlation reflect continuous behavior over a period of time. This can effectively avoid interference from instantaneous actions and accidental images, thereby improving the reliability of behavior determination.
[0116] Data statistics are performed based on the historical target behavior recognition results corresponding to each of the multiple historical images to determine the number of times the same target object has performed a target behavior.
[0117] The target behavior is mobile phone usage behavior, which includes mobile phone entertainment behavior and phone call behavior.
[0118] In some embodiments, the number of times each type of mobile phone use behavior is counted separately, such as the number of times a call is made or the number of times a user plays on their phone; or the total number of mobile phone use behaviors can be counted.
[0119] An anomaly report is generated when the number of times the target behavior occurs on the same target object exceeds a threshold.
[0120] The anomaly report may include information such as the target object's identifier, the type of target behavior, the time of occurrence, the statistical period, and the number of statistical occurrences.
[0121] In other embodiments, the mobile phone contact information of each target object is pre-stored. After an anomaly report is generated, the corresponding mobile phone contact information is determined according to the identifier of the target object in the anomaly report, and a prompt message is sent to the mobile phone of the corresponding target object according to the contact information. At the same time, starting from the time of sending the prompt message, multiple consecutive checks are performed on the target object within a preset time period. If the target behavior is still detected in multiple checks, a management prompt message is sent to the management personnel so that the management personnel can make manual intervention according to the management prompt message.
[0122] In this embodiment, statistical analysis is carried out by combining the historical recognition results of multiple frames of historical images with temporal correlation, counting the frequency of target object behavior and comparing it with a preset threshold, thereby triggering the generation of anomaly reports, which can avoid misjudgment caused by single-frame recognition.
[0123] To facilitate understanding, a specific embodiment will be used as an example below to illustrate the concept of safety behavior detection in a factory.
[0124] First, raw image data is collected in real time based on a multi-source camera network deployed inside the factory building. The raw image data is then denoised to obtain target image data, which is the image to be detected in the aforementioned embodiment. This helps to improve image quality and the effectiveness and accuracy of subsequent image recognition.
[0125] In some embodiments, due to continuous monitoring by multiple cameras, the operating temperature of the cameras is relatively high, and the field of view of the cameras is dark and the brightness is uneven, resulting in low quality of the original image data. Therefore, noise processing of the original image can be performed based on Gaussian filtering. For example, the Gaussian filtering function in OpenCV (Open Source Computer Vision Library) can be used to denoise the original image.
[0126] Secondly, based on object detection algorithms, the coordinates of key points such as ears, nose, and elbows are selected to obtain the bounding box coordinates of each person in the image. and the coordinates of the left and right ears are , The coordinates of the left and right elbows are , Left and right hand coordinates are , and the coordinates of the nose are .
[0127] Based on the coordinates of the human detection frame The coordinates of the left and right elbows are , and the coordinates of the nose are Construct the perception threshold of each human body detection frame ,in, = , = , = , =max( , ).
[0128] Next, based on the left and right hand coordinates, it is determined whether there are mobile phone behavior features within the perceptual domain. Specifically, if... or If any hand is located within the perceptual field, it is considered that the perceptual threshold of the human body detection box contains mobile phone behavior characteristics.
[0129] Next, human bounding boxes with mobile phone behavior characteristics are selected based on the perception threshold. The original image is then cropped based on the coordinates of the human bounding boxes. For example, a 640×640 canvas is created, which is a completely black image. The human bounding box area is placed in the upper left corner. The final detected image is a black image with the upper left corner as the human bounding box. The constructed image is sent to an image detection algorithm for recognition. The target is the hand and the mobile phone bounding box. The ratio of the intersection of the mobile phone bounding box and the hand bounding box to the area of the mobile phone bounding box is determined. If the ratio is greater than 90%, it is considered that a mobile phone has been identified to avoid false alarms. If a mobile phone is present, it is considered that the person is engaging in mobile phone behavior. Further determination is needed to determine whether the person is playing on the mobile phone or making a call. The mobile phone bounding box is mapped to the original image, and the mapped coordinates are the coordinates of the mobile phone bounding box plus the coordinates of the upper left corner of the human bounding box.
[0130] Finally, based on the characteristic representation of phone call behavior, mobile phone usage behavior is divided into playing on the phone and making phone calls, specifically:
[0131] Mobile phone usage behavior is categorized into playing on the phone and making a call. The rule for determining making a call is as follows: Select the target bounding box of the mobile phone that matches the perceptual domain with mobile phone behavior characteristics. Based on the coordinates of the left and right ears of the corresponding human posture point and the coordinates of the target bounding box, determine whether the ear point coordinates exist in the target bounding box. If they exist, the mobile phone behavior in the video frame is determined to be playing on the phone; if they do not exist, the mobile phone behavior in the video frame is determined to be making a call.
[0132] In some embodiments, when a frame is determined to be a phone call or mobile phone use, the video is continuously monitored by extracting frames, with 5 frames extracted per second for real-time monitoring; if within 30 seconds, i.e., within 150 frames, there are 10 non-consecutive frames that all show the same mobile phone behavior (using the phone or making a phone call), then it is marked as having the mobile phone behavior, and a report is made based on the determined mobile phone behavior.
[0133] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0134] Based on the same inventive concept, this application also provides an adaptive perceptual domain-based target behavior recognition device for implementing the aforementioned adaptive perceptual domain-based target behavior recognition method. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the adaptive perceptual domain-based target behavior recognition device provided below can be found in the limitations of the adaptive perceptual domain-based target behavior recognition method described above, and will not be repeated here.
[0135] In one exemplary embodiment, such as Figure 4 As shown, a schematic diagram of a target behavior recognition device based on an adaptive perceptual domain is provided. The target behavior recognition device 500 based on an adaptive perceptual domain includes:
[0136] The acquisition module 501 is used to acquire the image to be detected.
[0137] The detection module 502 is used to perform target detection based on the image to be detected, and to determine the human body detection box and the set of human body key points for each target object in the image to be detected.
[0138] The perceptual domain determination module 503 is used to determine the perceptual domain of each target object in the image to be detected based on the human body detection box and the set of human body key points; the perceptual domain of the target object is used to describe the potential area where the target behavior of the target object may occur.
[0139] The first recognition module 504 is used to determine the initial behavior recognition result based on the first positional relationship between the mobile phone object and the hand of the target object when the hand key point in the human body key point set is located within the perceptual domain and the mobile phone object is recognized from the image area corresponding to the human body detection box.
[0140] The second recognition module 505 is used to determine the target behavior recognition result of the target object based on the second positional relationship between the mobile phone detection frame of the mobile phone object and the ear key point in the human body key point set, when the initial behavior recognition result indicates that the first positional relationship meets the preset rules.
[0141] In an exemplary embodiment, the first recognition module 504 is configured to, when the hand key points in the human body key point set are located within the perceptual domain, crop the image of the image to be detected based on the human body detection box to obtain a cropped image; perform hand recognition on the cropped image to obtain a hand detection box; perform mobile phone object recognition on the cropped image to obtain a mobile phone detection box; and determine the initial behavior recognition result of the target object based on the area intersection ratio between the hand detection box and the mobile phone detection box.
[0142] In an exemplary embodiment, the first recognition module 504 is configured to determine that the initial behavior recognition result is that there is mobile phone usage behavior when the area intersection ratio between the hand detection box and the mobile phone detection box is greater than a preset ratio threshold; and to determine that the initial behavior recognition result is that there is no mobile phone usage behavior when the area intersection ratio between the hand detection box and the mobile phone detection box is not greater than the preset ratio threshold.
[0143] In an exemplary embodiment, the perceptual domain determination module 503 is used to determine a first boundary for each target object in the image to be detected based on the maximum ordinate values of the left elbow keypoint and the right elbow keypoint in the human body keypoint set; determine a second boundary of the perceptual domain based on the ordinate of the nose keypoint in the human body keypoint set; and determine the perceptual domain of the target object in the image to be detected based on the boundary of the human body detection box, the first boundary, and the second boundary.
[0144] In an exemplary embodiment, the second recognition module 505 is configured to determine the target behavior recognition result of the target object as making a phone call when the initial behavior recognition result indicates that the first positional relationship meets a preset rule and the ear key point is located in the mobile phone detection frame of the mobile phone object; the preset rule is that the positional intersection between the mobile phone object and the target object's hand is greater than a preset ratio threshold; and to determine the target behavior recognition result of the target object as mobile phone entertainment behavior when the initial behavior recognition result indicates that the positional relationship between the target object's hand and the mobile phone object meets the preset rule and the ear key point is not located in the mobile phone detection frame of the mobile phone object.
[0145] In an exemplary embodiment, the detection module 502 is used to perform human body recognition based on the image to be detected, and obtain human body detection boxes for each target object in the image to be detected; perform human body key point detection on the image to be detected based on a preset key point type, and obtain an initial set of human body key points for each target object in the image to be detected; for each target object, if there are missing key points in the initial set of human body key points, perform key point prediction based on the initial set of human body key points and the image to be detected, and obtain a set of human body key points for the target object.
[0146] In an exemplary embodiment, the target behavior recognition device 500 based on the adaptive perception domain further includes a reporting module, which is used to obtain the historical target behavior recognition results corresponding to each of the multiple historical images that have temporal correlation; perform data statistics based on the historical target behavior recognition results corresponding to each of the multiple historical images to determine the number of times the same target object has performed a target behavior; and generate an anomaly report when the number of times the same target object has performed a target behavior is greater than a threshold.
[0147] Each module in the aforementioned target behavior recognition device based on adaptive perceptual domain can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0148] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data related to target behavior recognition. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a target behavior recognition method based on an adaptive perceptual domain.
[0149] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0150] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described above.
[0151] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0152] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method described above.
[0153] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0154] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0155] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0156] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A target behavior recognition method based on adaptive perceptual domain, characterized in that, The method includes: Acquire the image to be detected; Target detection is performed based on the image to be detected to determine the human body detection box and the set of human body key points for each target object in the image to be detected. For each target object in the image to be detected, the perceptual domain of the target object in the image to be detected is determined based on the human body detection box and the human body key point set; the perceptual domain of the target object is used to describe the potential area where the target behavior of the target object may occur. When the hand key point in the human body key point set is located within the perceptual domain and a mobile phone object is identified from the image region corresponding to the human body detection box, the initial behavior recognition result is determined based on the first positional relationship between the mobile phone object and the hand of the target object. If the initial behavior recognition result indicates that the first positional relationship satisfies the preset rule, the target behavior recognition result of the target object is determined based on the second positional relationship between the mobile phone detection frame of the mobile phone object and the ear key point in the human body key point set.
2. The method according to claim 1, characterized in that, When the hand key point in the human body key point set is located within the perceptual domain and a mobile phone object is identified from the image region corresponding to the human body detection box, an initial behavior recognition result is determined based on the first positional relationship between the mobile phone object and the hand of the target object, including: When the hand key points in the human body key point set are located within the perceptual domain, the image to be detected is cropped based on the human body detection box to obtain a cropped image. Hand recognition is performed on the cropped image to obtain a hand detection box; The cropped image is used to perform mobile phone object recognition to obtain a mobile phone detection box; The initial behavior recognition result of the target object is determined based on the area intersection ratio between the hand detection frame and the mobile phone detection frame.
3. The method according to claim 2, characterized in that, The step of determining the initial behavior recognition result of the target object based on the area intersection ratio between the hand detection frame and the mobile phone detection frame includes at least one of the following: If the area overlap ratio between the hand detection frame and the mobile phone detection frame is greater than a preset ratio threshold, the initial behavior recognition result is determined to be that there is mobile phone usage behavior. If the area overlap ratio between the hand detection frame and the mobile phone detection frame is not greater than a preset ratio threshold, the initial behavior recognition result is determined to be that there is no mobile phone usage behavior.
4. The method according to claim 1, characterized in that, For each target object in the image to be detected, determining the receptive domain of the target object in the image to be detected based on the human body detection box and the set of human body key points includes: For each target object in the image to be detected, a first boundary is determined based on the maximum ordinate of the left elbow key point and the right elbow key point in the human body key point set; The second boundary of the perceptual domain is determined based on the ordinate of the nose key point in the human body key point set. Based on the boundaries of the human detection box, the first boundary, and the second boundary, the perceptual domain of the target object in the image to be detected is determined.
5. The method according to claim 1, characterized in that, When the initial behavior recognition result indicates that the first positional relationship satisfies a preset rule, the target behavior recognition result of the target object is determined based on the second positional relationship between the mobile phone detection frame of the mobile phone object and the ear key point in the human body key point set, including: If the initial behavior recognition result indicates that the first positional relationship satisfies a preset rule, and the ear key point is located in the mobile phone detection frame of the mobile phone object, the target behavior recognition result of the target object is determined to be making a phone call; the preset rule is that the positional intersection between the mobile phone object and the hand of the target object is greater than a preset ratio threshold. If the initial behavior recognition result indicates that the positional relationship between the target object's hand and the mobile phone object meets the preset rules, and the ear key point is not located in the mobile phone detection frame of the mobile phone object, then the target behavior recognition result of the target object is determined to be mobile phone entertainment behavior.
6. The method according to claim 1, characterized in that, The step of performing target detection based on the image to be detected, and determining the human body detection bounding box and the set of human body key points for each target object in the image to be detected, includes: Human body recognition is performed on the image to be detected to obtain human body detection boxes for each target object in the image to be detected. Human key point detection is performed on the image to be detected based on the preset key point type to obtain the initial set of human key points of each target object in the image to be detected. For each target object, if key points are missing in the initial human key point set, key point prediction is performed based on the initial human key point set and the image to be detected to obtain the human key point set of the target object.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Obtain the historical target behavior recognition results corresponding to each of multiple historical images that have temporal correlation; Based on the historical target behavior recognition results corresponding to each of the multiple historical images, data statistics are performed to determine the number of times the same target object has performed a target behavior; An anomaly report is generated if the number of times the target object performs the target behavior exceeds a threshold.
8. A target behavior recognition device based on an adaptive perceptual domain, characterized in that, The device includes: The acquisition module is used to acquire the image to be detected; The detection module is used to perform target detection based on the image to be detected, and to determine the human body detection box and the set of human body key points for each target object in the image to be detected; The perceptual domain determination module is used to determine the perceptual domain of each target object in the image to be detected, based on the human body detection box and the human body key point set; the perceptual domain of the target object is used to describe the potential occurrence area of the target object's target behavior; The first recognition module is used to determine an initial behavior recognition result based on a first positional relationship between the mobile phone object and the hand of the target object when the hand key point in the human body key point set is located within the perception domain and a mobile phone object is recognized from the image area corresponding to the human body detection box. The second recognition module is used to determine the target behavior recognition result of the target object based on the second positional relationship between the mobile phone detection frame of the mobile phone object and the ear key point in the human body key point set, when the initial behavior recognition result indicates that the first positional relationship meets the preset rules.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.