Method and Device for Identifying Mobile Phone Usage Behavior

By performing dual behavior recognition on the area of interest images of the target person, the problem of low accuracy in mobile phone recognition in the prior art is solved, and accurate recognition and reminder in key scenarios is achieved, and user security is improved.

CN115147818BActive Publication Date: 2025-07-22BOE TECHNOLOGY GROUP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210764212.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-07-22
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

In the prior art, the accuracy of mobile phone behavior recognition is low, and it is impossible to identify and remind users in time to play mobile phone behavior in certain scenarios, resulting in potential adverse effects, such as the increase in car accidents.

Method used

By obtaining the image to be identified, the image of the area of interest of the target person is extracted, and the first behavior recognition model and the second behavior recognition model are used for double recognition. If the results are inconsistent, the behavior recognition of the target person is further processed to improve the accuracy.

Benefits of technology

It improves the accuracy of mobile phone recognition, can issue reminders in a timely manner to avoid adverse effects, and is suitable for scenarios such as vehicle driving, guard posts, office areas and classrooms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147818B_ABST
    Figure CN115147818B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method and apparatus for identifying mobile phone usage behaviors, relating to the field of computer vision technology, and are used to improve the accuracy of identifying mobile phone usage behaviors. The method includes: obtaining an image to be identified, and extracting a region-of-interest image containing a target person from the image to be identified. Inputting the region-of-interest image into a first behavior recognition model to obtain a first behavior recognition result of the target person, where the first behavior recognition result is used to indicate whether the target person has a mobile phone usage behavior. Inputting the region-of-interest image into a second behavior recognition model to obtain a second behavior recognition result of the target person, where the second behavior recognition result is used to indicate whether the target person has a mobile phone usage behavior. Comparing the first behavior recognition result and the second behavior recognition result, if the first behavior result is inconsistent with the second behavior result, then based on the region-of-interest image, perform behavior recognition processing on the target person to determine whether the target person has a mobile phone usage behavior.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular, to a method and device for identifying mobile phone playing behaviors. Background Art

[0002] With the popularization of mobile phones, mobile phones have become an indispensable part of people's daily lives. While mobile phones bring convenience to people's lives, people's dependence on mobile phones has become increasingly serious. Playing with mobile phones in certain scenarios is likely to have a certain impact on people's lives. For example, in a driving scenario, if a driver plays with a mobile phone while driving, the probability of a car accident may increase. Therefore, in certain scenarios, it is necessary to accurately identify whether people are playing with mobile phones to issue real-time warnings. However, in related technologies, the accuracy of identifying mobile phone playing behaviors is relatively low. Summary of the Invention

[0003] Embodiments of the present disclosure provide a method and device for identifying mobile phone playing behaviors, which are used to improve the accuracy of identifying mobile phone playing behaviors.

[0004] On the one hand, a method for identifying mobile phone playing behaviors is provided. The method includes: obtaining an image to be recognized, and extracting a region of interest image containing a target person from the image to be recognized. Inputting the region of interest image into a first behavior recognition model to obtain a first behavior recognition result of the target person, where the first behavior recognition result is used to indicate whether the target person is playing with a mobile phone. Inputting the region of interest image into a second behavior recognition model to obtain a second behavior recognition result of the target person, where the second behavior recognition result is used to indicate whether the target person is playing with a mobile phone. Comparing the first behavior recognition result and the second behavior recognition result. If the first behavior result is inconsistent with the second behavior result, then based on the region of interest image, perform behavior recognition processing on the target person to determine whether the target person is playing with a mobile phone.

[0005] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: Based on the region of interest image containing the target person, the first behavior recognition model and the second behavior recognition model are used to perform dual recognition on whether the target person is playing with a mobile phone, which improves the accuracy of identifying mobile phone playing behaviors. And when the first behavior recognition result output by the first behavior recognition model is inconsistent with the second behavior recognition result of the second behavior recognition model, perform behavior recognition processing on the target person again based on the region of interest image containing the target person, so as to determine whether the target person is playing with a mobile phone. It can be seen that a method for identifying mobile phone playing behaviors provided by the embodiments of the present disclosure performs multiple behavior recognitions on whether the user is playing with a mobile phone, which improves the accuracy of identifying mobile phone playing behaviors. So as to issue a reminder message in time when it is recognized that the target person is playing with a mobile phone, and avoid the occurrence of adverse effects caused by the target person playing with a mobile phone.

[0006] In some embodiments, the above method further includes: if the recognition result of the first behavior is the same as that of the second behavior, determining whether the target person has the behavior of playing with a mobile phone based on the recognition result of the first behavior or the recognition result of the second behavior.

[0007] In other embodiments, the above-mentioned behavior recognition processing of the target person based on the region of interest image to determine whether the target person has the behavior of playing with a mobile phone includes: inputting the region of interest image into a mobile phone detection model and inputting the region of interest image into a person detection model; if no mobile phone is detected from the region of interest image, determining that the target person does not have the behavior of playing with a mobile phone; if a mobile phone is detected from the region of interest image, then determining whether the target person has the behavior of playing with a mobile phone according to the mobile phone frame output by the mobile phone detection model and the person frame output by the person detection model.

[0008] In other embodiments, when the person detection model outputs only one person frame, the above-mentioned determining whether the target person has the behavior of playing with a mobile phone according to the mobile phone frame output by the mobile phone detection model and the person frame output by the person detection model includes: determining the degree of overlap between the mobile phone frame and the person frame; if the degree of overlap is greater than or equal to a preset overlap threshold, determining that the target person has the behavior of playing with a mobile phone; if the degree of overlap is less than the preset overlap threshold, determining that the target person does not have the behavior of playing with a mobile phone.

[0009] In other embodiments, the above-mentioned determining the degree of overlap between the mobile phone frame and the person frame includes: determining the area of the overlapping region between the mobile phone frame and the person frame in the region of interest image; using the ratio between the area of the overlapping region and the area of the region occupied by the mobile phone frame in the region of interest as the degree of overlap.

[0010] In other embodiments, before the above-mentioned determining the degree of overlap between the mobile phone frame and the person frame, the above method further includes: determining the distance between the target person and the mobile phone based on the mobile phone frame and the person frame; when the distance between the target person and the mobile phone is greater than a preset distance threshold, determining that the target person does not have the behavior of playing with a mobile phone; the above-mentioned determining the degree of overlap between the mobile phone frame and the person frame includes: when the distance between the target person and the mobile phone is less than or equal to the preset distance threshold, determining the degree of overlap between the mobile phone frame and the person frame.

[0011] In some other embodiments, when the person detection model outputs multiple person bounding boxes, determining whether the target person has a behavior of playing with a mobile phone based on the mobile phone bounding box output by the mobile phone detection model and the person bounding boxes output by the person detection model includes: determining the distance between the target person and the mobile phone based on the person bounding box of the target person, the mobile phone bounding box, and the region of interest image; determining the distance between the non-target person and the mobile phone based on the person bounding box of the non-target person, the mobile phone bounding box, and the region of interest image; when the distance between the target person and the mobile phone is less than the distances between all non-target persons and the mobile phone, determining that the target person has a behavior of playing with a mobile phone; when the distance between the target person and the mobile phone is greater than or equal to the distance between the target person and any non-target person and the mobile phone, determining that the target person does not have a behavior of playing with a mobile phone.

[0012] In some other embodiments, determining the distance between the target person and the mobile phone based on the person bounding box of the target person, the mobile phone bounding box, and the region of interest image includes: performing hand recognition on the target person based on the person bounding box of the target person and the region of interest image to determine the central position of the hand of the target person; determining the central position of the mobile phone based on the mobile phone bounding box and the region of interest image; and determining the distance between the target person and the mobile phone according to the central position of the hand of the target person and the central position of the mobile phone.

[0013] In some other embodiments, the first behavior recognition model is an Inception network model, and the second behavior recognition model is a ResNet model.

[0014] In another aspect, a behavior recognition device is provided. The behavior recognition device includes: a communication unit configured to obtain an image to be recognized; a processing unit configured to: extract a region of interest image including a target person from the image to be recognized; input the region of interest image into a first behavior recognition model to obtain a first behavior recognition result of the target person, where the first behavior recognition result is used to indicate whether the target person has a behavior of playing with a mobile phone; input the region of interest image into a second behavior recognition model to obtain a second behavior recognition result of the target person, where the second behavior recognition result is used to indicate whether the target person has a behavior of playing with a mobile phone; and if the first behavior recognition result is inconsistent with the second behavior recognition result, perform behavior recognition processing on the target person based on the region of interest image to determine whether the target person has a behavior of playing with a mobile phone.

[0015] In some embodiments, the processing unit is further configured to, if the first behavior recognition result is consistent with the second behavior recognition result, determine whether the target person has a behavior of playing with a mobile phone based on the first behavior recognition result or the second behavior recognition result.

[0016] In some other embodiments, the above-mentioned processing unit is specifically configured to: input the image of the region of interest into a mobile phone detection model and a person detection model; if no mobile phone is detected in the image of the region of interest, determine that the target person is not engaged in the behavior of playing with the mobile phone; if a mobile phone is detected in the image of the region of interest, then determine whether the target person is engaged in the behavior of playing with the mobile phone according to the mobile phone frame output by the mobile phone detection model and the person frame output by the person detection model.

[0017] In some other embodiments, when the person detection model outputs only one person frame, the above-mentioned processing unit is specifically configured to: determine the degree of overlap between the mobile phone frame and the person frame; if the degree of overlap is greater than or equal to a preset overlap threshold, determine that the target person is engaged in the behavior of playing with the mobile phone; if the degree of overlap is less than the preset overlap threshold, determine that the target person is not engaged in the behavior of playing with the mobile phone.

[0018] In some other embodiments, the above-mentioned processing unit is specifically configured to: determine the area of the overlapping region between the mobile phone frame and the person frame in the image of the region of interest; use the ratio between the area of the overlapping region and the area of the region occupied by the mobile phone frame in the region of interest as the degree of overlap.

[0019] In some other embodiments, the above-mentioned processing unit is further configured to: determine the distance between the target person and the mobile phone based on the mobile phone frame and the person frame; when the distance between the target person and the mobile phone is greater than a preset distance threshold, determine that the target person is not engaged in the behavior of playing with the mobile phone; the above-mentioned processing unit is specifically configured to determine the degree of overlap between the mobile phone frame and the person frame when the distance between the target person and the mobile phone is less than or equal to the preset distance threshold.

[0020] In some other embodiments, when the person detection model outputs multiple person frames, the above-mentioned processing unit is specifically configured to: determine the person frame of the target person and the person frames of non-target persons from the multiple person frames, where non-target persons are other persons in the image of the region of interest except the target person; determine the distance between the target person and the mobile phone based on the person frame of the target person, the mobile phone frame, and the image of the region of interest; determine the distance between the non-target persons and the mobile phone based on the person frames of the non-target persons, the mobile phone frame, and the image of the region of interest; when the distance between the target person and the mobile phone is less than the distances between all non-target persons and the mobile phone, determine that the target person is engaged in the behavior of playing with the mobile phone; when the distance between the target person and the mobile phone is greater than or equal to the distance between the target person and any one of the non-target persons and the mobile phone, determine that the target person is not engaged in the behavior of playing with the mobile phone.

[0021] In some other embodiments, the above-mentioned processing unit is specifically configured to perform hand recognition on the target person based on the person frame of the target person and the image of the region of interest, to determine the central position of the hand of the target person; determine the central position of the mobile phone based on the mobile phone frame and the image of the region of interest; and determine the distance between the target person and the mobile phone according to the central position of the hand of the target person and the central position of the mobile phone.

[0022] In some other embodiments, the above-mentioned first behavior recognition model is an Inception network model, and the above-mentioned second behavior recognition model is a ResNet model.

[0023] On the other hand, a behavior recognition device is provided. The behavior recognition device includes a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions. Wherein, when the processor executes the computer instructions, the behavior recognition device is caused to execute the mobile phone playing behavior recognition method as described in any of the above embodiments.

[0024] On the other hand, a non-transitory computer-readable storage medium is provided. The computer-readable storage medium stores computer program instructions, and when the computer program instructions run on a processor, the processor is caused to execute one or more steps in the mobile phone playing behavior recognition method as described in any of the above embodiments.

[0025] On the other hand, a computer program product is provided. The computer program product includes computer program instructions, and when the computer program instructions are executed on a computer, the computer program instructions cause the computer to execute one or more steps in the mobile phone playing behavior recognition method as described in any of the above embodiments.

[0026] On the other hand, a computer program is provided. When the computer program is executed on a computer, the computer program causes the computer to execute one or more steps in the mobile phone playing behavior recognition method as described in any of the above embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the present disclosure, the following will briefly introduce the drawings required to be used in some embodiments of the present disclosure. Obviously, the drawings in the following description are only the drawings of some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings. In addition, the drawings in the following description can be regarded as schematic diagrams, and are not limitations on the actual sizes of the products, the actual processes of the methods, the actual timings of the signals, etc. involved in the embodiments of the present disclosure.

[0028] Figure 1 FIG. is a composition diagram of a mobile phone playing behavior recognition system according to some embodiments;

[0029] Figure 2 Hardware structure diagram of a behavior recognition device according to some embodiments;

[0030] Figure 3 Flowchart of a method for recognizing mobile phone usage behavior according to some embodiments Figure 1 ;

[0031] Figure 4 Architecture of the inception structure according to some embodiments Figure 1 ;

[0032] Figure 5 Architecture of the inception structure according to some embodiments Figure 2 ;

[0033] Figure 6 Architecture of the resnet18 model according to some embodiments Figure 1 ;

[0034] Figure 7 Architecture of the resnet18 model according to some embodiments Figure 2 ;

[0035] Figure 8 Flowchart of a method for recognizing mobile phone usage behavior according to some embodiments Figure 2 ;

[0036] Figure 9 Flowchart of a method for recognizing mobile phone usage behavior according to some embodiments Figure 3 ;

[0037] Figure 10 Flowchart of a method for recognizing mobile phone usage behavior according to some embodiments Figure 4 ;

[0038] Figure 11 Schematic diagram of an image of an area of interest according to some embodiments;

[0039] Figure 12 Flowchart of a method for recognizing mobile phone usage behavior according to some embodiments Figure 5 ;

[0040] Figure 13 Flowchart of a method for recognizing mobile phone usage behavior according to some embodiments Figure 6 ;

[0041] Figure 14 Flowchart of a process for recognizing mobile phone usage behavior according to some embodiments;

[0042] Figure 15 Structure diagram of a behavior recognition device according to some embodiments. Detailed implementation manners

[0043] The following will describe the technical solutions in some embodiments of the present disclosure clearly and completely with reference to the accompanying drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present disclosure. Based on the embodiments provided by the present disclosure, all other embodiments obtained by those of ordinary skill in the art fall within the protection scope of the present disclosure.

[0044] Unless otherwise required by the context, throughout the specification and claims, the term "comprise" and its other forms, such as the third-person singular form "comprises" and the present participle form "comprising", are interpreted as open and inclusive, that is, "including, but not limited to". In the description of the specification, the terms "one embodiment", "some embodiments", "exemplary embodiments", "example", "specific example" or "some examples", etc. are intended to indicate that the specific features, structures, materials or characteristics related to the embodiment or example are included in at least one embodiment or example of the present disclosure. The schematic representations of the above terms do not necessarily refer to the same embodiment or example. In addition, the described specific features, structures, materials or characteristics may be included in any one or more embodiments or examples in any appropriate manner.

[0045] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, unless otherwise specified, the meaning of "a plurality" is two or more.

[0046] "At least one of A, B, and C" has the same meaning as "at least one of A, B, or C", and both include the following combinations of A, B, and C: only A, only B, only C, the combination of A and B, the combination of A and C, the combination of B and C, and the combination of A, B, and C.

[0047] "A and / or B" includes the following three combinations: only A, only B, and the combination of A and B.

[0048] As used herein, depending on the context, the term "if" is optionally interpreted to mean "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if it is determined that..." or "if [the stated condition or event] is detected" is optionally interpreted to mean "when it is determined that..." or "in response to determining that..." or "when [the stated condition or event] is detected" or "in response to detecting [the stated condition or event]".

[0049] As used herein, the use of "configured to" or "adapted to" means open and inclusive language that does not exclude devices that are adapted to or configured to perform additional tasks or steps.

[0050] Additionally, the use of "based on" is open and inclusive because a process, step, calculation, or other action "based on" one or more of the stated conditions or values can, in practice, be based on additional conditions or values beyond those stated.

[0051] As used herein, "about", "substantially", or "approximately" includes the stated value and an average within an acceptable deviation range of the specific value, where the acceptable deviation range is determined by a person of ordinary skill in the art considering the measurement being discussed and the errors associated with the measurement of a particular quantity (i.e., the limitations of the measurement system).

[0052] With the increasing intelligence of mobile phones, people's dependence on mobile phones in aspects such as clothing, food, housing, and transportation is getting higher and higher. To avoid the adverse effects caused by people playing with mobile phones in certain scenarios, it is necessary to identify people's mobile phone playing behaviors in a timely manner to remind people not to play with mobile phones in certain scenarios and avoid the occurrence of adverse effects. Taking the vehicle driving scenario as an example, the adverse effects caused by people's mobile phone playing behaviors can be that the probability of a car accident increases when the vehicle driver plays with a mobile phone during the driving process.

[0053] In the method for identifying mobile phone playing behaviors provided by the related art, for the identification of people's mobile phone playing behaviors, the image of the environment where people are located is input as a whole into the behavior recognition model for mobile phone playing behavior recognition. Since the image of the environment where people are located contains a large amount of redundant data and is only recognized once by the behavior recognition model, the accuracy of mobile phone playing behavior recognition is low, and it is impossible to timely identify whether people have mobile phone playing behaviors, and thus it is impossible to give a timely reminder when people have mobile phone playing behaviors.

[0054] Based on this, embodiments of the present disclosure provide a method for identifying the behavior of playing with a mobile phone. This method obtains an image to be identified, extracts a region of interest image containing the target person from the image, and determines whether the target person has the behavior of playing with a mobile phone based on the region of interest image containing the target person, rather than determining whether the target person has the behavior of playing with a mobile phone based on the image to be identified that contains a large amount of redundant data. This reduces the interference of the redundant data in the image to be identified on the identification of the behavior of playing with a mobile phone and improves the accuracy of the identification of the behavior of playing with a mobile phone.

[0055] In addition, regarding the problem of relatively low accuracy of identifying the behavior of playing with a mobile phone caused by performing a single identification of the behavior of playing with a mobile phone through a behavior recognition model in the method for identifying the behavior of playing with a mobile phone provided in the related art, the method for identifying the behavior of playing with a mobile phone provided in the embodiments of the present disclosure first identifies whether the target person has the behavior of playing with a mobile phone through a first behavior recognition model and a second behavior recognition model respectively. When the first behavior recognition result of the first behavior recognition model is consistent with the second behavior recognition result of the second behavior recognition model, the first behavior recognition result or the second behavior recognition result is used as the result of whether the target person has the behavior of playing with a mobile phone. Thus, a double identification of whether the target person has the behavior of playing with a mobile phone is performed, improving the accuracy of the identification of the behavior of playing with a mobile phone. And when the first behavior recognition result of the first behavior recognition model is inconsistent with the second behavior recognition result of the second behavior recognition model, the behavior of the target person is recognized and processed again based on the region of interest image containing the target person in the image to be identified, so as to determine whether the target person has the behavior of playing with a mobile phone. It can be seen that the method for identifying the behavior of playing with a mobile phone provided in the embodiments of the present disclosure performs multiple identifications of the behavior of playing with a mobile phone on the region of interest image containing the target person, and the accuracy of the recognition result of the user's behavior of playing with a mobile phone is higher, improving the accuracy of the identification of the behavior of playing with a mobile phone. Furthermore, when the target person has the behavior of playing with a mobile phone, a reminder can be sent in a timely manner to avoid the occurrence of adverse effects caused by the target person having the behavior of playing with a mobile phone.

[0056] The method for identifying the behavior of playing with a mobile phone provided in the embodiments of the present disclosure can be applied to scenarios such as vehicle driving, sentry duty in a sentry box, office areas, and classrooms.

[0057] Taking the application of the method for identifying the behavior of playing with a mobile phone in the vehicle driving scenario as an example, after the behavior recognition device determines that the vehicle driver has the behavior of playing with a mobile phone based on the method for identifying the behavior of playing with a mobile phone provided in the embodiments of the present disclosure, the behavior recognition device can upload the image in the vehicle terminal and the recognition result of the behavior of playing with a mobile phone to the background management server of the vehicle terminal for the management personnel to view. Further, after the behavior recognition device determines that the vehicle driver has the behavior of playing with a mobile phone for a period of time, the behavior recognition device can control the vehicle terminal to send an alarm message to prompt the vehicle driver to prohibit playing with the mobile phone and pay attention to driving safety.

[0058] Taking the application of the mobile phone playing behavior recognition method in the classroom scenario as an example, after the behavior recognition device determines that there are students playing with mobile phones in the classroom based on the mobile phone playing behavior recognition method provided by the embodiments of the present disclosure, the behavior recognition device can upload the image in the classroom at this time and the mobile phone playing behavior recognition result to the teacher's terminal device for the teacher to view, so that the teacher can maintain the classroom teaching environment according to the mobile phone playing behavior recognition result displayed on the terminal device.

[0059] As Figure 1 shown, the embodiments of the present disclosure provide a composition diagram of a mobile phone playing behavior recognition system. The mobile phone playing behavior recognition system includes: a behavior recognition device 10 and a shooting device 20. Among them, the behavior recognition device 10 and the shooting device 20 can be connected by wired or wireless means.

[0060] The shooting device 20 can be arranged near the supervision area. For example, taking the supervision area as the vehicle cab, the shooting device 20 can be installed on the top of the vehicle cab. The embodiments of the present disclosure do not limit the specific installation method and specific installation position of the shooting device 20.

[0061] The shooting device 20 can be used to shoot the image to be recognized in the supervision area.

[0062] In some embodiments, the shooting device 20 can use a color camera to shoot color images.

[0063] Exemplarily, the color camera can be an RGB camera. Among them, the RGB camera adopts the RGB color model, and obtains various colors through the changes of the three color channels of red (R), green (G), and blue (B) and their superposition with each other. Usually, the RGB camera gives three basic color components by three different cables, and three independent charge coupled device (CCD) sensors are used to obtain three color signals.

[0064] In some embodiments, the shooting device can use a depth camera to shoot depth images.

[0065] Exemplarily, the depth camera can be a time of flight (TOF) camera. The TOF camera adopts TOF technology, and the imaging principle of the TOF camera is as follows: According to the pulsed infrared light emitted by the laser light source, after encountering an object, it is reflected, and the light source detector receives the light source reflected by the object. By calculating the time difference or phase difference between the emission and reflection of the light source, the distance between the TOF camera and the object to be photographed is converted, and then according to the distance between the TOF camera and the object to be photographed, the depth values of each point in the scene are obtained.

[0066] The behavior recognition device 10 is configured to obtain the image to be recognized captured by the imaging device 20, and determine whether there is a behavior of playing with a mobile phone by the person in the supervised area based on the image to be recognized captured by the imaging device 20.

[0067] In some embodiments, the behavior recognition device 10 may be an independent server, or a server cluster or a distributed system composed of multiple servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data service networks.

[0068] In some embodiments, the behavior recognition device 10 may be a mobile phone, a tablet computer, a desktop type, a laptop, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, etc. Alternatively, the behavior recognition device 10 may be a vehicle terminal. The vehicle terminal is a front-end device for vehicle communication and management and can be installed in various vehicles.

[0069] In some embodiments, the behavior recognition device 10 may communicate with other terminal devices in a wired or wireless manner. For example, it communicates with the terminal device of the vehicle administrator in a vehicle driving scenario, or communicates with the terminal device of the teacher in a classroom scenario.

[0070] Exemplarily, in a classroom scenario, after the behavior recognition device 10 determines the result of the recognition of the behavior of playing with a mobile phone in the classroom based on the image to be recognized captured by the imaging device 20, the result of the recognition of the behavior of playing with a mobile phone may be sent to the terminal device of the teacher in the form of voice, text, or video for the teacher to view.

[0071] In some embodiments, the behavior recognition device 10 may be integrated with the imaging device 20.

[0072] Figure 2 This is a hardware structure diagram of a behavior recognition device provided by an embodiment of the present disclosure. Refer to Figure 2 The behavior recognition device may include a processor 41, a memory 42, a communication interface 43, and a bus 44. The processor 41, the memory 42, and the communication interface 43 may be connected through the bus 44.

[0073] The processor 41 is the control center of the behavior recognition device, which can be a single processor or a collective term for multiple processing elements. For example, the processor 41 can be a general-purpose CPU or other general-purpose processors. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0074] As an embodiment, the processor 41 may include one or more CPUs, such as Figure 2 the CPUs 0 and 1 shown in

[0075] The memory 42 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0076] In a possible implementation, the memory 42 can exist independently of the processor 41. The memory 42 can be connected to the processor 41 through the bus 44 and is used to store instructions or program codes. When the processor 41 calls and executes the instructions or program codes stored in the memory 42, the method for recognizing the behavior of playing with a mobile phone provided in the following embodiments of the present disclosure can be implemented.

[0077] In another possible implementation, the memory 42 can also be integrated with the processor 41.

[0078] The communication interface 43 is used for the behavior recognition device to be connected to other devices through a communication network. The communication network can be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 43 can include a receiving unit for receiving data and a sending unit for sending data.

[0079] The bus 44 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 2 it is only represented by a thick line in the figure, but it does not mean that there is only one bus or one type of bus.

[0080] It should be noted that Figure 2 the structure shown in the figure does not constitute a limitation on the behavior recognition device. Except Figure 2 for the components shown, the behavior recognition device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0081] The following specifically introduces the embodiments provided by the present disclosure in conjunction with the accompanying drawings of the specification.

[0082] The method for recognizing the behavior of playing with a mobile phone provided by the embodiments of the present disclosure is applied to a behavior recognition device. The behavior recognition device may be the behavior recognition device 10 in the above-mentioned behavior recognition system for playing with a mobile phone, or the processor of the behavior recognition device 10. As Figure 3 shown, the method includes the following steps:

[0083] S101. Obtain an image to be recognized.

[0084] Among them, the image to be recognized is an image obtained by a photographing device photographing a supervised area. The supervised area is an area where it is necessary to supervise whether the user has the behavior of playing with a mobile phone. For example, a vehicle cab, a classroom, an office area, a sentry box, etc.

[0085] In some embodiments, the supervised area can be determined by the behavior recognition device. For example, there are multiple photographing devices connected to the behavior recognition device, and the behavior recognition device can consider the area where each photographing device is located as the supervised area.

[0086] In some embodiments, the supervised area can be determined by the user in a direct or indirect manner. For example, in an application scenario of a classroom, a school has M classrooms, and a corresponding photographing device is installed in each classroom. When there are no students in N classrooms out of the M classrooms, the user can choose to turn off the photographing devices in the N classrooms. Then, the behavior recognition device can select each of the M - N classrooms as the supervised area. In this way, the behavior recognition device does not need to perform the recognition of mobile phone playing behaviors in the N classrooms, so as to save computing resources. Here, both M and N are positive integers.

[0087] The image to be recognized is used to record the images of K persons included in the supervised area at the current moment. Here, K is a positive integer.

[0088] In some embodiments, after the behavior recognition device turns on the mobile phone playing behavior recognition function, it executes the mobile phone playing behavior recognition method provided in the embodiments of the present disclosure. Correspondingly, if the behavior recognition device turns off the mobile phone playing behavior recognition function, the behavior recognition device does not execute or stops executing the mobile phone playing behavior recognition method provided in the embodiments of the present disclosure.

[0089] In an optional implementation manner, the behavior recognition device defaults to turn on the mobile phone playing behavior recognition function.

[0090] In another optional implementation manner, the behavior recognition device periodically turns on the mobile phone playing behavior recognition function. For example, in a classroom scenario, the behavior recognition device automatically turns on the mobile phone playing behavior recognition function between 8:00 am and 5:30 pm, and automatically turns off the mobile phone playing behavior recognition function between 5:30 pm and 8:00 am.

[0091] In another optional implementation manner, the behavior recognition device determines whether to turn on / off the mobile phone playing behavior recognition function according to the instruction of the terminal device.

[0092] For example, in an application scenario of vehicle driving, during the process of a driver driving a vehicle, a vehicle manager issues an instruction to turn on the mobile phone playing behavior recognition function to the behavior recognition device through the terminal device. In response to this instruction, the behavior recognition device turns on the mobile phone playing behavior recognition function. Or, after the driver stops driving the vehicle, the vehicle manager issues an instruction to turn off the mobile phone playing behavior recognition function to the behavior recognition device through the terminal device. In response to this instruction, the behavior recognition device turns off the mobile phone playing behavior recognition function.

[0093] In some embodiments, when a preset condition is satisfied, the behavior recognition device obtains the image to be recognized of the supervised area through the photographing device.

[0094] Optionally, when applied to the vehicle driving scenario, the above preset conditions include: the shooting device detects that there is a person in the vehicle cab. In this way, the behavior recognition device only needs to recognize the behavior of playing with the mobile phone when there is a person in the vehicle cab, and does not need to recognize the behavior of playing with the mobile phone when there is no person in the vehicle cab, which helps to reduce the computing amount of the behavior recognition device.

[0095] In some embodiments, the behavior recognition device obtains the image to be recognized in the supervised area through the shooting device, which can be specifically implemented as: the behavior recognition device sends a shooting instruction to the shooting device, and the shooting instruction is used to instruct the shooting device to shoot the image of the supervised area; after that, the behavior recognition device receives the image to be recognized in the supervised area from the shooting device.

[0096] Optionally, the image to be recognized can be the one shot by the shooting device before receiving the shooting instruction, or the one shot by the shooting device after receiving the shooting instruction.

[0097] S102. Extract the region of interest image containing the target person from the image to be recognized.

[0098] In some embodiments, after the behavior recognition device receives the image to be recognized sent by the shooting device, it can perform human body recognition processing on the image to be recognized, and then determine the target person from the K persons in the image to be recognized. Among them, the target person can be any one of the K persons, or a specific person among the K persons, and K is a positive integer.

[0099] It can be understood that in some scenarios, the behavior recognition device only needs to recognize the behavior of playing with the mobile phone for a specific person in the supervised area, and does not need to recognize the behavior of playing with the mobile phone for each person in the supervised area. For example, in the vehicle driving scenario, the behavior recognition device only needs to recognize the behavior of playing with the mobile phone for the vehicle driver, and does not need to recognize the behavior of playing with the mobile phone for other passengers in the vehicle, which can reduce the computing amount of the behavior recognition device.

[0100] In some embodiments, after the behavior recognition device receives the image to be recognized, the behavior recognition device can perform identity recognition processing on the image to be recognized to recognize the identity of each person among the K persons included in the supervised area, and then send the identity recognition results of the K persons to the terminal device for the user of the terminal device to view. If the user of the terminal device selects to recognize the behavior of playing with the mobile phone for a certain person among the K persons according to the identity recognition results of the K persons, the behavior recognition device determines that this person is the target person. If the user of the terminal device selects to recognize the behavior of playing with the mobile phone for the K persons according to the identity recognition results of the K persons, the behavior recognition device determines any one of the K persons as the target person.

[0101] Optionally, the behavior recognition device performs identity recognition processing on the image to be recognized to identify the identity of each of the K persons included in the supervised area. This can be specifically implemented as: inputting the image to be recognized into an identity recognition model to obtain the identity recognition result of each person.

[0102] In some embodiments, a trained identity recognition model is pre-stored in the memory of the behavior recognition device. After obtaining the image to be recognized, the behavior recognition device can input the image to be recognized into the identity recognition model to obtain the identity recognition result of each of the K persons included in the supervised area.

[0103] In some embodiments, the above identity recognition model can be a convolutional neural network model (CNN). For example, the model structure of VGG-16 can be used to implement it.

[0104] In some embodiments, after the behavior recognition device determines the target person, in order to remove the influence of redundant information in the image to be recognized on the accuracy of subsequent mobile phone usage behavior recognition, the behavior recognition device can perform image segmentation processing on the image to be recognized, and then extract the region of interest image containing the target person from the image to be recognized.

[0105] It can be understood that after the image to be recognized is segmented, the target person is presented in the form of a detection box in the image to be recognized. The image of the region formed by equally enlarging and expanding the detection box of the target person in the image to be recognized is used as the region of interest image containing the target person.

[0106] Optionally, extracting the region of interest image containing the target person from the image to be recognized can be specifically implemented as: inputting the image to be recognized into an image segmentation model to obtain the region of interest image corresponding to each person.

[0107] In some embodiments, a trained image segmentation model is pre-stored in the memory of the behavior recognition device. After obtaining the image to be recognized, the behavior recognition device can input the image to be recognized into the trained image segmentation model to obtain the region of interest image corresponding to each of the K persons included in the supervised area.

[0108] In some embodiments, the above image segmentation model can be a deep neural network (DNN) model.

[0109] It is easy to understand that deep neural networks can automatically extract and learn more essential features in images from a vast amount of training data. Applying deep neural networks to image segmentation will significantly enhance the classification effect and further improve the accuracy of subsequent recognition of mobile phone usage behaviors.

[0110] In some embodiments, the above-mentioned image segmentation model can be constructed based on the Deeplab v3+ semantic segmentation algorithm.

[0111] Optionally, the image of the region of interest of the target person can be an image of the region of interest after restoration processing to ensure the accuracy of the subsequent recognition result of the mobile phone usage behavior of the target person based on the image of the region of interest of the target person.

[0112] S103: Input the image of the region of interest into the first behavior recognition model to obtain the first behavior recognition result of the target person.

[0113] In some embodiments, the trained first behavior recognition model is pre-stored in the memory of the behavior recognition device. To recognize whether the target person has the behavior of using a mobile phone, after obtaining the image of the region of interest of the target person, the image of the region of interest of the target person can be input into the first behavior recognition model to obtain the first behavior recognition result of the target person. Among them, the first behavior recognition result is used to indicate whether the target person has the behavior of using a mobile phone.

[0114] Among them, the behavior of using a mobile phone includes the target person holding the mobile phone in hand to send text messages or voice messages, placing the mobile phone on an object such as a table to send text messages or voice messages, and putting the mobile phone against the ear to make a call or listen to voice messages, etc.

[0115] Optionally, the above-mentioned first behavior recognition model is an inception network model, for example, it can be an inception-v3 model. Among them, the inception structure in the inception-v3 model combines different convolutional layers in a parallel manner. It should be understood that the first behavior recognition mode can adopt the inception structure in related technologies (such as Figure 4 the shown inception structure), or it can adopt the improved inception structure provided in the embodiments of the present application (such as Figure 5 the shown inception structure).

[0116] Figure 4 Shows a schematic diagram of an inception structure. As shown in Figure 4As shown, in the related art, the Inception structure includes an output layer, a fully connected layer, and four learning paths located between the output layer and the fully connected layer. The first learning path includes a 1*1 convolutional kernel, a 3*3 convolutional kernel, and a 3*3 convolutional kernel connected in sequence. The second learning path includes a 1*1 convolutional kernel and a 3*3 convolutional kernel connected in sequence. The third learning path includes a pool and a 1*1 convolutional kernel. The fourth learning path includes a 1*1 convolutional kernel.

[0117] In some embodiments, in order to accelerate the training and convergence speed of the Inception-v3 model, the improved Inception structure provided by the embodiments of the present application uses 1*7 convolutional kernels and 7*1 convolutional kernels to replace the originally used 3*3 convolutional kernels.

[0118] Exemplarily, as Figure 5 shown, the embodiments of the present disclosure provide a schematic diagram of an improved Inception structure. The improved Inception structure includes an output layer, a fully connected layer, and ten learning paths located between the output layer and the fully connected layer. The first learning path includes a 1*7 convolutional kernel, a 7*7 convolutional kernel, and a 1*7 convolutional kernel connected in sequence. The second learning path includes a 7*1 convolutional kernel, a 7*7 convolutional kernel, and a 7*1 convolutional kernel connected in sequence. The third learning path includes a 1*1 convolutional kernel and a 1*7 convolutional kernel connected in sequence. The fourth learning path includes a 1*1 convolutional kernel and a 7*1 convolutional kernel connected in sequence. The fifth learning path includes a Pool and a 1*7 convolutional kernel connected in sequence. The sixth learning path includes a Pool and a 7*1 convolutional kernel connected in sequence. The seventh learning path includes a 1*7 convolutional kernel. The eighth learning path includes a 7*1 convolutional kernel. The ninth learning path includes a Pool and a 1*7 convolutional kernel connected in sequence. The tenth learning path includes a Pool and a 7*1 convolutional kernel connected in sequence.

[0119] Exemplarily, if the recognition result of the first row is "yes", it means that the behavior recognition result of the first row recognition model for the target person based on the region of interest image of the target person is that the target person has the behavior of playing with the mobile phone; if the recognition result of the first row is "no", it means that the behavior recognition result of the first row recognition model for the target person based on the region of interest image of the target person is that the target person does not have the behavior of playing with the mobile phone.

[0120] S104. Input the region of interest image into the second behavior recognition model to obtain the second behavior recognition result of the target person.

[0121] In some embodiments, a trained second behavior recognition model is pre-stored in the memory of the behavior recognition device. In order to recognize whether the target person has the behavior of playing with a mobile phone, after obtaining the image of the region of interest of the target person, the image of the region of interest of the target person can be input into the second behavior recognition model to obtain the second behavior recognition result of the target person. The second behavior recognition result is used to indicate whether the target user has the behavior of playing with a mobile phone.

[0122] Optionally, the second behavior recognition model is a residual network model, such as the resnet18 model. The resnet18 model is a serial network structure based on basicblock, which cleverly uses shortcut connections to solve the problem of model degradation in deep networks. It should be understood that the above second behavior recognition model can adopt the resnet18 model in related technologies (such as Figure 6 the resnet18 model shown), or can adopt the improved resnet18 model provided by the embodiments of the present application (such as Figure 7 the resnet18 model shown).

[0123] Figure 6 shows an architecture diagram of a resnet18 model. As Figure 6 shown, the resnet18 model in related technologies includes an output layer, a 7*7 convolutional layer, a maximum pooling (Maxpool) layer, a 3*3 convolutional layer, a 3*3 convolutional layer, a 3*3 convolutional layer, a 3*3 convolutional layer, a 3*3 convolutional layer, a 3*3 convolutional layer, a 3*3 convolutional layer, a 3*3 convolutional layer, a 3*3 convolutional layer, a 3*3 convolutional layer, a 3*3 convolutional layer, a 3*3 convolutional layer, a 3*3 convolutional layer, an average pooling (Argpool) layer, and an output layer connected in sequence.

[0124] In some embodiments, in order to accelerate the training and convergence speed of the resnet18 model, the improved resnet18 model provided by the embodiments of the present application adds at least one batch normalization (BN) layer. Optionally, the newly added BN layer can be located between two 3*3 convolutional layers.

[0125] As Figure 7 shown, the embodiments of the present disclosure provide an architecture diagram of an improved resnet18 model. Refer to Figure 7, the improved ResNet18 model includes an output layer, a 7×7 convolutional layer, a max pooling layer, a 3×3 convolutional layer, a 3×3 convolutional layer, a BN layer, a 3×3 convolutional layer, a 3×3 convolutional layer, a BN layer, a 3×3 convolutional layer, a 3×3 convolutional layer, a 3×3 convolutional layer, a BN layer, a 3×3 convolutional layer, a 3×3 convolutional layer, a BN layer, a 3×3 convolutional layer, a 3×3 convolutional layer, a 3×3 convolutional layer, a 3×3 convolutional layer, a BN layer, an average pooling layer, and an output layer connected in sequence.

[0126] Exemplarily, if the recognition result of the second row is "yes", it means that the second row recognition model's recognition result of the target person's behavior based on the target person's region of interest image is that the target person has a behavior of playing with a mobile phone; if the recognition result of the second row is "no", it means that the second row recognition model's recognition result of the target person's behavior based on the target person's region of interest image is that the target person does not have a behavior of playing with a mobile phone.

[0127] It should be noted that the present disclosure embodiment does not limit the execution order between step S103 and step S104. For example, step S103 can be executed first, and then step S104; or, step S104 can be executed first, and then step S103; or, step S103 and step S104 can be executed simultaneously.

[0128] It should be understood that the advantage of selecting the Inception-v3 model as the first behavior recognition model for mobile phone playing behavior recognition is that: the Inception-v3 model introduces the method of splitting a larger two-dimensional convolution into two smaller one-dimensional convolutions. For example, a 7×7 convolution kernel can be split into a 1×7 convolution kernel and a 7×1 convolution kernel. Of course, a 3×3 convolution kernel can also be split into a 1×3 convolution kernel and a 3×1 convolution kernel, which is called the factorization into small convolutions idea. The effect of this asymmetric convolution structure splitting in dealing with more and richer spatial features and increasing feature diversity can be better than that of symmetric convolution structure splitting, and at the same time, it can reduce the computational amount. For example, replacing 1 5×5 convolution with 2 3×3 convolutions can reduce the computational amount by 28%.

[0129] Similarly, the advantage of selecting the ResNet18 model as the second behavior recognition model for mobile phone playing behavior recognition is that: compared with the traditional VGG model, the complexity of the ResNet18 model is reduced, the required number of parameters decreases, and the network depth is deeper, and there will be no gradient disappearance phenomenon, solving the problem of deep network degradation, being able to accelerate network convergence, and preventing overfitting.

[0130] S105. If the recognition result of the first row is inconsistent with that of the second row, then based on the image of the region of interest, perform behavior recognition processing on the target person to determine whether the target person has the behavior of playing with the mobile phone.

[0131] It can be understood that the recognition model of the first row and the recognition model of the second row are two different recognition models. Therefore, there may be different behavior recognition results for the image of the region of interest of the same target user. For example, the recognition result of the first row indicates that the target person has the behavior of playing with the mobile phone, and the recognition result of the second row indicates that the target does not have the behavior of playing with the mobile phone; or, the recognition result of the first row indicates that the target person does not have the behavior of playing with the mobile phone, and the recognition result of the second row indicates that the target person has the behavior of playing with the mobile phone.

[0132] Based on Figure 3 The embodiments shown bring at least the following beneficial effects: Based on the image of the region of interest including the target person, double recognition is performed on whether the target person has the behavior of playing with the mobile phone through the recognition model of the first row and the recognition model of the second row, improving the accuracy of the recognition of the behavior of playing with the mobile phone. And when the first behavior recognition result output by the first behavior recognition model is inconsistent with the second behavior recognition result of the second behavior recognition model, perform behavior recognition processing on the target person again based on the image of the region of interest including the target person, so as to determine whether the target person has the behavior of playing with the mobile phone. It can be seen that a method for recognizing the behavior of playing with the mobile phone provided by the embodiments of the present disclosure performs multiple behavior recognitions on whether the user has the behavior of playing with the mobile phone, improving the accuracy of the recognition of the behavior of playing with the mobile phone. So as to send a reminder message in time when it is recognized that the target person has the behavior of playing with the mobile phone, and avoid the occurrence of adverse effects caused by the target person having the behavior of playing with the mobile phone.

[0133] In some embodiments, as Figure 8 shown, after step S104, the method further includes the following steps:

[0134] S106. If the recognition result of the first row is consistent with that of the second row, then based on the recognition result of the first row or the recognition result of the second row, determine whether the target person has the behavior of playing with the mobile phone.

[0135] It can be understood that if the recognition result of the first row is consistent with that of the second row, it means that the recognition models of the first row and the second row have consistent recognition results on whether the target person has the behavior of playing with the mobile phone. Since the recognition models of the first row and the second row are behavior recognition models based on different algorithms, and the behavior recognition models based on different algorithms output consistent recognition results, the accuracy of the recognition results is high. Then, it can be determined whether the target person has the behavior of playing with the mobile phone based on the recognition result of the first row or the second recognition result.

[0136] Exemplarily, if the first line of the recognition result indicates that the target person has the behavior of playing with a mobile phone, and the second line of the recognition result indicates that the target person has the behavior of playing with a mobile phone, then it is determined that the target person has the behavior of playing with a mobile phone. If the first line of the recognition result indicates that the target person does not have the behavior of playing with a mobile phone, and the second line of the recognition result indicates that the target person does not have the behavior of playing with a mobile phone, then it is determined that the target person does not have the behavior of playing with a mobile phone.

[0137] In some embodiments, as Figure 9 shown, the above step S105 can be specifically implemented as the following steps:

[0138] S1051. Input the region of interest image into the mobile phone detection model, and input the region of interest image into the person detection model.

[0139] It can be understood that if the target person has the behavior of playing with a mobile phone, it is necessary that there is a mobile phone in the area where the target person is located, that is, there is a mobile phone in the region of interest image of the target person. If there is no mobile phone in the region of interest image of the target person, that is, there is no mobile phone in the area where the target person is located, then the target person has no possibility of playing with a mobile phone.

[0140] In some embodiments, a trained mobile phone detection model is pre-stored in the memory of the behavior recognition device. In order to identify whether the target person has the possibility of playing with a mobile phone, the region of interest image can be input into the mobile phone detection model to detect whether there is a mobile phone in the region of interest image.

[0141] Specifically, after the region of interest image of the target person is input into the mobile phone detection model, if the mobile phone detection model outputs at least one mobile phone frame, it means that there is a mobile phone in the region of interest image, and the target person has the possibility of playing with a mobile phone. If the mobile phone detection model outputs 0 mobile phone frames, it means that there is no mobile phone in the region of interest image, and the target person does not have the possibility of playing with a mobile phone.

[0142] As can be seen from the above, the region of interest image of the target person is the image of the area where the detection frame of the target person is located. The region of interest image of the target person can not only contain the target person, but also contain other people (also called non-target people) and items (such as walls, mobile phones, etc.) other than the target person due to the shooting angle of the shooting device. For example, in a case where a non-target person stands relatively close to the target person, the region of interest image of the target person can also include this non-target person.

[0143] It can be understood that in the case where the region of interest image of the target person contains non-target people other than the target person, the non-target people other than the target person in the region of interest image will interfere with the recognition result of whether the target person has the behavior of playing with a mobile phone.

[0144] In some embodiments, a trained person detection model is pre-stored in the memory of the behavior recognition device. To identify whether there are non-target persons other than the target person in the region of interest image, the region of interest image can be input into the person detection model to detect whether there are non-target persons in the region of interest image.

[0145] Specifically, after the region of interest image of the target person is input into the person detection model, if the person detection model outputs only one person bounding box, then this person bounding box is also the person bounding box of the target person, indicating that there are no non-target persons in the region of interest image, that is, there are no non-target persons in the region where the target person is located. If the person detection model outputs at least one person bounding box, it means that there are non-target persons in the region of interest image, that is, there are non-target persons in the region where the target person is located.

[0146] In some embodiments, the above-mentioned mobile phone detection model includes: yolov5 model, yolox model.

[0147] In some embodiments, the above-mentioned pedestrian detection model includes: yolov5 model, yolov4 model, yolov3 model, mobilenetv1_ssd model, mobilenetv2_ssd model, and mobilenetv3_ssd model.

[0148] S1052. If no mobile phone is detected from the region of interest image, it is determined that the target person does not have the behavior of playing with the mobile phone.

[0149] It can be understood that the region of interest image reflects the region where the target person is located. If no mobile phone is detected from the region of interest image, it means that to a certain extent, there is no mobile phone in the region where the target person is located. If there is no mobile phone in the region where the target person is located, then there is no possibility that the target person has the behavior of playing with the mobile phone. Therefore, if no mobile phone is detected from the region of interest image, it is determined that the target person does not have the behavior of playing with the mobile phone.

[0150] It should be understood that the advantage of step S1052 is that: according to whether there is a mobile phone in the region of interest image, it is directly determined whether the target person has the behavior of playing with the mobile phone. The behavior recognition device does not need to perform cumbersome calculations, and can improve the accuracy of identifying the target user's behavior of playing with the mobile phone while reducing the calculation amount of the behavior recognition device.

[0151] S1053. If a mobile phone is detected from the region of interest image, then according to the mobile phone bounding box output by the mobile phone detection model and the person bounding box output by the person detection model, it is determined whether the target person has the behavior of playing with the mobile phone.

[0152] It is understandable that if a mobile phone is detected in the image of the region of interest, it means that there is a mobile phone in the region where the target person is located, that is, there is a possibility that the target person is engaged in the behavior of playing with the mobile phone.

[0153] In some embodiments, after detecting a mobile phone in the image of the region of interest, it is possible to determine whether the target person is engaged in the behavior of playing with the mobile phone based on the mobile phone bounding box output by the mobile phone detection model and the person bounding box output by the person detection model.

[0154] Exemplarily, determining whether the target person is engaged in the behavior of playing with the mobile phone based on the mobile phone bounding box output by the mobile phone detection model and the person bounding box output by the person detection model may specifically include the following situations.

[0155] Situation 1: The person detection model outputs only one person bounding box.

[0156] As can be seen from the above S1051, when the person detection model outputs only one person bounding box, it means that there are no non-target persons in the region where the target person is located. In the case of Situation 1, as Figure 10 shown, step S1053 can be specifically implemented as the following steps:

[0157] S201. Determine the overlap degree between the mobile phone bounding box and the person bounding box.

[0158] Among them, the overlap degree between the mobile phone bounding box and the person bounding box is positively correlated with the possibility that the target person is engaged in the behavior of playing with the mobile phone, that is, the higher the overlap degree, the higher the possibility that the target person is engaged in the behavior of playing with the mobile phone.

[0159] It is understandable that usually, if the target person is engaged in the behavior of playing with the mobile phone, the mobile phone should be present around the target person. The closer the mobile phone is to the target person, the higher the possibility that the target person is engaged in the behavior of playing with the mobile phone. In the image, both the target person and the mobile phone exist in the form of detection bounding boxes, and the overlap degree between the mobile phone bounding box and the person bounding box can reflect the distance between the target person and the mobile phone. Therefore, the overlap degree between the mobile phone bounding box and the person bounding box is positively correlated with the possibility that the target person is engaged in the behavior of playing with the mobile phone.

[0160] Exemplarily, the process of determining the overlap degree between the mobile phone bounding box and the person bounding box is as follows:

[0161] Step 1. Determine the area of the overlapping region between the mobile phone bounding box and the person bounding box in the image of the region of interest.

[0162] It is easy to understand that when the distance between the mobile phone and the target person is within a certain range, there is an overlapping region between the corresponding mobile phone bounding box and the corresponding person bounding box of the target person.

[0163] As Figure 11As shown, based on the upper boundary, lower boundary, left boundary, and right boundary of the mobile phone frame, the shape and coordinates of the pixel region corresponding to the mobile phone frame in the region of interest image can be determined. Among them, the shape of the pixel region corresponding to the mobile phone frame in the region of interest image is a rectangle, and the coordinates of the pixel region corresponding to the mobile phone frame in the region of interest image are (X a min, Y a min, X a max, Y a max), where X a min is the minimum abscissa of the mobile phone frame in the pixel region, Y a min is the minimum ordinate of the mobile phone frame in the pixel region, X a max is the maximum abscissa of the mobile phone frame in the pixel region, and Y a max is the maximum ordinate of the mobile phone frame in the pixel region. Furthermore, based on the coordinates of the pixel region corresponding to the mobile phone frame in the region of interest image, the area occupied by the mobile phone frame in the region of interest image can be obtained.

[0164] Similarly, based on the upper boundary, lower boundary, left boundary, and right boundary of the person frame, the shape and coordinates of the pixel region corresponding to the person frame in the region of interest image can be determined. Among them, the shape of the pixel region corresponding to the person frame in the region of interest image is a rectangle, and the coordinates of the pixel region corresponding to the person frame in the region of interest image are (X b min, Y b min, X b max, Y b max), where X b min is the minimum abscissa of the person frame in the pixel region, Y b min is the minimum ordinate of the person frame in the pixel region, X b max is the maximum abscissa of the person frame in the pixel region, and Y b max is the maximum ordinate of the person frame in the pixel region. Furthermore, based on the coordinates of the pixel region corresponding to the person frame in the region of interest image, the area occupied by the person frame in the region of interest image can be obtained. Among them, Figure 11 the dashed box shown on the left in the figure is the mobile phone frame, and the dashed box shown on the right is the person frame.

[0165] After obtaining the coordinates of the pixel region corresponding to the mobile phone frame in the region of interest image and the coordinates of the pixel region corresponding to the person frame in the region of interest image, based on the coordinates of the pixel region corresponding to the mobile phone frame in the region of interest image and the coordinates of the pixel region corresponding to the person frame in the region of interest image, the overlapping region between the mobile phone frame and the person frame in the region of interest image can be obtained, and then the area of the overlapping region can be obtained.

[0166] Exemplarily, the relationship between the coordinates of the mobile phone frame in the image of the region of interest, the coordinates of the person frame in the image of the region of interest, and the overlapping region can be shown as in the following formula (1):

[0167] A = renwu ∩ shouji Formula (1)

[0168] Wherein, A is used to represent the overlapping region, renwu is used to represent the coordinates of the person frame in the image of the region of interest, and shouji is used to represent the coordinates of the mobile phone frame in the image of the region of interest.

[0169] Step 2: Use the ratio between the area of the overlapping region and the area of the region occupied by the mobile phone frame in the region of interest as the overlapping degree.

[0170] Exemplarily, the relationship between the overlapping degree, the area of the overlapping region, and the area of the region occupied by the mobile phone frame in the image of the region of interest can be shown as in the following formula (2):

[0171]

[0172] Wherein, B is used to represent the overlapping degree, A sq is used to represent the area of the overlapping region, and shouji sq is used to represent the area of the region occupied by the mobile phone frame in the image of the region of interest.

[0173] S202: If the overlapping degree is greater than or equal to the preset overlapping degree threshold, determine that the target person has the behavior of playing with the mobile phone.

[0174] Wherein, the preset overlapping degree threshold can be set in advance by the management personnel according to manual experience. For example, the preset overlapping degree threshold is 80%. That is, when the ratio between the area of the overlapping region between the mobile phone frame and the person frame and the area of the mobile phone frame is greater than or equal to 80%, it is determined that the target person has the behavior of playing with the mobile phone.

[0175] It should be understood that generally, when the mobile phone exists around the target person, there is a possibility that the target person has the behavior of playing with the mobile phone. However, even when the mobile phone exists around the target person, the target person does not necessarily have the behavior of playing with the mobile phone. Therefore, a method for identifying the behavior of playing with the mobile phone provided by the embodiments of the present disclosure determines that the target person has the behavior of playing with the mobile phone based on the overlapping degree being greater than or equal to the preset overlapping degree threshold, improving the accuracy of identifying the behavior of playing with the mobile phone.

[0176] S203: If the overlapping degree is less than the preset overlapping degree threshold, determine that the target person does not have the behavior of playing with the mobile phone.

[0177] It can be understood that if the overlapping degree is less than the preset overlapping degree threshold, it means that the possibility that the target person has the behavior of playing with the mobile phone is low. Therefore, it can be determined that the target person does not have the behavior of playing with the mobile phone.

[0178] As a possible implementation, in order to reduce the computational load of the behavior recognition device, for example Figure 12 as shown, the above-mentioned mobile phone playing behavior recognition method may further include step S301 before step S201, and step S201 may be specifically implemented as step S303.

[0179] S301. Based on the mobile phone frame and the person frame, determine the distance between the target person and the mobile phone.

[0180] The above steps S201 to S203 are described under the default condition that there is an overlapping area between the mobile phone frame and the person frame. It can be understood that if the target person does not have the behavior of playing with the mobile phone, there is no overlapping area between the mobile phone frame and the person frame. If the overlapping degree between the mobile phone frame and the person frame is continuously calculated when there is no overlapping area between the mobile phone frame and the person frame, it will increase the computational load of the behavior recognition device and cause waste of the computational resources of the behavior recognition device.

[0181] Based on this, before determining the overlapping degree between the mobile phone frame and the person frame, the behavior recognition device may obtain the coordinates of the pixel area corresponding to the center position of the mobile phone frame in the image of the region of interest according to the coordinates of the pixel area corresponding to the mobile phone frame in the image of the region of interest abbreviated as the coordinates of the center position of the mobile phone frame. And according to the coordinates of the pixel area corresponding to the person frame in the image of the region of interest, obtain the coordinates of the pixel area corresponding to the center position of the person frame in the image of the region of interest abbreviated as the coordinates of the center position of the person frame.

[0182] Furthermore, according to the coordinates of the center position of the mobile phone frame and the coordinates of the center position of the person frame, the distance between the center position of the mobile phone frame and the center position of the person frame can be obtained. Take the distance between the center position of the mobile phone frame and the center position of the person frame as the distance between the target person and the mobile phone.

[0183] S302. When the distance between the target person and the mobile phone is greater than the preset distance threshold, determine that the target person does not have the behavior of playing with the mobile phone.

[0184] It can be understood that when the distance between the mobile phone and the target person is greater than the preset distance threshold, it means that there is no intersection between the mobile phone frame and the person frame, that is, there is no overlapping area between the mobile phone frame and the person frame. When there is no overlapping area between the mobile phone frame and the person frame, it means that the mobile phone is far away from the target person, and the possibility that the target person has the behavior of playing with the mobile phone is relatively low. It can be directly determined that the target person does not have the behavior of playing with the mobile phone. Furthermore, there is no need to calculate the overlapping degree between the mobile phone frame and the person frame, reducing the computational load of the behavior recognition device and reducing the waste of the computational resources of the behavior recognition device.

[0185] Among them, the distance threshold is used to indicate the distance threshold when the mobile phone frame and the person frame do not intersect.

[0186] As a possible implementation, the preset distance threshold can be calculated in real time by the behavior recognition device based on the resolution of the image of the region of interest.

[0187] Exemplarily, the embodiments of the present disclosure provide a method for determining a preset distance threshold. The behavior recognition device uses the distance from the center position of the mobile phone frame to any corner of the upper left, upper right, lower left, and lower right corners of the mobile phone frame, and the distance from the center position of the person frame to any corner of the upper left, upper right, lower left, and lower right corners of the person frame, and takes the sum of the two distances as the preset distance threshold.

[0188] As another possible implementation, the preset distance threshold can be pre-set by the management personnel according to manual experience.

[0189] S303. When the distance between the target person and the mobile phone is less than or equal to the preset distance threshold, determine the degree of overlap between the mobile phone frame and the person frame.

[0190] As a possible implementation, the above step S201 can be specifically implemented as: when the distance between the target person and the mobile phone is less than or equal to the preset distance threshold, determine the degree of overlap between the mobile phone frame and the person frame.

[0191] It can be understood that when the distance between the target person and the mobile phone is less than or equal to the preset distance threshold, it means that there is an intersection between the mobile phone frame and the person frame, that is, there is an overlapping area between the mobile phone frame and the person frame. When there is an overlapping area between the mobile phone frame and the person frame, it means that there is a mobile phone in the area where the target person represented by the person frame is located, that is, there is a possibility that the target person has the behavior of playing with the mobile phone. Whether the target person has the behavior of playing with the mobile phone can be further determined according to the degree of overlap between the person frame and the mobile phone frame.

[0192] Regarding the specific implementation of determining the degree of overlap between the mobile phone frame and the person frame, reference can be made to the description of the above step S201, which will not be elaborated here.

[0193] The above embodiments mainly introduce the situation when the person detection model outputs only one person frame. In some embodiments, the method for recognizing the behavior of playing with a mobile phone provided by the embodiments of the present disclosure further includes the following situations:

[0194] Situation 2: The person detection model outputs multiple person frames.

[0195] As can be seen from the above S1051, when the person detection model outputs multiple person frames, it means that there are non-target persons in the area where the target person is located. In the case of Situation 2, such as Figure 13As shown, step S1053 can also be specifically implemented as the following steps:

[0196] S401. Determine the person box of the target person and the person boxes of non-target persons from multiple person boxes.

[0197] In some embodiments, in step S102 above, when the behavior recognition device performs image segmentation on the image to be recognized based on the image segmentation model and obtains the region-of-interest images corresponding to each person, the behavior recognition device establishes an identity identifier corresponding to each person for each person, and one identity identifier is used to uniquely indicate one person.

[0198] In the case where the person detection model outputs multiple person boxes, the behavior recognition device can determine the person box of the target person and the person boxes of non-target persons from the multiple person boxes based on the identity identifiers of the persons corresponding to each person box in the multiple person boxes.

[0199] S402. Determine the distance between the target person and the mobile phone based on the person box of the target person, the mobile phone box, and the region-of-interest image.

[0200] Optionally, determining the distance between the target person and the mobile phone based on the person box of the target person, the mobile phone box, and the region-of-interest image may include one or more of the following methods:

[0201] Method 1. The behavior recognition device determines the distance between the target person and the mobile phone based on the central position of the target person and the central position of the mobile phone.

[0202] Exemplarily, the behavior recognition device can determine the shape and coordinates of the pixel region corresponding to the person box of the target person in the region-of-interest image according to the upper boundary, lower boundary, left boundary, and right boundary of the person box of the target person. Among them, the shape of the pixel region corresponding to the person box of the target person in the region-of-interest image is a rectangle.

[0203] Similarly, the behavior recognition device can determine the shape and coordinates of the pixel region corresponding to the mobile phone box in the region-of-interest image according to the upper boundary, lower boundary, left boundary, and right boundary of the mobile phone box. Among them, the shape of the pixel region corresponding to the mobile phone box in the region-of-interest image is a rectangle.

[0204] After the behavior recognition device obtains the coordinates of the pixel region corresponding to the person box of the target person in the region-of-interest image, it can obtain the coordinates of the pixel region corresponding to the central position of the target person in the region-of-interest image.

[0205] Similarly, after the behavior recognition device obtains the coordinates of the pixel region corresponding to the mobile phone box in the region-of-interest image, it can also obtain the coordinates of the pixel region corresponding to the central position of the mobile phone in the region-of-interest image.

[0206] Based on the coordinates of the pixel region corresponding to the center position of the target person in the region of interest image and the coordinates of the pixel region corresponding to the center position of the mobile phone in the region of interest image, the distance between the center position of the mobile phone and the center position of the target person can be obtained. Furthermore, the distance between the center position of the mobile phone and the center position of the target person is used as the distance between the target person and the mobile phone.

[0207] Method 2: The behavior recognition device determines the distance between the target person and the mobile phone based on the center position of the target person's hand and the center position of the mobile phone.

[0208] In the above Method 1, the distance between the center position of the target person and the center position of the mobile phone is used as the distance between the target person and the mobile phone. It can be understood that usually, if the target person is engaged in the behavior of playing with the mobile phone, the target person plays with the mobile phone through the hand. In order to improve the accuracy of recognizing the behavior of playing with the mobile phone, the embodiments of the present disclosure propose to use the distance between the center position of the target person's hand and the center position of the mobile phone as the distance between the target person and the mobile phone.

[0209] Specifically, the above Method 2 may include the following steps:

[0210] S1. Based on the person frame of the target person and the region of interest image, perform hand recognition on the target person to determine the center position of the target person's hand.

[0211] In some embodiments, a trained hand recognition model is pre-stored in the memory of the service area. The service area can input the region of interest image including the person frame of the target person into the hand recognition model to obtain the hand frame of the target person.

[0212] According to the upper boundary, lower boundary, left boundary, and right boundary of the hand frame of the target person, the shape and coordinates of the pixel region corresponding to the hand frame of the target person in the region of interest image can be determined. Among them, the shape of the pixel region corresponding to the target person in the region of interest image is a rectangle. Furthermore, according to the coordinates of the pixel region corresponding to the hand frame of the target person in the region of interest image, the center position of the target person's hand can be obtained.

[0213] In some embodiments, the above hand recognition model may be a hand recognition model based on the Faster R-CNN algorithm.

[0214] S2. Based on the mobile phone frame and the region of interest image, determine the center position of the mobile phone.

[0215] Regarding determining the center position of the mobile phone based on the mobile phone frame and the region of interest image, reference may be made to the method for confirming the center position of the mobile phone in the above Method 1, which will not be elaborated here.

[0216] S3. Determine the distance between the target person and the mobile phone based on the central position of the target person's hand and the central position of the mobile phone.

[0217] Optionally, the distance between the central position of the target person's hand and the central position of the mobile phone can be obtained according to the coordinates of the pixel area corresponding to the central position of the target person's hand in the region of interest image and the coordinates of the pixel area corresponding to the central position of the mobile phone in the region of interest image. Furthermore, the distance between the central position of the target person's hand and the central position of the mobile phone is used as the distance between the target person and the mobile phone.

[0218] Method 3. The behavior recognition device determines the distance between the target person and the mobile phone based on the central position of the target person's eye and the central position of the mobile phone.

[0219] It should be understood that usually, when the target person is performing a mobile phone usage behavior, the target person's eyes will look at the mobile phone. Therefore, in the embodiments of the present disclosure, the distance between the central position of the target person's eye and the central position of the mobile phone is used as the distance between the target person and the mobile phone.

[0220] Specifically, the above Method 3 may include the following steps:

[0221] P1. Perform eye recognition on the target person based on the person frame of the target person and the region of interest image to determine the central position of the target person's eye.

[0222] In some embodiments, an eye recognition model is pre-stored in the memory of the behavior recognition device. The region of interest image including the person frame of the target person can be input into the eye recognition model to obtain the eye frame of the target person.

[0223] The method for obtaining the central position of the target person's eye based on the eye frame of the target person can refer to the method for obtaining the central position of the target person's hand based on the hand frame of the target person in the above S1, which will not be elaborated here.

[0224] In some embodiments, the above eye recognition model may be an eye recognition model based on the scale-invariant feature transform (SIFT) algorithm.

[0225] P2. Determine the central position of the mobile phone based on the mobile phone frame and the region of interest image.

[0226] P3. Determine the distance between the target person and the mobile phone based on the central position of the target person's eye and the central position of the mobile phone.

[0227] For the descriptions of P2 and P3, reference may be made to the descriptions of S2 and S3 above, which will not be elaborated here.

[0228] S403. Based on the person frame, mobile phone frame, and region of interest image of the non-target person, determine the distance between the non-target person and the mobile phone.

[0229] For the description of step S403, reference may be made to the description of step S402 above, which will not be elaborated here.

[0230] In some embodiments, when there are multiple non-target persons in the region of interest image of the target person, the behavior recognition device performs the above calculations for each non-target person among the multiple non-target persons to obtain the distance between each non-target person and the mobile phone.

[0231] It should be noted that to ensure the accuracy of mobile phone usage behavior recognition, if the behavior recognition device uses method 1 in the above S402 to determine the distance between the target person and the mobile phone, then the behavior recognition device also uses method 1 in the above S402 to determine the distance between the non-target person and the mobile phone. Similarly, if the behavior recognition device uses method 2 in the above S402 to determine the distance between the target person and the mobile phone, then the behavior recognition device also uses method 2 in the above S402 to determine the distance between the non-target person and the mobile phone.

[0232] The embodiments of the present disclosure do not limit the execution order between step S402 and step S403. For example, step S402 can be executed first, and then step S403; or, step S403 can be executed first, and then step S402; or, step S402 and step S403 can be executed simultaneously.

[0233] S404. When the distance between the target person and the mobile phone is less than the distances between all non-target persons and the mobile phone, determine that the target person has the behavior of using the mobile phone.

[0234] It can be understood that if the distance between the target person and the mobile phone is less than the distances between all non-target persons and the mobile phone, it means that the target person is the closest to the mobile phone, that is, the target person is the person among multiple persons with the highest possibility of using the mobile phone, so it is determined that the target person has the behavior of using the mobile phone.

[0235] S405. When the distance between the target person and the mobile phone is greater than or equal to the distance between any non-target person and the mobile phone, determine that the target person does not have the behavior of using the mobile phone.

[0236] It can be understood that when the distance between the target person and the mobile phone is greater than or equal to the distance between any non-target person and the mobile phone, it means that the target person is not the user closest to the mobile phone, and the possibility that the target user has the behavior of playing with the mobile phone is relatively low. To avoid misidentification, the behavior recognition device determines that the target person does not have the behavior of playing with the mobile phone.

[0237] Based on Figure 13 the embodiments shown, at least the following beneficial effects are brought: when the person detection model outputs multiple person bounding boxes, it means that there are not only the target person but also non-target persons in the area where the target person is located. To exclude the influence of non-target persons on the recognition of whether the target person has the behavior of playing with the mobile phone, according to the distance between each person and the mobile phone, when the distance between the target person and the mobile phone is the shortest, it is determined that the target person has the behavior of playing with the mobile phone, thereby excluding the influence of non-target persons on the recognition of whether the target person has the behavior of playing with the mobile phone and improving the accuracy of the recognition of the behavior of playing with the mobile phone.

[0238] The following uses a specific example to illustrate a method for recognizing the behavior of playing with a mobile phone provided by the embodiments of the present disclosure.

[0239] As Figure 14 shown, assume Figure 14 the image shown is the image to be recognized, and the image to be recognized includes Person 1 and Person 2.

[0240] First, perform image segmentation processing on the image to be recognized to obtain the region of interest image of Person 1 and the region of interest image of Person 2.

[0241] Input the region of interest image of Person 1 into the first behavior recognition model and the second behavior recognition model respectively to obtain the first behavior recognition result and the second behavior recognition result of Person 1. Input the region of interest image of Person 2 into the first behavior recognition model and the second behavior recognition model respectively to obtain the first behavior recognition result and the second behavior recognition result of Person 2.

[0242] Assume that the first behavior recognition result of Person 1 is consistent with the second behavior recognition result, and the first behavior recognition result indicates that the task machine has the behavior of playing with the mobile phone, then it is determined that Person 1 has the behavior of playing with the mobile phone.

[0243] Assume that the first behavior recognition result of Person 2 is inconsistent with the second behavior recognition result, which means that it is impossible to confirm whether Person 2 has the behavior of playing with the mobile phone. The region of interest image of Person 2 can be input into the person detection model and the mobile phone detection model to detect the persons existing in the area where Person 2 is located and whether there is a mobile phone in the area where Person 2 is located.

[0244] Assume that the person detection model only outputs one person bounding box, and the mobile phone detection model outputs one mobile phone bounding box, indicating that there is only one person in the area where person 2 is located, and there is a mobile phone in the area where person 2 is located. Then, based on the overlap degree between the mobile phone bounding box output by the mobile phone detection model and the person bounding box output by the person detection model, it can be determined whether person 2 is playing with the mobile phone.

[0245] Assume that the preset overlap degree threshold is 80%. If the overlap degree between the mobile phone bounding box and the person bounding box is 85%, it is determined that person 2 has the behavior of playing with the mobile phone. The behavior recognition device outputs the final recognition result, that is, person 1 has the behavior of playing with the mobile phone, and person 2 has the behavior of playing with the mobile phone.

[0246] The above mainly introduces the solution provided by the embodiments of the present disclosure from the perspective of methods. To implement the above functions, it includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0247] The embodiments of the present disclosure also provide a behavior recognition device. As Figure 15 shown, the behavior recognition device 300 may include: a communication unit 301 and a processing unit 302. In some embodiments, the above behavior recognition device 300 may further include a storage unit 303.

[0248] In some embodiments, the above communication unit 301 is used to obtain the image to be recognized.

[0249] The above processing unit 302 is used to: extract the region of interest image containing the target person from the image to be recognized; input the region of interest image into the first behavior recognition model to obtain the first behavior recognition result of the target person, and the first behavior recognition result is used to indicate whether the target person has the behavior of playing with the mobile phone; input the region of interest image into the second behavior recognition model to obtain the second behavior recognition result of the target person, and the second behavior recognition result is used to indicate whether the target person has the behavior of playing with the mobile phone; if the first behavior recognition result is inconsistent with the second behavior recognition result, then based on the region of interest image, perform behavior recognition processing on the target person to determine whether the target person has the behavior of playing with the mobile phone.

[0250] In some other embodiments, the above-mentioned processing unit 302 is further configured to determine whether the target person has the behavior of playing with a mobile phone based on the recognition result of the first behavior or the recognition result of the second behavior if the recognition result of the first behavior is consistent with the recognition result of the second behavior.

[0251] In some other embodiments, the above-mentioned processing unit 302 is specifically configured to: input the image of the region of interest into a mobile phone detection model and input the image of the region of interest into a person detection model; determine that the target person does not have the behavior of playing with a mobile phone if no mobile phone is detected from the image of the region of interest; and determine whether the target person has the behavior of playing with a mobile phone according to the mobile phone frame output by the mobile phone detection model and the person frame output by the person detection model if a mobile phone is detected from the image of the region of interest.

[0252] In some other embodiments, when the person detection model outputs only one person frame, the above-mentioned processing unit 302 is specifically configured to: determine the degree of overlap between the mobile phone frame and the person frame; determine that the target person has the behavior of playing with a mobile phone if the degree of overlap is greater than or equal to a preset overlap threshold; and determine that the target person does not have the behavior of playing with a mobile phone if the degree of overlap is less than the preset overlap threshold.

[0253] In some other embodiments, the above-mentioned processing unit 302 is specifically configured to: determine the area of the overlapping region between the mobile phone frame and the person frame in the image of the region of interest.

[0254] Use the ratio between the area of the overlapping region and the area of the region occupied by the mobile phone frame in the region of interest as the degree of overlap.

[0255] In some other embodiments, the above-mentioned processing unit is further configured to: determine the distance between the target person and the mobile phone based on the mobile phone frame and the person frame; determine that the target person does not have the behavior of playing with a mobile phone when the distance between the target person and the mobile phone is greater than a preset distance threshold; and the above-mentioned processing unit is specifically configured to determine the degree of overlap between the mobile phone frame and the person frame when the distance between the target person and the mobile phone is less than or equal to the preset distance threshold.

[0256] In some other embodiments, when the person detection model outputs multiple person bounding boxes, the processing unit 302 is specifically configured to: determine the person bounding box of the target person and the person bounding boxes of non-target persons from the multiple person bounding boxes, where the non-target persons are other persons in the region of interest image except the target person; determine the distance between the target person and the mobile phone based on the person bounding box of the target person, the mobile phone bounding box, and the region of interest image; determine the distance between the non-target persons and the mobile phone based on the person bounding boxes of the non-target persons, the mobile phone bounding box, and the region of interest image; determine that the target person has a behavior of playing with the mobile phone when the distance between the target person and the mobile phone is less than the distances between all non-target persons and the mobile phone; and determine that the target person does not have a behavior of playing with the mobile phone when the distance between the target person and the mobile phone is greater than or equal to the distance between the target person and any one of the non-target persons.

[0257] In some other embodiments, the processing unit 302 is specifically configured to: perform hand recognition on the target person based on the person bounding box of the target person and the region of interest image to determine the central position of the hand of the target person; determine the central position of the mobile phone based on the mobile phone bounding box and the region of interest image; and determine the distance between the target person and the mobile phone according to the central position of the hand of the target person and the central position of the mobile phone.

[0258] In some other embodiments, the storage unit 303 is configured to store the image to be recognized.

[0259] In some other embodiments, the storage unit 303 is configured to store the first behavior recognition model, the second behavior recognition model, the person detection model, the mobile phone detection model, the hand recognition model, the identity recognition model, and the image segmentation model.

[0260] Figure 15 The units in can also be referred to as modules. For example, the processing unit can be referred to as a processing module.

[0261] Figure 15When each unit in [the relevant context] is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a behavior recognition device, or a network device, etc.) or a processor to execute all or part of the steps of the methods in the various embodiments of this application. The storage media storing the computer software product include: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs, etc., which are all media that can store program codes.

[0262] Some embodiments of the present disclosure provide a computer-readable storage medium (for example, a non-transitory computer-readable storage medium). Computer program instructions are stored in this computer-readable storage medium. When the computer program instructions run on the processor of the computer, they cause the processor to execute the mobile phone behavior recognition method as described in any one of the above embodiments.

[0263] Exemplarily, the above computer-readable storage medium may include, but is not limited to: magnetic storage devices (such as hard disks, floppy disks, or magnetic tapes, etc.), optical discs (such as CDs (Compact Disks), DVDs (Digital Versatile Disks), etc.), smart cards, and flash memory devices (such as EPROMs (Erasable Programmable Read-Only Memories), cards, sticks, or key drives, etc.). The various computer-readable storage media described in the present disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media that can store, contain, and / or carry instructions and / or data.

[0264] Some embodiments of the present disclosure also provide a computer program product. For example, this computer program product is stored on a non-transitory computer-readable storage medium. This computer program product includes computer program instructions. When the computer program instructions are executed on the computer, they cause the computer to execute the mobile phone behavior recognition method as described in the above embodiments.

[0265] Some embodiments of the present disclosure also provide a computer program. When the computer program is executed on a computer, the computer program causes the computer to execute the mobile phone playing behavior recognition method as described in the above embodiments.

[0266] The beneficial effects of the above computer-readable storage medium, computer program product and computer program are the same as those of the mobile phone playing behavior recognition method described in some of the above embodiments, and will not be elaborated here.

[0267] The above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure who thinks of changes or substitutions should be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A method for identifying mobile phone usage behaviors, characterized in that, The method includes: Obtain an image to be recognized; Extract a region of interest image containing the target person from the image to be recognized; Input the region of interest image into a first behavior recognition model to obtain a first behavior recognition result of the target person, where the first behavior recognition result is used to indicate whether the target person has a behavior of playing with a mobile phone; Input the region of interest image into a second behavior recognition model to obtain a second behavior recognition result of the target person, where the second behavior recognition result is used to indicate whether the target person has a behavior of playing with a mobile phone; If the first behavior recognition result is inconsistent with the second behavior recognition result, then based on the region of interest image, perform behavior recognition processing on the target person to determine whether the target person has a behavior of playing with a mobile phone; wherein, based on the region of interest image, performing behavior recognition processing on the target person to determine whether the target person has a behavior of playing with a mobile phone includes: Input the region of interest image into a mobile phone detection model and input the region of interest image into a person detection model; If no mobile phone is detected from the region of interest image, determine that the target person does not have a behavior of playing with a mobile phone; If a mobile phone is detected from the region of interest image, then based on the mobile phone frame output by the mobile phone detection model and the person frame output by the person detection model, determine whether the target person has a behavior of playing with a mobile phone.

2. The method according to claim 1, wherein The method further includes: If the first behavior recognition result is consistent with the second behavior recognition result, then based on the first behavior recognition result or the second behavior recognition result, determine whether the target person has a behavior of playing with a mobile phone.

3. The method according to claim 1, characterized in that, When the person detection model outputs only one person frame, the determining whether the target person has a behavior of playing with a mobile phone based on the mobile phone frame output by the mobile phone detection model and the person frame output by the person detection model includes: Determine the overlap degree between the mobile phone frame and the person frame; If the overlap degree is greater than or equal to a preset overlap degree threshold, determine that the target person has a behavior of playing with a mobile phone; If the overlap degree is less than the preset overlap degree threshold, determine that the target person does not have a behavior of playing with a mobile phone.

4. The method according to claim 3, characterized in that, The determining the overlap degree between the mobile phone frame and the person frame includes: Determine the area of the overlapping region between the mobile phone frame and the person frame in the region of interest image; Use the ratio between the area of the overlapping region and the area of the region occupied by the mobile phone frame in the region of interest as the overlap degree.

5. The method according to claim 3, characterized in that, Before the determining the overlap degree between the mobile phone frame and the person frame, the method further includes: Based on the mobile phone frame and the person frame, determine the distance between the target person and the mobile phone; When the distance between the target person and the mobile phone is greater than a preset distance threshold, determine that the target person does not have a behavior of playing with a mobile phone; The determining the overlap degree between the mobile phone frame and the person frame includes: When the distance between the target person and the mobile phone is less than or equal to the preset distance threshold, determine the overlap degree between the mobile phone frame and the person frame.

6. The method according to claim 1, wherein When the person detection model outputs multiple person bounding boxes, determining whether the target person has a behavior of playing with a mobile phone based on the mobile phone bounding box output by the mobile phone detection model and the person bounding boxes output by the person detection model includes: Determining the person bounding box of the target person and the person bounding boxes of non-target persons from the multiple person bounding boxes, where the non-target persons are other persons in the region of interest image except the target person; Based on the person bounding box of the target person, the mobile phone bounding box, and the region of interest image, determining the distance between the target person and the mobile phone; Based on the person bounding boxes of the non-target persons, the mobile phone bounding box, and the region of interest image, determining the distances between the non-target persons and the mobile phone; When the distance between the target person and the mobile phone is less than the distances between all non-target persons and the mobile phone, determining that the target person has a behavior of playing with a mobile phone; When the distance between the target person and the mobile phone is greater than or equal to the distance between any one of the non-target persons and the mobile phone, determining that the target person does not have a behavior of playing with a mobile phone.

7. The method according to claim 6, wherein The determining the distance between the target person and the mobile phone based on the person bounding box of the target person, the mobile phone bounding box, and the region of interest image includes: Based on the person bounding box of the target person and the region of interest image, performing hand recognition on the target person to determine the central position of the hands of the target person; Based on the mobile phone bounding box and the region of interest image, determining the central position of the mobile phone; According to the central position of the hands of the target person and the central position of the mobile phone, determining the distance between the target person and the mobile phone.

8. The method according to any one of claims 1 to 7, characterized in that, The first behavior recognition model is an Inception network model, and the second behavior recognition model is a Residual network model.

9. An action recognition device, characterized in that, The behavior recognition device includes: A communication unit, configured to obtain an image to be recognized; A processing unit, configured to: extract a region of interest image including the target person from the image to be recognized; input the region of interest image into the first behavior recognition model to obtain a first behavior recognition result of the target person, where the first behavior recognition result is used to indicate whether the target person has a behavior of playing with a mobile phone; input the region of interest image into the second behavior recognition model to obtain a second behavior recognition result of the target person, where the second behavior recognition result is used to indicate whether the target person has a behavior of playing with a mobile phone; if the first behavior recognition result is inconsistent with the second behavior recognition result, perform behavior recognition processing on the target person based on the region of interest image to determine whether the target person has a behavior of playing with a mobile phone; Wherein, the processing unit is specifically configured to: input the region of interest image into a mobile phone detection model, and input the region of interest image into a person detection model; if no mobile phone is detected from the region of interest image, determine that the target person does not have a behavior of playing with a mobile phone; if a mobile phone is detected from the region of interest image, determine whether the target person has a behavior of playing with a mobile phone according to the mobile phone bounding box output by the mobile phone detection model and the person bounding boxes output by the person detection model.

10. A behavior recognition device, characterized in that, The behavior recognition device includes a memory and a processor; The memory is coupled to the processor; the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the behavior recognition device is caused to execute the mobile phone playing behavior recognition method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium storing a computer program; characterized in that, When the computer program runs on the behavior recognition device, the behavior recognition device is caused to implement the mobile phone playing behavior recognition method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for detecting a mobile phone playing behavior of a driver based on deep learning

    CN109871799A

  • Object recognition method and device based on graph embedding and medium

    CN111652286A

  • Distracted driving monitoring method, distracted driving monitoring system and electronic equipment

    CN111661059A