Target behavior detection method and device and storage medium
Patent Information
- Application Number
- CN202510359809.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明提供一种目标行为检测方法、装置及存储介质,用以解决现有技术中对目标行为进行检测具有滞后性的缺陷,实现通过人员的抓拍人脸以及监控视频对人员之间是否存在目标行为进行检测,使得可以提前获知人员之间是否存在目标行为,实现预防目标行为发生的目的
[0016]本发明还提供一种计算机可读存储介质,其上存储有计算机程序,该计算机程序被处理器执行时实现如上述任一种所述目标行为检测方法。
Smart Images

Figure CN122842004A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image detection technology, and in particular to a target behavior detection method, apparatus and storage medium. Background Technology
[0002] Targeted actions typically occur between two parties, where one party performs a targeted action against the other, who passively accepts it. Once a targeted action occurs between these parties, it may impact the person passively receiving the action; therefore, it is necessary to detect whether targeted actions exist between them.
[0003] In related technologies, the detection of target behaviors between people is generally achieved by monitoring and detecting the faces of the people involved in the target behaviors, and then performing facial segmentation and detection on the face segments to obtain information about these people.
[0004] However, the above-mentioned detection of target behavior has a lag. Summary of the Invention
[0005] This invention provides a target behavior detection method, device, and storage medium to overcome the shortcomings of existing target behavior detection methods, which have a lag. It enables the detection of target behavior between people by capturing facial images and monitoring video, allowing for early detection of target behavior and thus preventing its occurrence.
[0006] This invention provides a target behavior detection method, comprising: A first face image is acquired, and an emotion category recognition process is performed on the first person in the first face image to determine the first emotion category corresponding to the first person; the first face image includes a face image of the first person taken from the front. If the first emotion category is an emotion category related to the target behavior, then the direction of the first person's gaze and / or the direction of the first person's movement trajectory are determined based on the first facial image and / or the surveillance video associated with the first person; Based on the direction of the first person's gaze and / or the direction of the first person's movement trajectory, a set of second persons is determined from the surveillance video associated with the first person; the set of second persons includes at least one second person and a second face image associated with the second person, and the probability that the second person has a target behavior with the first person is greater than a preset probability; A first association relationship is established between the first person and the second person based on the first face image and the second face image. The first association relationship and the number of times the first association relationship appears are stored in a preset relational database. Based on the number of times the first association relationship appears in the relational database, the presence of a target behavior between the second person and the first person is detected.
[0007] According to a target behavior detection method provided by the present invention, determining the gaze direction and / or movement trajectory direction of the first person based on a first face image and / or a surveillance video associated with the first person includes: Based on the first face image, determine the pupil position and face angle of the first person in the first face image, and determine the direction of the first person's gaze based on the pupil position and face angle; Based on the surveillance video associated with the first person, at least one frame of the first surveillance image adjacent to the first face image is determined, and the movement trajectory direction of the first person is determined based on the first face image and the first surveillance image; the first surveillance image includes the first person.
[0008] According to a target behavior detection method provided by the present invention, the determination of a second set of persons from the surveillance video associated with the first person based on the gaze direction and / or movement trajectory direction of the first person includes: Obtain the second surveillance image corresponding to the first face image from the surveillance video associated with the first person, and determine the first initial set of people based on the people in the area pointed to by the gaze in the second surveillance image; the second surveillance image is the large surveillance image corresponding to the first face image; Based on the movement trajectory direction of the first person, the personnel in the area opposite to the movement trajectory direction are obtained in the second monitoring image to obtain the second initial personnel set; Determine the second set of personnel based on the first initial set of personnel and / or the second initial set of personnel.
[0009] According to a target behavior detection method provided by the present invention, determining the second set of personnel based on a first initial set of personnel and / or a second initial set of personnel includes: Take the union of the first initial set of personnel and the second initial set of personnel, and obtain the third monitoring image corresponding to the personnel in the union from the monitoring videos associated with the personnel in the union; The hand gestures of people in the third surveillance image are identified to obtain the gesture recognition results; The emotion categories of people in the third surveillance image are identified, and the emotion category identification results are obtained; The direction of people's eyes in the third surveillance image is analyzed to obtain the results of the eye direction analysis; Based on at least one of the gesture recognition results, emotion category recognition results, and eye direction analysis results, determine the second set of people from the union of the people.
[0010] According to a target behavior detection method provided by the present invention, the method for determining a first initial set of persons based on the direction of gazes in a second surveillance image towards persons in the area being pointed to includes: If the first person's gaze is directed toward the target side, then the first effective visible range corresponding to the direction toward the target side is determined in the second monitoring image; the aforementioned direction toward the target side is not toward the exact center. Determine the second effective visible range corresponding to the orientation towards the center in the second surveillance image; Based on the first and second effective visible ranges, a third effective visible range is determined, and based on the personnel in the area corresponding to the third effective visible range, a first initial personnel set is determined.
[0011] According to a target behavior detection method provided by the present invention, the method further includes: Retrieve the second association between the first person and the third person from the relational database, as well as the frequency of occurrence of the second association marker; the third person is someone whose probability of having a target behavior with the first person is greater than a preset probability; Based on the first association, the number of times the first association marker appears, the second association, and the number of times the second association marker appears, generate the relationship graphs of the first person, the second person, and the third person.
[0012] According to a target behavior detection method provided by the present invention, the method further includes: Obtain the trajectory of personnel within a historical time period; the trajectory of personnel within the aforementioned historical time period includes multiple fourth parties; Peer analysis is performed on the trajectory of personnel within a historical time period to identify at least one peer group for multiple fourth persons, and the same group label is assigned to multiple peer fourth persons in each peer group; the aforementioned peer group includes multiple peer fourth persons. If any target fourth person in the peer group has a relationship with the first person, then all fourth people in the corresponding peer group are obtained based on the group tag of the target fourth person, and the relationship graph of the first person and the relationship graph of each of the fourth people are displayed.
[0013] According to a target behavior detection method provided by the present invention, before performing emotion category recognition processing on a first person in a first face image to determine the first emotion category corresponding to the first person, the method further includes: Obtain a pre-defined personnel database; the personnel database includes facial images of at least one candidate person, and the personnel database is obtained by performing emotion category recognition and individual analysis on personnel in historical surveillance videos; The first face image is matched with the face images of candidates in the personnel database. If the match is successful, the process returns to the above steps of performing emotion category recognition processing on the first person in the first face image to determine the first emotion category corresponding to the first person.
[0014] The present invention also provides a target behavior detection device, comprising the following modules: An emotion recognition module is used to acquire a first face image and perform emotion category recognition processing on the first person in the first face image to determine the first emotion category corresponding to the first person; the first face image includes a face image of the first person taken from the front. The eye gaze and trajectory determination module is used to determine the eye gaze and / or movement trajectory direction of the first person based on the first facial image and / or the surveillance video associated with the first person if the first emotion category is an emotion category related to the target behavior. The personnel set determination module is used to determine a second personnel set in the surveillance video associated with the first person based on the direction of the first person's gaze and / or the direction of the first person's movement trajectory; the second personnel set includes at least one second person and a second face image associated with the second person, and the probability that the second person has a target behavior with the first person is greater than a preset probability; The target behavior detection module is used to establish a first association relationship between a first person and a second person based on a first face image and a second face image, and to store the first association relationship and the number of times the first association relationship appears in a preset relational database, and to detect whether there is a target behavior between the second person and the first person based on the number of times the first association relationship mark appears in the relational database.
[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the target behavior detection method as described above.
[0016] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the target behavior detection method as described above.
[0017] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the target behavior detection method as described above.
[0018] The target behavior detection method, apparatus, and storage medium provided by this invention acquire a first facial image of a first person taken from the front, and perform emotion recognition processing on the first facial image to determine a first emotion category corresponding to the first person. If the first emotion category is an emotion category related to target behavior, the orientation of the first person's gaze and / or the direction of the first person's movement trajectory are determined based on the first facial image and / or the surveillance video associated with the first person. Then, based on the orientation of the first person's gaze and / or the direction of the first person's movement trajectory, a second set of people, including at least one second person machine-associated second facial image, is determined in the surveillance video associated with the first person. A first association relationship is established between the first person and the second person based on the first facial image and the second facial image, and the occurrence frequency of the first association relationship is stored and marked in a relational database. Then, the existence of target behavior between the second person and the first person is detected based on the occurrence frequency of the first association relationship marking in the relational database, wherein the probability of target behavior between the second person and the first person is greater than a preset probability. In this method, the emotion category, eye direction, and movement trajectory direction of the first person's facial image can be identified and analyzed in advance before the target behavior occurs between the second and first persons. This allows for early detection of whether there is a target behavior between the second and first persons, rather than detecting it at the moment the target behavior occurs. This avoids detection lag and prevents the target behavior from going undetected, while also improving the accuracy of target behavior detection. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0020] Figure 1 This is one of the flowcharts of the target behavior detection method provided in the embodiments of the present invention.
[0021] Figure 2 This is the second flowchart of the target behavior detection method provided in the embodiments of the present invention.
[0022] Figure 3 This is a schematic diagram of different pupil positions in the human eye provided in an embodiment of the present invention.
[0023] Figure 4 This is a schematic diagram of the effective visual range of the pupil when the human eye looks straight ahead, provided by an embodiment of the present invention.
[0024] Figure 5 This is a schematic diagram of the effective visual range of the pupil when the human eye looks to the left front, provided by an embodiment of the present invention.
[0025] Figure 6 This is a schematic diagram of determining a second initial set of personnel based on the movement trajectory direction of a first person, provided by an embodiment of the present invention.
[0026] Figure 7 This is the third flowchart of the target behavior detection method provided in the embodiments of the present invention.
[0027] Figure 8 This is a schematic diagram of the relationship map of the first person provided in an embodiment of the present invention.
[0028] Figure 9 This is a flowchart illustrating the process of performing emotion category analysis on individuals, provided in an embodiment of the present invention.
[0029] Figure 10 This is a flowchart illustrating the process of performing individual analysis on personnel, provided in an embodiment of the present invention.
[0030] Figure 11 This is a schematic diagram of the target behavior detection device provided in an embodiment of the present invention.
[0031] Figure 12 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0033] Current target behavior detection methods generally detect the target behavior at the moment it occurs, thus exhibiting a lag. This invention provides a target behavior detection method, apparatus, and storage medium that can solve this technical problem.
[0034] The following describes the technical terms used in the embodiments of the present invention: Trajectory: Captures included in the archive, such as facial captures included in the facial archive.
[0035] Real-name profile: Taking a face as an example, the capture of the same face forms a face profile for that person. If it is linked to the ID card database, it will include relevant information such as the ID card number, which can be called a real-name profile.
[0036] The following is combined with Figures 1-10 This invention describes a target behavior detection method according to an embodiment of the present invention.
[0037] It should be noted that the execution subject in the embodiments of the present invention can be a target behavior detection device, an electronic device including the target behavior detection device, or other devices or apparatuses. The following embodiments will use an electronic device as an example for illustration. The electronic device can be a terminal or a server, and is not specifically limited thereto.
[0038] Figure 1 This is one of the flowcharts illustrating the target behavior detection method provided in this embodiment of the invention, such as... Figure 1 As shown, the method includes the following steps: Step 102: Obtain the first face image and perform emotion category recognition processing on the first person in the first face image to determine the first emotion category corresponding to the first person; the first face image includes a face image of the first person taken from the front.
[0039] In this context, "capturing" refers to the process by which surveillance cameras identify and capture images of people, bodies, motor vehicles, non-motorized vehicles, and other objects in a video feed. Surveillance cameras within a specific area can monitor and record personnel in that area in real time, performing face capture processing to obtain facial images and video footage of individuals at different times. A captured face image typically contains only one person's face, a small image of the face. The video footage can include surveillance videos captured at various times, comprising multiple frames. Each frame can include full-body or partial images of multiple different individuals. The specific area can be the area covered by the surveillance cameras, such as a school campus. There can be one or more surveillance cameras.
[0040] In this step, when capturing real-time facial images of people appearing in a specific area, the facial image of each person at a certain capture moment can be obtained. Here, we take one person as an example. Let's assume this person is the first person. The facial image of the first person captured at the current moment may be a frontal facial image or a side facial image. If the captured facial image of the first person is a frontal view, it can be directly used as the first person's first facial image. If the captured image is a side view, it may be difficult to identify the person's emotions. In this case, it's necessary to first determine if the person's face is tilted down (a side view makes it easier to identify tilting the head). If so, the surveillance video associated with the person (i.e., the large surveillance image captured simultaneously with the facial image) is used to find the image of the person before they tilted their head down. This image is then used to extract the person's face before they tilted their head down, which is usually a frontal view and can be used as the first person's first facial image. In short, the first facial image described above is a frontal view of the person, which facilitates accurate identification of the person's emotion category.
[0041] After obtaining the first facial image of the first person, the emotion category of the first person in the facial image can be directly identified to obtain the current emotion category of the first person, which is recorded as the first emotion category. The first emotion category can include categories such as happy, natural, depressed, fearful, and angry. Among them, from the perspective of the first person, fear and depressed may be emotion categories related to the target behavior.
[0042] Alternatively, after obtaining the facial image of the first person captured at the current moment, or after obtaining the first facial image of the first person, it can be determined whether the first person is suspected of engaging in the target behavior. If so, the emotion category of the first person in the first facial image is identified to obtain the current first emotion category of the first person. If not, the emotion category identification process is not performed on the first person in the first facial image, thereby reducing the amount of computation and improving the detection accuracy.
[0043] In addition, when identifying the emotion category of a person in a facial image, a pre-trained large model can be used for identification or emotion analysis using a large model. The specific architecture and type of the large model are not specifically limited here.
[0044] Step 104: If the first emotion category is an emotion category related to the target behavior, then determine the direction of the first person's gaze and / or the direction of the first person's movement trajectory based on the first face image and / or the surveillance video associated with the first person.
[0045] In this step, after obtaining the first person's current first emotion category, it can be determined whether the first person's current first emotion category is an emotion category related to the target behavior, such as whether it is depression or fear. If so, it is determined that the first person may / is suspected of having the target behavior and further analysis is required.
[0046] If the first person's current first emotion category is an emotion category related to the target behavior, the direction of the first person's eyes in the first face image can be analyzed to obtain the current direction of the first person's eyes. The direction of the eyes can include eyes facing the center or eyes facing a side other than the center, such as facing the left or right.
[0047] Simultaneously, the surveillance video (i.e., the large surveillance image) associated with the first face image can be obtained. Then, by combining the surveillance video and the first face image, the movement trajectory direction of the first person in the first face image can be identified to obtain the movement trajectory direction of the first person, such as moving towards the left front or the right front.
[0048] Step 106: Based on the direction of the first person's gaze and / or the direction of the first person's movement trajectory, determine the second set of persons in the surveillance video associated with the first person; the second set of persons includes at least one second person and a second face image associated with the second person, and the probability that the second person has a target behavior with the first person is greater than a preset probability.
[0049] In this step, after obtaining the first person's current gaze direction and movement trajectory direction, one or more surveillance images associated with the first person, taken at the same or adjacent times as the first person's face image, can be found. Then, the person in the area pointed to by the first person's gaze can be located within these images. Alternatively, one or more surveillance images associated with the first person, taken at the same or adjacent times as the first person's face image, can be found within these images. Then, the person in the area pointed to by the first person's movement trajectory direction, or in an area opposite or on the opposite side of that movement trajectory direction, can be located within these images.
[0050] Then, the person found by the direction of the first person's gaze and the person found by the direction of their movement can both be considered as the second person. Alternatively, the person found by the direction of the gaze and the person found by the direction of their movement can be considered as the second person. Or, a subset of the people found by the direction of the gaze and the person found by the direction of their movement can be considered as the second person. Or, other methods can be used. In short, as long as the second person can be found, that is fine.
[0051] When identifying the second person, facial masking can be performed simultaneously on the surveillance image containing the second person to obtain their corresponding facial image, which is then recorded as the second facial image (small face image). This second person and their second facial image can then be placed into a set, forming a second person set. This second person set can include one or more second persons and their second facial images. Furthermore, the facial image of the second person can be associated with their identification database to include their identification number and other information, forming a real-name profile for the second person. Similarly, the facial image of the first person can be associated with their identification database to include their identification number and other information, forming a real-name profile for the first person. The real-name profile for each person should include their facial image, body data, and image.
[0052] It is understandable that the individuals identified through the first person's gaze and movement trajectory are all closely related to the first person's current behavior. The probability that these individuals (i.e., the second person) and the first person have a target-related interaction is greater than the preset probability; that is, these individuals are very likely to have a target-related interaction with the first person. The preset probability can be a value set according to the actual situation, such as 70%, 80%, etc.
[0053] Step 108: Establish a first association relationship between the first person and the second person based on the first face image and the second face image, store the first association relationship and the number of times the first association relationship appears in the preset relational database, and detect whether there is a target behavior between the second person and the first person based on the number of times the first association relationship mark appears in the relational database.
[0054] In this step, after obtaining the first facial image of the first person and the second facial image of the second person, the human body image of the first person can be obtained by cutting out the human body from the associated surveillance video using the first facial image. Then, the first person's facial image or human body image can be used to search the real-name file to obtain the specific information of the first person (such as name, ID number, etc.). Similarly, the human body image of the second person can be obtained by cutting out the human body from the associated surveillance video using the second facial image. Then, the second person's facial image or human body image can be used to search the real-name file to obtain the specific information of the second person. If the image of a person is a back view, the person can also be found by searching for their back.
[0055] Then, the specific personnel information of the first person and the second person can be bound together to obtain the association relationship between them. This association refers to the possible / suspected association between the first person and the second person regarding the target behavior, and is denoted as the first association relationship. The roles of the first person and the second person can be marked within this association relationship: the first person's role is the affected person, and the second person's role is the person exhibiting the target behavior. After establishing the first association relationship, it can be checked in a preset relational database to see if it exists. If it does not exist, the first association relationship is stored in the relational database, and its occurrence count is marked as 1. Subsequent occurrences of the first association relationship will continue to accumulate from the previously marked occurrence count. If the first association relationship exists in the relational database, the current occurrence count of the first association relationship can be obtained from the relational database, and then the current occurrence count of the first association relationship can be incremented by one. For example, when the first association between a first person and a second person in the same group is established and stored for the first time, the occurrence count of the marker is 1. If, after executing steps 102-108 again, the same first association between the first person and the second person in the same group is established again, then the occurrence count of the marker for this first association in the relational database is incremented by 1, becoming 2. Furthermore, when storing the first association and its occurrence count in the relational database, the time when the first association was first established and stored can also be stored, so that the time of its initial establishment can be used to determine whether the first association needs to be deleted. The relational database can store various associations between people and the occurrence count of each association.
[0056] After continuously performing steps 102-108 above on a specific area over a period of time, the occurrence count of the current marker for the first association can be obtained. This occurrence count can then be used to detect whether a target behavior truly exists between the second and first individuals. For example, if the occurrence count of the first association marker in the relational database reaches a set number (e.g., 5 times, meaning the same association appears 5 times within a certain period), then a target behavior is determined to exist between the second and first individuals. If the occurrence count does not reach the set number, then no target behavior exists between the second and first individuals. Alternatively, for example, if the occurrence count of the first association marker is greater than 1, then a target behavior is determined to exist between the second and first individuals, and the priority of the target behavior between the first and second individuals is determined by the occurrence count of the marker; for example, the higher the occurrence count, the higher the priority of the target behavior.
[0057] Furthermore, after detecting and obtaining the occurrence count of the current marker for the first association over a period of time, if the occurrence count reaches a set number, the priority of the target behavior between the first person and the second person is marked as the highest. The first association, along with the specific personnel information of the first and second persons, is then pushed to relevant superiors (such as the homeroom teacher or other teachers) via alerts to remind them to pay closer attention to the first and second persons and prevent subsequent target behaviors from occurring. Alternatively, if the occurrence count does not reach the set number, the priority of the target behavior between the first person and the second person is marked as a lower priority, and no alerts are pushed to relevant superiors; only a record is made.
[0058] In addition, relevant superiors can view the first established relationship and the number of times its tag appears at any time. Within a certain period after the first relationship is established (e.g., 1 month), the current number of times the tag of the first relationship appears can be detected. If the number of times the tag appears does not reach the set number within this period, the first relationship can be deleted to avoid misleading relevant superiors.
[0059] Furthermore, the aforementioned target behavior can be a target behavior that may occur or exist between people in a specific area, or it can be a target behavior that may occur or exist in an area outside the specific area. The specific area here can be, for example, the area corresponding to a school. The target behavior can be, for example, a controlled or affected behavior between people, such as a second person seriously affecting the first person's work or study, or a second person controlling the first person.
[0060] In this embodiment, a first facial image of a first person is acquired by taking a frontal shot, and emotion recognition processing is performed on the first facial image to determine the first emotion category corresponding to the first person. If the first emotion category is an emotion category related to the target behavior, the direction of the first person's gaze and / or the direction of the first person's movement trajectory are determined based on the first facial image and / or the surveillance video associated with the first person. Then, based on the direction of the first person's gaze and / or the direction of the first person's movement trajectory, a set of second persons including at least one second person machine-associated second facial image is determined in the surveillance video associated with the first person. A first association relationship between the first person and the second person is established based on the first facial image and the second facial image, and the occurrence frequency of the first association relationship is stored and marked in the relational database. Then, the occurrence frequency of the first association relationship mark in the relational database is used to detect whether there is a target behavior between the second person and the first person, wherein the probability that there is a target behavior between the second person and the first person is greater than a preset probability. In this method, the emotion category, eye direction, and movement trajectory direction of the first person's facial image can be identified and analyzed in advance before the target behavior occurs between the second and first persons. This allows for early detection of whether there is a target behavior between the second and first persons, rather than detecting it at the moment the target behavior occurs. This avoids detection lag and prevents the target behavior from going undetected, while also improving the accuracy of target behavior detection.
[0061] The following examples illustrate the process of determining the direction of a person's gaze and movement trajectory, and on this basis, determining the second group of people.
[0062] Figure 2 This is a second schematic flowchart of the target behavior detection method provided in this embodiment of the invention, as shown below. Figure 2 As shown, step 104 above, which determines the direction of the first person's gaze and / or the direction of the first person's movement trajectory based on the first face image and / or the surveillance video associated with the first person, may include the following steps: Step 202: Based on the first face image, determine the pupil position and face angle of the first person in the first face image, and determine the direction of the first person's gaze based on the pupil position and face angle.
[0063] In this step, the process of determining the facial angle of the first person in the first facial image may include: after obtaining the first facial image, detecting each point in the first facial image to obtain the position of each point, and finding the point in the middle of the forehead or brow and the point in the middle of the chin. Then, a straight line can be obtained by connecting the point in the middle of the forehead or brow and the point in the middle of the chin, and then compared with the straight line of a standard frontal face to calculate the angle between the two, thereby obtaining the facial angle of the first person in the first facial image.
[0064] The process of determining the pupil position of the first person in the first facial image may include: see Figure 3 The diagram shows different pupil positions in the human eye. Generally, there are three pupil positions: looking straight ahead, looking to the left front, and looking to the right front. After obtaining the first face image, the region where any eye of the first person is located can be identified. After identifying the region where the eye is located, the position of the pupil in the eye (such as the circular area in the figure) can be identified. Then, the relative position of the pupil in the eye can be obtained by using the identified pupil position. This relative position is recorded as the final pupil position. The pupil position can be, for example, in the middle of the eye, on the left side of the eye, or on the right side of the eye.
[0065] After obtaining the facial angle and pupil position of the first person, the direction of their gaze can be determined by combining these two factors. For example, if both the facial angle and pupil position indicate that the first person is looking to the left, then their gaze is directed to the left front. Determining the first person's gaze direction using both pupil position and facial angle is more accurate.
[0066] Step 204: Based on the surveillance video associated with the first person, determine at least one frame of the first surveillance image adjacent to the first face image, and determine the movement trajectory direction of the first person based on the first face image and the first surveillance image; the first surveillance image includes the first person.
[0067] In this step, after obtaining the first person's facial image and the associated surveillance video, which includes multiple frames of surveillance images captured at various times, we first locate the frame containing / corresponding to the first facial image. Then, we find one or more frames adjacent to the captured time of that first frame, all of which are recorded as the first surveillance image, and each includes the first person. Next, we perform position detection processing on both the frame containing / corresponding to the first facial image and the first surveillance images to obtain the first person's position within each of these frames. Then, we use these positions to form a trajectory to determine the direction of the first person's movement. For example, if the first person is in the middle of the frame at one moment and on the right side at another, it indicates that the first person's movement is to the right front or right rear.
[0068] After determining the current gaze direction and movement trajectory direction of the first person, the second group of people can be determined accordingly. Optionally, step 106 above, which determines the second group of people from the surveillance video associated with the first person based on the gaze direction and / or movement trajectory direction of the first person, can include the following scenarios: Scenario 1: If the first face image is a frontal face image of the first person before they lower their head, obtained from a side-view image directly captured from the first person, and the first emotion category of the first person is identified as an emotion category related to the target behavior based on this frontal face image before they lower their head, then the second set of people can be determined from the surveillance video associated with the first person based on the direction of the first person's gaze and the direction of their movement trajectory. Specifically, this may include the following steps: Step A1: Obtain the second surveillance image corresponding to the first face image from the surveillance video associated with the first person, and determine the first initial set of people based on the people in the area pointed to by the gaze in the second surveillance image; the second surveillance image is the large surveillance image corresponding to the first face image.
[0069] Step A2: Based on the movement trajectory direction of the first person, obtain the personnel in the area on the opposite side of the movement trajectory direction in the second monitoring image to obtain the second initial personnel set.
[0070] Step A3: Determine the second set of personnel based on the first initial set of personnel and / or the second initial set of personnel.
[0071] After obtaining the first facial image of the first person and the associated surveillance video, which includes multiple frames captured at various times, the process involves first identifying the frame containing / corresponding to the first facial image and designating it as the second surveillance image. Then, the area pointed to by the first person's gaze is located within the second surveillance image, and facial masking and detection are performed on the individuals within that area to obtain their facial images and specific information. Finally, these individuals and their facial images are used to determine the initial set of personnel.
[0072] As mentioned above, the direction of a person's gaze includes looking directly to the center, to the left front, and to the right front. When observing, the human eye corresponds to a visual range, which refers to the total area that the eye can clearly see while moving. The eye also has an effective visual range for each gaze direction, which is smaller than the aforementioned visual range but falls within it. For an example of looking directly forward, see [link to relevant documentation]. Figure 4 The diagram illustrates the effective visual range of the pupil when the human eye is looking straight ahead. The effective visual range refers to the area that can be clearly seen without eye movement. This effective visual range serves as the perceived range of the area the gaze is directed towards. For example, when looking directly to the center, within the length of the eye, the dotted line in the diagram represents the effective visual range, while the solid line represents the total visual range even when the eye is moving.
[0073] Optionally, if the first person's gaze is directed toward the center, then the people and their facial images in the effective visible area corresponding to the direction of the gaze toward the center are directly used as the first initial set of people.
[0074] Optionally, if the first person's gaze is not directed towards the center, that is, if the first person's gaze is directed towards the target side, and the target side is not directed towards the center, it may be directed towards the left front or the right front of the target side, then the first effective visible range corresponding to the direction towards the target side is determined in the second monitoring image, that is, the effective visible range area corresponding to the direction towards the target side is determined in the second monitoring image; at the same time, the second effective visible range corresponding to the direction towards the center can be determined in the second monitoring image, that is, the effective visible range area corresponding to the direction towards the center is determined in the second monitoring image; then, based on the first and second effective visible ranges, a third effective visible range is determined, and based on the people in the area corresponding to the third effective visible range, the first initial set of people is determined. For example, the union of the first and second effective visible ranges (i.e., the third effective visible range) can be taken, and the people and their facial images in the corresponding area of the union in the second monitoring image can be used as the first initial set of people; or, the intersection of the first and second effective visible ranges can be taken, and the people and their facial images in the corresponding area of the intersection in the second monitoring image can be used as the first initial set of people.
[0075] For example, taking looking to the left front as an example (looking to the right front is similar), see [link to example]. Figure 5 The diagram illustrates the effective visual range of the pupil when a person is looking to the left front. In this case, the pupil also has an effective visual range. If we use this effective visual range to select people within it as the first initial set of people, the number of people selected will be too large. To narrow down the range, we can also obtain the effective visual range when looking directly forward, and then superimpose the effective visual range when looking to the left front / left (i.e., the area corresponding to the solid line in the diagram) and the effective visual range when looking directly forward (i.e., the area corresponding to the dashed line in the diagram). After removing the overlapping areas, the remaining effective visual range when looking to the left front (i.e., the filled area in the diagram) is the range for people whose gaze is directed to the left front. Finally, we select the people within this range and their facial images as the first initial set of people.
[0076] While determining the first initial set of people, a second initial set of people can also be determined based on the movement trajectory direction of the first people. This can be achieved by first acquiring the set of all people on the horizontal plane directly in front of the current capture image, then finding the area on the side opposite to the movement trajectory direction in the second monitoring image, and performing face image segmentation and recognition processing on the faces of the people in that area to obtain the people and their faces in that area, which serves as the second initial set of people. For example, see [link to example]. Figure 6The diagram shown illustrates the determination of a second initial set of personnel based on the movement trajectory direction of a first person. For example, if the movement trajectory direction of the first person is towards the right front, the opposite side can be towards the left front. For instance, the area corresponding to the dashed box in the diagram is the area towards the left front. The personnel and their facial images identified in this area can be used as the second initial set of personnel.
[0077] After obtaining the first initial set of personnel and the second initial set of personnel, the final second set of personnel can be determined using the first initial set of personnel and the second initial set of personnel. This process can be implemented using any of the following methods: Method 1: Take the intersection of the people in the first initial personnel set and the second initial personnel set, and use the people and their facial images corresponding to the intersection as the second personnel set.
[0078] Method 2: Take the union of the people in the first initial set of people and the second initial set of people, and use the people and their face images corresponding to the union as the second set of people.
[0079] Method 3: Take the union of the first initial set of people and the second initial set of people, and obtain the third monitoring image corresponding to the people in the union from the monitoring videos associated with the people in the union; identify the gestures of the people in the third monitoring image to obtain gesture recognition results; identify the emotion categories of the people in the third monitoring image to obtain emotion category recognition results; analyze the eye direction of the people in the third monitoring image to obtain eye direction analysis results; determine the second set of people from the people in the union based on at least one of the gesture recognition results, emotion category recognition results, and eye direction analysis results.
[0080] Specifically, in method three, the union of the first and second initial personnel sets can be taken, and the corresponding large-scale surveillance images of the personnel in the union set can be found in the associated surveillance videos. Then, human body cutout is performed on the personnel in the large-scale surveillance images to obtain human body images of the personnel in the union set. Next, gesture recognition processing is performed on the human body images of the personnel in the union set to determine whether these personnel have abnormal gesture behaviors (such as pointing downwards at the first person). Simultaneously, the large-scale model can be used to identify the emotion categories of these personnel in the large-scale surveillance images and analyze their eye direction to determine whether these personnel have emotion categories related to the target behavior and / or whether their eye direction is towards the first person (the eye direction analysis method can adopt the same eye direction analysis method as in the previous embodiments). If these personnel have abnormal gesture behaviors, and / or if these personnel have emotion categories related to the target behavior, and / or if these personnel's eye direction is towards the first person, then these personnel and their facial images are included as the second personnel set to improve the accuracy of determining the two parties involved in the target behavior, thereby improving the accuracy of target behavior detection. From the perspective of the second personnel, happiness and anger are likely emotion categories related to the target behavior.
[0081] Scenario 2: If the first facial image is a frontal image directly captured from the first person, and the first emotion category of the first person is identified as an emotion category related to the target behavior, the second set of people can be determined from the surveillance video associated with the first person solely based on the direction of the first person's gaze. Specifically, this can include the following steps: Obtain the second surveillance image corresponding to the first face image from the surveillance video associated with the first person, and determine the first initial set of people based on the people in the area pointed to by the gaze in the second surveillance image; the second surveillance image is the large surveillance image corresponding to the first face image; Use the first initial set of people as the second set of people.
[0082] The specific process of determining the initial group of personnel by observing the direction of the first person's gaze can be the same as in Scenario 1 above. For details, please refer to the explanation in Scenario 1 above, and it will not be repeated here.
[0083] In this embodiment, the direction of a person's gaze is determined by the position of their pupils and the angle of their face, thus improving the accuracy of the determined gaze direction. Simultaneously, the direction of the person's movement trajectory is determined by using at least one large monitoring frame adjacent to the first face image, ensuring the accuracy of the determined movement trajectory direction. Furthermore, the final second set of people is determined by combining the set of people whose gaze direction is the same as the set of people whose movement trajectory direction is opposite to the first person's. This avoids missing second people and improves the accuracy of the final determined second set of people. Further, when the gaze direction is not exactly in the center, the range of the gaze direction towards the center and the range towards the target side can be used to determine the acquisition range of the second person, thus narrowing the acquisition range and improving the efficiency of determining the second set of people.
[0084] The following examples illustrate the process of generating a relationship graph of the relationships between a first person and multiple other people.
[0085] Figure 7 This is the third flowchart of the target behavior detection method provided in this embodiment of the invention, as shown below. Figure 7 As shown, the above method may also include the following steps: Step 302: Obtain the second association relationship between the first person and the third person and the number of occurrences of the second association relationship marker in the relational database; the third person is a person whose probability of having a target behavior with the first person is greater than a preset probability.
[0086] In the process of detecting whether there is a target behavior between the first person and the second person, the detection of whether there is a target behavior between the first person and the third person is also performed simultaneously. The third person and the second person are different people detected at different times. After the third person is detected, a second association relationship can be established between the first person and the third person in the manner described above, and the second association relationship can be stored in the relational database. The second association relationship indicates that there is a suspected target behavior between the first person and the third person. At the same time, the role information of the first person and the third person can be marked in the relational database.
[0087] At the same time, the occurrence frequency of the second association can be marked in the relational database, and the occurrence frequency of the second association mark can be obtained at a set time.
[0088] It is understandable that the aforementioned third party is someone who is closely related to the first party's current behavior. The probability that these third parties and the first party have a target behavior is greater than the preset probability, that is, these third parties are very likely to have a target behavior with the first party.
[0089] Step 304: Generate the relationship graph of the first person, the relationship graph of the second person, and the relationship graph of the third person based on the first relationship, the number of occurrences of the first relationship marker, the second relationship, and the number of occurrences of the second relationship marker.
[0090] In this step, during the actual detection process, there may be multiple third-party relationships between different third parties and the first party. That is, the above method can establish all the relationships between the first party and other parties and obtain the occurrence number of each relationship marker. Then, taking the first party as the center, the parties with relationships with it (first relationship and second relationship) can be connected with the first party to generate the relationship graph of the first party.
[0091] For example, see Figure 8 The diagram shown illustrates the relationship graph of the first person, illustrating the first relationship between the first person and the second person, the second relationship between the first person and a third person, and the frequency of each relationship marker (e.g., occurrence 1, occurrence 2). Furthermore, this relationship graph can also indicate the priority or severity of each relationship of the first person, as well as the role information of the first person and other individuals. This relationship graph allows for a clear and quick identification of all individuals suspected of engaging in target behavior with the first person, improving the efficiency and accuracy of target behavior detection.
[0092] The relationship graphs for the second and third personnel can also be generated in the same way as described above, so we will not go into details here.
[0093] In addition, when the number of occurrences of a certain relationship marker of the first person reaches a preset number and an alarm is pushed to the relevant superior, the relevant superior can also obtain other relationships of the first person through the relationship graph of the first person, but the alarm push has a lower priority.
[0094] Furthermore, after generating the relationship graph for each person, the relationship graph can be checked periodically. If the number of times a certain relationship is marked within a set time (e.g., 30 days) does not reach the set number (e.g., 5 times), the relationship can be deleted and will no longer be displayed in the relationship graph to avoid misleading relevant superiors.
[0095] In this embodiment, by obtaining the association between the first person and other persons and the frequency of occurrence of the markers, and generating the corresponding relationship graph of the first person and the relationship graph of other persons based on all the associations of the first person, it is possible for relevant superiors to clearly and quickly identify all persons who are suspected of having target behavior with the first person through the relationship graph, thereby improving the efficiency and accuracy of target behavior detection.
[0096] In actual testing, there may be situations where a group of multiple people engages in target behavior with the first person. The following examples illustrate this situation.
[0097] In some embodiments, the above method may further include the following steps: Obtain the trajectory of personnel within a historical time period; the trajectory of personnel within the aforementioned historical time period includes multiple fourth parties; Peer analysis is performed on the trajectory of personnel within a historical time period to identify at least one peer group for multiple fourth persons, and the same group label is assigned to multiple peer fourth persons in each peer group; the aforementioned peer group includes multiple peer fourth persons. If any target fourth person in the peer group has a relationship with the first person, then all fourth people in the corresponding peer group are obtained based on the group tag of the target fourth person, and the relationship graph of the first person and the relationship graph of each of the fourth people are displayed.
[0098] Among them, it can obtain the trajectory of personnel from each surveillance camera in a specific area within a preset historical time period (such as 30 days). The trajectory of personnel can be the trajectory of personnel's real-name files. The cameras are grouped according to different surveillance cameras and sorted in ascending order according to the shooting time. Then, peer analysis is performed starting from the first face image data in each group.
[0099] The specific peer analysis process includes: starting with the first facial image data (hereinafter referred to as "data"), determining whether other people appear in multiple data sets within a certain time period. If other people appear, these people are considered peers, and their peer relationships are temporarily marked. That is, each group of peers is formed into a separate set, obtaining at least one initial set. Then, the number of times the same group of peers in all the obtained initial sets is accumulated to obtain the peer count. When the peer count exceeds a set value, this group of peers is considered a true peer, and this group is formed into a peer set. The same group label / gang label is set for the real-name profile of each person in the peer set; different peer sets can have different group labels; each peer set contains at least two people. For example, accumulating the peer counts of person A and person B yields the same peer count for both A and B. Only those with a peer count > 5 are considered true peers, and their peer set is obtained.
[0100] During the target behavior detection process, if a relationship is detected between any fourth person in the peer group and the first person, the peer group can be marked as a group / clique with target behavior. Simultaneously, when an alert is sent to relevant superiors regarding the potential target behavior between the fourth person and the first person, all fourth persons in the peer group can be found through the fourth person's group tag. The specific personnel information and relationship graphs (which can be generated using the methods described in the above embodiments) of these fourth persons, along with the relationship graph of the first person, are then sent to relevant superiors for review or investigation to ensure comprehensive detection.
[0101] In this embodiment, by analyzing the historical trajectory of a person and finding the set of people in the same group, once any person in the set of people ...
[0102] When actually detecting whether there is a target behavior among various personnel, the large number of personnel may result in a large amount of detection data and low detection accuracy. Based on this, the present invention provides a technical solution to solve this problem, and the following embodiments will illustrate this.
[0103] In some embodiments, before performing emotion category recognition processing on the first person in the first face image in step 102 above to determine the first emotion category corresponding to the first person, the above method may further include the following steps: Obtain a pre-defined personnel database; the personnel database includes facial images of at least one candidate person, and the personnel database is obtained by performing emotion category recognition and individual analysis on personnel in historical surveillance videos; The first face image is matched with the face images of candidates in the personnel database. If the match is successful, the process returns to the above steps of performing emotion category recognition processing on the first person in the first face image to determine the first emotion category corresponding to the first person.
[0104] This involves pre-establishing a personnel database, which typically includes individuals potentially influenced by the target behavior. This database can be obtained by performing emotion category identification / analysis and solitary behavior analysis on individuals captured in surveillance videos of a specific area. The processes of emotion category analysis and solitary behavior analysis are explained below.
[0105] Emotion category analysis: See Figure 9 The flowchart shown illustrates the process of analyzing the emotion categories of individuals. It involves accessing facial capture data from surveillance cameras in a specific area, acquiring each facial image captured by the cameras, and using a large-scale model algorithm to identify the emotion category of each person in the facial image. If the emotion category is related to the target behavior (such as depression or fear), the facial image of that emotion category is associated with the individual's real-name profile to obtain their specific information. These individuals are suspected of being affected by the target behavior, and those whose emotion category was detected can be directly added to the personnel database.
[0106] Alternatively, to more accurately identify individuals potentially affected by targeted behavior and build a comprehensive personnel database, each person captured by surveillance cameras can be analyzed individually. For a specific individual, emotion category recognition can be performed on each frame of their facial image data in real time to obtain the emotion category (i.e., emotion field) for each frame. This individual's facial image data and the emotion category of each frame are then stored in a big data database. Subsequently, a clustering algorithm (specifically, selecting a number of facial images from the facial image data to form an individual's facial profile and trajectory) is used to analyze the facial image data, creating the individual's facial profile and trajectory. Simultaneously, this is correlated with the individual's identification documents in a document database to obtain their real-name profile. Since the trajectory is formed from the facial image data, it also contains an emotion field corresponding to the emotion category.
[0107] Next, we can obtain the trajectories from facial profiles within a certain analysis range (e.g., a specific time period or a specific area). We process these trajectories separately for each face, calculating the number of times each person experiences different emotions each day, and calculating the percentage of each person's daily emotions that are related to the target behavior (e.g., depression and fear). The formula is: Percentage = (Number of trajectories with depression + Number of trajectories with fear) / Total number of trajectories. When a person's percentage is greater than 50%, they are temporarily labeled. We can then continue calculating this percentage over several consecutive days. If a person is temporarily labeled for a set number of consecutive days (e.g., 3 consecutive days), they are identified as someone whose emotion category is related to the target behavior, specifically someone experiencing depression or fear.
[0108] The above analysis can identify all individuals whose emotional categories are related to the target behavior, and these individuals can be considered as potential targets.
[0109] The following section explains the analysis of solo play: See [link / reference] Figure 10 The flowchart shown illustrates the process of performing individual analysis on individuals. It involves accessing facial capture data from surveillance cameras in a specific area and storing it in big data. The facial capture data is then analyzed using a clustering algorithm to create facial profiles and trajectories of individuals within the facial capture data. Simultaneously, the individual's identity document is linked to the document database to obtain their real-name profile.
[0110] Next, we can acquire all face capture data within a certain analysis range (e.g., a certain time period, such as 7 days of history), group them according to the surveillance camera, and then sort them in ascending order by capture time within each group to obtain the sorted face capture data (hereinafter referred to as data). Then, we can process each data in each group one by one, starting from the first data. We can determine whether the capture time interval between the data and the previous and next data is greater than a time limit (e.g., 30 seconds). If the capture time interval between the data and the previous data is greater than the time limit and the capture time interval between the data and the next data is also greater than the time limit, then the data is considered to be a single capture.
[0111] If the time interval between the capture of this data and the previous data is less than or equal to the duration threshold, or the time interval between the capture of this data and the next data is less than or equal to the duration threshold, it is associated with the trajectory data of the real-name profile to query and confirm whether the person in the previous or next data belongs to the same person's profile. If they do not belong to the same person, it is determined that it is not a solo capture and no processing is required for this data. If they belong to the same person, the process continues to search the data before the previous data and the data after the next data to confirm whether the person in those data belongs to the same person as the person in this data, until all data within the duration threshold has been checked. If a capture containing only the same person is found within the duration threshold, then the current data is also a solo capture. If no capture containing only the same person is found within the duration threshold, it is determined that it is not a solo capture and no processing is required for this data. Afterwards, all the acquired solo captures can be associated with the trajectory data of the real-name profile again to obtain the solo trajectory corresponding to the solo capture (i.e., to obtain the solo face profile).
[0112] Next, the percentage of each person's daily solo travel can be calculated using the formula: Percentage = Number of solo travels (or number of solo travel tracks) / Total number of tracks. If a person's percentage is greater than 50%, that person is temporarily marked. This percentage can then be calculated for several consecutive days. If a person is marked with a temporary tag for a set number of consecutive days (e.g., 3 consecutive days), that person is identified as a solo traveler, and a solo traveler's facial profile is obtained.
[0113] The above analysis identifies all individuals traveling alone. These individuals can then be combined with the suspected individuals identified through the emotion category analysis. The individuals in this intersection can be considered as candidates, and their facial images are added to the personnel database to form the final database. This database can be adaptively updated as the duration of surveillance video captures increases.
[0114] After constructing the personnel database, the first facial image of the aforementioned first person can be matched against the database. If a matching candidate's facial image is found, it indicates that the first person may be a potential target of the target behavior. In this case, step 102 above can be performed to identify the emotion category of the first person in the first facial image, determining the first emotion category and subsequent target behavior detection. Alternatively, after each facial capture image is obtained, it can be matched against the personnel database. If a match is successful, the first facial image of that person is then acquired, and step 102 above can be performed to identify the emotion category of the first person in the first facial image, determining the first emotion category and subsequent target behavior detection.
[0115] In this embodiment, by matching the people in the captured face images with a preset personnel database, and then continuing to detect the target behavior of the person after a successful match, blind detection of all people can be avoided, reducing the amount of detection data and improving the accuracy of detection.
[0116] The target behavior detection device provided by the present invention is described below. The target behavior detection device described below and the target behavior detection method described above can be referred to in correspondence.
[0117] Figure 11 This is a schematic diagram of the target behavior detection device provided in an embodiment of the present invention. See also: Figure 11 As shown, the device may include: The emotion recognition module 410 is used to acquire a first face image and perform emotion category recognition processing on the first person in the first face image to determine the first emotion category corresponding to the first person; the first face image includes a face image of the first person taken from the front. The eye direction and trajectory determination module 420 is used to determine the eye direction and / or the movement trajectory direction of the first person based on the first face image and / or the surveillance video associated with the first person if the first emotion category is an emotion category related to the target behavior. The personnel set determination module 430 is used to determine a second personnel set in the surveillance video associated with the first person based on the direction of the first person's gaze and / or the direction of the first person's movement trajectory; the second personnel set includes at least one second person and a second face image associated with the second person, and the probability that the second person has a target behavior with the first person is greater than a preset probability. The target behavior detection module 440 is used to establish a first association relationship between a first person and a second person based on a first face image and a second face image, store the first association relationship and the number of times the first association relationship appears in a preset relational database, and detect whether there is a target behavior between the second person and the first person based on the number of times the first association relationship mark appears in the relational database.
[0118] In some embodiments, the above-mentioned eye direction and trajectory determination module 420 is specifically used to determine the pupil position and face angle of the first person in the first face image based on the first face image, and determine the eye direction of the first person based on the pupil position and face angle; determine at least one frame of the first monitoring image adjacent to the first face image based on the monitoring video associated with the first person, and determine the movement trajectory direction of the first person based on the first face image and the first monitoring image; the first monitoring image includes the first person.
[0119] In some embodiments, the personnel set determination module 430 is specifically used to obtain a second monitoring image corresponding to the first face image in the monitoring video associated with the first person, and determine a first initial personnel set based on the personnel in the area pointed to by the gaze in the second monitoring image; the second monitoring image is a large monitoring image corresponding to the first face image; based on the movement trajectory direction of the first person, obtain the personnel in the area on the opposite side of the movement trajectory direction in the second monitoring image to obtain a second initial personnel set; and determine the second personnel set based on the first initial personnel set and / or the second initial personnel set.
[0120] Optionally, the personnel set determination module 430 is specifically used to determine the first effective visible range corresponding to the direction of the first person's gaze towards the target side in the second monitoring image if the first person's gaze is directed towards the target side; the direction of the target side is not directed towards the center; determine the second effective visible range corresponding to the direction of the center in the second monitoring image; determine the third effective visible range based on the first and second effective visible ranges, and determine the first initial personnel set based on the personnel in the area corresponding to the third effective visible range.
[0121] In some embodiments, the above-described apparatus further includes: The relationship graph generation module is used to obtain the second association relationship between the first person and the third person and the number of occurrences of the second association relationship marker in the relational database; the third person is a person whose probability of having a target behavior with the first person is greater than a preset probability; based on the first association relationship, the number of occurrences of the first association relationship marker, the second association relationship, and the number of occurrences of the second association relationship marker, the relationship graph of the first person, the relationship graph of the second person, and the relationship graph of the third person are generated.
[0122] In some embodiments, the above-described apparatus further includes: The peer identification and display module is used to acquire the trajectory of personnel within a historical time period; the trajectory of personnel within the historical time period includes multiple fourth persons; peer analysis is performed on the trajectory of personnel within the historical time period to determine at least one peer set corresponding to the multiple fourth persons, and the same group label is assigned to the multiple peer fourth persons in each peer set; the peer set includes multiple peer fourth persons; if any target fourth person in the peer set has a relationship with the first person, then all fourth persons in the corresponding peer set are obtained according to the group label of the target fourth person, and the relationship graph of the first person and the relationship graph of each of the four fourth persons are displayed.
[0123] In some embodiments, before the emotion recognition module 410 performs emotion category recognition processing on the first person in the first face image and determines the first emotion category corresponding to the first person, the device further includes: The database acquisition module is used to acquire a preset personnel database; the personnel database includes the facial image of at least one candidate person, and the personnel database is obtained by performing emotion category recognition and individual analysis on people in historical surveillance videos; The matching and execution module is used to match the first face image with the face images of candidates in the personnel database. If the match is successful, it returns to the above steps of performing emotion category recognition processing on the first person in the first face image to determine the first emotion category corresponding to the first person.
[0124] It should be noted that the apparatus provided in this embodiment of the invention can implement all the method steps implemented in the above method embodiment and can achieve the same technical effect. Therefore, the parts and beneficial effects that are the same as those in the method embodiment will not be described in detail here.
[0125] Figure 12 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 12As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communications bus 540. The processor 510 can call logical instructions in the memory 530 to execute a target behavior detection method. This method includes: acquiring a first face image and performing emotion category recognition processing on a first person in the first face image to determine a first emotion category corresponding to the first person; the first face image includes a frontal view of the first person; if the first emotion category is an emotion category related to the target behavior, then determining the direction of the first person's gaze and / or the direction of the first person's movement trajectory based on the first face image and / or the surveillance video associated with the first person; determining a set of second persons in the surveillance video associated with the first person based on the direction of the first person's gaze and / or the direction of the first person's movement trajectory; the set of second persons includes at least one second person and a second face image associated with the second person, wherein the probability that the second person has a target behavior with the first person is greater than a preset probability; establishing a first association relationship between the first person and the second person based on the first face image and the second face image, storing the first association relationship and marking the occurrence frequency of the first association relationship in a preset relational database, and detecting whether there is a target behavior between the second person and the first person based on the occurrence frequency of the first association relationship marking in the relational database.
[0126] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0127] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is capable of executing the target behavior detection method provided by the above methods, the method comprising: acquiring a first face image, and performing emotion category recognition processing on a first person in the first face image to determine a first emotion category corresponding to the first person; the first face image includes a face image of the first person taken from the front; if the first emotion category is an emotion category related to the target behavior, then determining the direction of the first person's gaze and / or the direction of the first person's gaze based on the first face image and / or the surveillance video associated with the first person. The first person's movement trajectory direction; based on the first person's gaze and / or movement trajectory direction, a second set of persons is determined from the surveillance video associated with the first person; the second set of persons includes at least one second person and a second face image associated with the second person, the probability that the second person has a target behavior with the first person is greater than a preset probability; a first association relationship is established between the first person and the second person based on the first face image and the second face image, and the first association relationship and the number of occurrences of the first association relationship are stored in a preset relational database, and the number of occurrences of the first association relationship is marked based on the number of occurrences of the first association relationship mark in the relational database is used to detect whether there is a target behavior between the second person and the first person.
[0128] In another aspect, the present invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the target behavior detection method provided by the methods described above. The method includes: acquiring a first face image, and performing emotion category recognition processing on a first person in the first face image to determine a first emotion category corresponding to the first person; the first face image includes a frontal view of the first person; if the first emotion category is an emotion category related to the target behavior, then determining the direction of the first person's gaze and / or the direction of the first person's movement trajectory based on the first face image and / or a surveillance video associated with the first person; based on... The direction of the first person's gaze and / or the direction of the first person's movement trajectory are used to determine a set of second persons in the surveillance video associated with the first person; the set of second persons includes at least one second person and a second face image associated with the second person, and the probability that the second person has a target behavior with the first person is greater than a preset probability; a first association relationship is established between the first person and the second person based on the first face image and the second face image, and the first association relationship and the number of occurrences of the first association relationship are stored in a preset relational database, and the presence of a target behavior between the second person and the first person is detected based on the number of occurrences of the first association relationship marker in the relational database.
[0129] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0130] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A target behavior detection method, characterized in that, include: A first face image is acquired, and an emotion category recognition process is performed on the first person in the first face image to determine the first emotion category corresponding to the first person; the first face image includes a face image of the first person taken from the front. If the first emotion category is an emotion category related to the target behavior, then the direction of the first person's gaze and / or the direction of the first person's movement trajectory are determined based on the first facial image and / or the surveillance video associated with the first person. Based on the direction of the first person's gaze and / or the direction of the first person's movement trajectory, a second set of persons is determined in the surveillance video associated with the first person; the second set of persons includes at least one second person and a second face image associated with the second person, and the probability that there is a target behavior between the second person and the first person is greater than a preset probability; A first association relationship is established between the first person and the second person based on the first face image and the second face image. The first association relationship and the number of times the first association relationship appears are stored in a preset relationship database. Based on the number of times the first association relationship tag appears in the relationship database, it is detected whether there is a target behavior between the second person and the first person.
2. The target behavior detection method according to claim 1, characterized in that, The step of determining the direction of the first person's gaze and / or the direction of the first person's movement trajectory based on the first facial image and / or the surveillance video associated with the first person includes: Based on the first face image, determine the pupil position and face angle of the first person in the first face image, and determine the direction of the first person's gaze based on the pupil position and face angle; Based on the surveillance video associated with the first person, at least one frame of the first surveillance image adjacent to the first face image is determined, and the movement trajectory direction of the first person is determined based on the first face image and the first surveillance image; the first surveillance image includes the first person.
3. The target behavior detection method according to claim 1 or 2, characterized in that, The step of determining the second set of persons from the surveillance video associated with the first person based on the direction of the first person's gaze and / or the direction of the first person's movement trajectory includes: Obtain the second surveillance image corresponding to the first face image from the surveillance video associated with the first person, and determine the first initial set of people based on the people in the area pointed to by the gaze in the second surveillance image; the second surveillance image is the large surveillance image corresponding to the first face image; Based on the movement trajectory direction of the first person, people in the area opposite to the movement trajectory direction are obtained in the second monitoring image to obtain a second initial set of people; The second set of personnel is determined based on the first initial set of personnel and / or the second initial set of personnel.
4. The target behavior detection method according to claim 3, characterized in that, The step of determining the second set of personnel based on the first initial set of personnel and / or the second initial set of personnel includes: Take the union of the first initial set of personnel and the second initial set of personnel, and obtain the third monitoring image corresponding to the personnel in the union from the monitoring videos associated with the personnel in the union; The gestures of the people in the third surveillance image are identified to obtain the gesture recognition results; The emotion categories of the people in the third surveillance image are identified to obtain the emotion category identification results; The direction of the eyes of the people in the third surveillance image is analyzed to obtain the result of the eye direction analysis; Based on at least one of the gesture recognition results, the emotion category recognition results, and the eye direction analysis results, the second set of people is determined from the people in the union set.
5. The target behavior detection method according to claim 3, characterized in that, The step of determining the first initial set of people based on the area pointed to by the gaze in the second surveillance image includes: If the first person's gaze is directed toward the target side, then the first effective visible range corresponding to the direction toward the target side is determined in the second monitoring image; the direction toward the target side is not toward the exact center; In the second monitoring image, determine the second effective visible range corresponding to the orientation towards the exact center; Based on the first effective visible range and the second effective visible range, a third effective visible range is determined, and based on the people in the area corresponding to the third effective visible range, the first initial set of people is determined.
6. The target behavior detection method according to claim 1, characterized in that, The method further includes: The second association relationship between the first person and the third person, as well as the frequency of occurrence of the second association relationship marker, are obtained from the relational database; the third person is a person whose probability of having a target behavior with the first person is greater than a preset probability. Based on the first association, the number of times the first association marker appears, the second association, and the number of times the second association marker appears, a relationship graph of the first person, a relationship graph of the second person, and a relationship graph of the third person are generated.
7. The target behavior detection method according to claim 1, characterized in that, The method further includes: Obtain the trajectory of personnel within a historical time period; the trajectory of personnel within the historical time period includes multiple fourth persons; Peer analysis is performed on the trajectory of personnel within the historical time period to determine at least one set of peers corresponding to the multiple fourth persons, and the same group label is assigned to the multiple peers in each set of peers; the set of peers includes multiple peers. If any target fourth person in the set of peers is associated with the first person, then all fourth persons in the corresponding set of peers are obtained according to the group tag corresponding to the target fourth person, and the relationship graph of the first person and the relationship graph of each of the fourth persons are displayed.
8. The target behavior detection method according to claim 1, characterized in that, Before performing emotion category recognition processing on the first person in the first facial image to determine the first emotion category corresponding to the first person, the method further includes: Obtain a preset personnel database; the personnel database includes facial images of at least one candidate person, and the personnel database is obtained by performing emotion category recognition and individual analysis on personnel in historical surveillance videos; The first face image is matched with the face images of candidates in the personnel database. If the match is successful, the process returns to the step of performing emotion category recognition processing on the first person in the first face image to determine the first emotion category corresponding to the first person.
9. A target behavior detection device, characterized in that, include: An emotion recognition module is used to acquire a first face image and perform emotion category recognition processing on a first person in the first face image to determine the first emotion category corresponding to the first person; the first face image includes a face image of the first person taken from the front. The eye direction and trajectory determination module is used to determine the eye direction and / or the movement trajectory direction of the first person based on the first face image and / or the surveillance video associated with the first person if the first emotion category is an emotion category related to the target behavior. The personnel set determination module is used to determine a second personnel set in the surveillance video associated with the first person based on the direction of the first person's gaze and / or the direction of the first person's movement trajectory; the second personnel set includes at least one second person and a second face image associated with the second person, and the probability that there is a target behavior between the second person and the first person is greater than a preset probability; The target behavior detection module is used to establish a first association relationship between the first person and the second person based on the first face image and the second face image, store the first association relationship and the number of times the first association relationship appears in a preset relationship database, and detect whether there is a target behavior between the second person and the first person based on the number of times the first association relationship mark appears in the relationship database.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the target behavior detection method as described in any one of claims 1 to 8.