Bystander detection system and bystander detection method
By using personnel detection and bystander assessment modules, and leveraging distance information from images and various evaluation methods, the problem of false alarms in bystander detection in existing technologies has been solved, achieving more accurate safety classification and judgment.
Patent Information
- Application Number
- CN202410631142.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-21
- Publication Date
- 2025-11-21
AI Technical Summary
In existing technologies, bystander detection systems based on depth map detection are prone to false alarms and cannot accurately distinguish whether a non-user exists or determine its security classification.
The system employs a personnel detection module and a bystander assessment module. By receiving images and obtaining distance information of personnel, it determines whether there are non-users and assesses their safety classification based on various information, including bystander category.
It improves the accuracy and precision of bystander detection, enabling a more comprehensive assessment of people's situation in camera footage and reducing false alarms.
Smart Images

Figure CN120997872A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of information display security. In particular, the present disclosure relates to a technique of applying a non-user feature in a video to determine a security classification of the non-user. BACKGROUND
[0002] In recent years, with the increasing awareness of information security, many systems and applications have introduced the function of detecting onlookers using depth maps to ensure the privacy and security of users. However, when someone passes behind the user without engaging in any peeping behavior, false positives may occur due to the use of only the depth information in the depth map for judgment. SUMMARY
[0003] Therefore, some embodiments of the present disclosure provide an onlooker detection system and an onlooker detection method to improve the problems of the prior art.
[0004] Some embodiments of the present disclosure provide an onlooker detection system, comprising: a person detection module configured to receive a video and, in response to the presence of a person in the video, obtain person information of each person, wherein the person information comprises distance information relative to the device; and an onlooker judgment module configured to perform: judging whether at least one non-user is within a range based on the distance information of the person information of each person; and in response to the presence of at least one non-user within the range, judging a security classification to which each non-user belongs based on the person information of each non-user, wherein the security classification comprises an onlooker category.
[0005] Some embodiments of the present disclosure provide an onlooker detection method, comprising: receiving a video by a person detection module and, in response to the presence of at least one person in the video, obtaining person information of each person, wherein the person information comprises distance information relative to the device; and performing by an onlooker judgment module: judging whether at least one non-user is within a range based on the distance information of the person information of each person; and in response to the presence of a non-user within the range, judging a security classification to which each non-user belongs based on the person information of each non-user, wherein the security classification comprises an onlooker category.
[0006] Based on the above, the onlooker detection system and the onlooker detection method provided by the embodiments of the present disclosure obtain multiple information through vision to more comprehensively evaluate the situation of the person in the video obtained by the lens to increase the accuracy of the judgment. BRIEF DESCRIPTION OF DRAWINGS
[0007] Figure 1 is a block diagram of an onlooker detection system according to some embodiments of the present disclosure. Figures 2-5 This is a schematic diagram of the operation of a personnel detection module according to some embodiments of the present invention. Figure 6 These are schematic diagrams of detection based on some embodiments of the present invention. Figure 7 These are schematic diagrams depicting key points according to some embodiments of the present invention. Figure 8 This is a schematic diagram of the angle information corresponding to human face information, as depicted in some embodiments of the present invention. Figure 9 This is a schematic diagram of distance measurement depicted according to some embodiments of the present invention. Figure 10 This is a block diagram of a neural network module depicted according to some embodiments of the present invention. Figure 11 This is a block diagram of the output feature tensor generation module described according to some embodiments of the present invention. Figure 12 This is a block diagram of a fusion module depicted according to some embodiments of the present invention. Figure 13 This is a schematic diagram of the prediction module structure depicted according to some embodiments of the present invention. Figure 14 This is a schematic diagram of an information tensor structure depicted according to some embodiments of the present invention. Figure 15 This is a block diagram of an electronic device system depicted according to some embodiments of the present invention. Figures 16-22 This is a flowchart of a bystander detection method based on some embodiments of the present invention. Detailed Implementation
[0008] The foregoing and other technical contents, features, and effects of the present invention will be clearly presented in the following detailed description of the embodiments with reference to the accompanying drawings. Any actions that do not affect the effects and objectives achieved by the present invention should still fall within the scope of the technical contents disclosed in the present invention.
[0009] Figure 1 This is a block diagram of a bystander detection system depicted according to some embodiments of the present invention. Figures 2-3 This is a schematic diagram illustrating the operation of a personnel detection module according to some embodiments of the present invention. Please refer to... Figures 1-3 The bystander detection system 100 includes a personnel detection module 101 and a bystander determination module 102. The personnel detection module 101 is configured to receive images 103 (e.g., ...). Figure 2 Image 200 or Figure 3 Image 300), and in response to the presence of at least one person (e.g., in image 103)Figure 2 Image 200 and Figure 3 The bystander determination module 102 is configured to receive the personnel information of the personnel detected by the personnel detection module 101 and perform further judgment and processing. This information includes distance information relative to a device, such as the display screen of an electronic device. The distance information included in the personnel information includes the distance of the person in the image relative to the display surface of the electronic device's screen.
[0010] In some embodiments of the present invention, the personnel detection module 101 first determines which person in the image 103 is the user. The personnel detection module 101 can determine that the person closest to the device is the user, or it can first find multiple people closest to the device and then find the one closest to the center and determine it as the user. The present invention does not limit the method of determining the user. Figure 2 Image 200 and Figure 3 Taking image 300 as an example, the personnel detection module 101 determines that personnel 201 is a user. In image 103, personnel who are not users are referred to as non-users.
[0011] The following detailed description, with reference to the accompanying drawings, illustrates how the bystander detection method and the various modules of the bystander detection system 100 of some embodiments of the present invention work together.
[0012] Figure 16 This is a flowchart illustrating a bystander detection method according to some embodiments of the present invention. Please refer to... Figures 1-3 , Figure 16 .exist Figure 16 In this embodiment, the bystander detection method includes steps S1601 to S1603. In step S1601, the personnel detection module 101 receives the image 103 and, in response to the presence of personnel in the image, obtains personnel information for each person, wherein the personnel information includes distance information relative to a device. In step S1602, the bystander judgment module 102 determines, based on the distance information of the personnel detected by the personnel detection module 101, whether there are any non-users within the range. Taking the aforementioned device as the display screen of an electronic device as an example, the aforementioned range is set to a default distance in front of the display screen of the electronic device.
[0013] In step S1603, in response to determining that at least one non-user is within the aforementioned range, the bystander determination module 102 determines the security category to which each non-user belongs based on the personnel information of each non-user, wherein the security category includes the bystander category. The bystander determination module 102's determination that a non-user belongs to the bystander category indicates that the bystander determination module 102 determines that this non-user poses a danger of being spied on.
[0014] Figure 17 is a flowchart of a method of detecting a bystander according to some embodiments of the present application. Please refer to Figures 2-3 、 Figure 17 In Figure 17 some embodiments, the aforementioned security classification further comprises a non-bystander category in addition to the bystander category. The bystander judging module 102 judges that the non-user belongs to the non-bystander category means that the non-user is not dangerous to the device. The aforementioned step S1603 comprises steps S1701-S1706. In step S1701, the bystander judging module 102 judges whether a current object in the non-users is facing the aforementioned device, if yes, step S1702 is executed, if not, step S1703 is executed. In step S1702, in response to the current object facing the device, the current object is judged to belong to the bystander category. In step S1703, in response to the current object not facing the device, the current object is judged to belong to the non-bystander category. For example, in Figure 2 and Figure 3 , the bystander judging module 102 judges that the person 201 is a user, the person 202 is a non-user and is within the range. In Figure 2 , the bystander judging module 102 judges that the person 202 belongs to the bystander category. In Figure 3 , the bystander judging module 102 judges that the person 202 belongs to the non-bystander category. In step S1704, the bystander judging module 102 judges whether there is an unselected object in the non-users, if yes, step S1706 is executed, if not, step S1705 is executed. In step S1705, since the bystander judging module 102 has judged the security category to which all the non-users belong, the present procedure is exited. In step S1706, in response to there being an unselected object in the non-users, the bystander judging module 102 selects one of the non-users that has not been selected as the current object and returns to step S1701.
[0015] Figures 4-5 is a schematic diagram of the operation of a person detecting module according to some embodiments of the present application. Figure 18 is a flowchart of a method of detecting a bystander according to some embodiments of the present application. Please refer to Figures 4-5 、 Figure 18 In Figure 18In some embodiments of the present application, the passerby judging module 102 calculates the distance between the current object and the user based on the distance information of each person relative to the device. In some embodiments of the present application, the passerby judging module 102 calculates the distance between the current object and the user based on the distance information of each person relative to the device.
[0016] In some embodiments of the present application, the passerby judging module 102 calculates the distance between the current object and the user based on the distance information of each person relative to the device. In Figure 4 In some embodiments of the present application, the passerby judging module 102 calculates the distance between the current object and the user based on the distance information of each person relative to the device. In Figure 5 In some embodiments of the present application, the passerby judging module 102 calculates the distance between the current object and the user based on the distance information of each person relative to the device. In
[0017] In step S1804, in response to the current object facing the device, the passerby judging module 102 judges that the current object belongs to the passerby category. In step S1805, in response to the current object not facing the device, the passerby judging module 102 judges that the current object belongs to the passerby category. In step S1806, the passerby judging module 102 judges whether there is an unselected object among the non-users, if yes, it executes step S1808, if not, it executes step S1807. In step S1807, since the passerby judging module 102 has judged the security category of all non-users, it exits the program. In step S1808, in response to there being an unselected object among the non-users, the passerby judging module 102 selects one of the non-users that has not been selected as the current object and returns to step S1801.
[0018] Figure 19 is a flowchart of a passerby detection method according to some embodiments of the present application. In Figure 19The aforementioned personnel information includes facial information, the corresponding angle information of the facial information, and key point information. The aforementioned facial information includes the bounding box of the detected person's face (e.g.,...). Figures 2-3 Face frames 2011, 2021 Figures 4-5 The position, height, and width of the face bounding boxes (5011, 5021, 5031) are specified. If a person is detected but the position, height, and width of the face bounding box are not specified, it indicates that no person's face has been detected. This will be further explained in later embodiments. The aforementioned step of determining whether the current object is facing the device includes steps S1901 to S1903 executed by the bystander judgment module 102. In step S1901, based on the face information of the current object, it is determined whether the face of the current object is detected. If yes, step S1902 is executed; otherwise, step S1903 is executed. In step S1902, in response to the detection of the face of the current object, it is determined whether the current object is facing the device based on the angle information of the corresponding face information of the current object. In step S1903, in response to the absence of a detected face of the current object, it is determined whether the current object is facing the device based on the key point information of the current object.
[0019] Figure 8 This is a schematic diagram illustrating the angle information corresponding to facial information according to some embodiments of the present invention. Please refer to... Figure 8 The head of the face in image 103 is shown as head 804. The head pitch angle of the face in image 103 is the angle rotated about the x-axis 801 with the center of the face facing the device as a reference, and its range is [-180°, 180°]. The head yaw angle of the face in image 103 is the angle rotated about the y-axis 802 with the center of the face facing the device as a reference, and its range is [-90°, 90°]. The head roll angle of the face in image 103 is the angle rotated about the z-axis 803 with the center of the face facing the device as a reference, and its range is [0°, 360°]. When the center of the face of head 804 is aligned with the device, the head pitch angle, head yaw angle, and head roll angle are all 0 degrees.
[0020] The pitch angle of the gaze point of the face in image 103 is the angle of rotation about the x-axis 801 with the gaze direction 805 of the head 804 facing the device as a reference, and its range is [-180°, 180°]. The yaw angle of the gaze point of the face in image 103 is the angle of rotation about the y-axis 802 with the gaze direction 805 of the head 804 facing the device as a reference, and its range is [-90°, 90°]. The roll angle of the gaze point of the face in image 103 is the angle of rotation about the z-axis 803 with the gaze direction 805 of the head 804 facing the device as a reference, and its range is [0°, 360°]. When the gaze direction 805 of the head 804 is aligned with the device, the pitch angle, yaw angle, and roll angle of the gaze point of the face are 0 degrees.
[0021] In some embodiments of the present disclosure, the angle information of the corresponding face information of the current object includes a head yaw angle, and the aforementioned step S1902 includes: determining that the current object faces the device in response to the head yaw angle being within an angle threshold range; and determining that the current object does not face the device in response to the head yaw angle not being within the angle threshold range. For example, in the case of the head yaw angle of the person 202 being 3°, the bystander judgment module 102 determines that the person 202 faces the device. In the case of the head yaw angle of the person 202 being 60°, the bystander judgment module 102 determines that the person 202 does not face the device. Figure 2 With Figure 3 For example, in the case of the head yaw angle of the person 202 being 3°, the bystander judgment module 102 determines that the person 202 faces the device. In the case of the head yaw angle of the person 202 being 60°, the bystander judgment module 102 determines that the person 202 does not face the device. Figure 2 With Figure 3 In the case of the head yaw angle of the person 202 being 3°, the bystander judgment module 102 determines that the person 202 faces the device. In the case of the head yaw angle of the person 202 being 60°, the bystander judgment module 102 determines that the person 202 does not face the device. Figure 2 In the case of the head yaw angle of the person 202 being 3°, the bystander judgment module 102 determines that the person 202 faces the device. In the case of the head yaw angle of the person 202 being 60°, the bystander judgment module 102 determines that the person 202 does not face the device. Figure 3 In the case of the head yaw angle of the person 202 being 3°, the bystander judgment module 102 determines that the person 202 faces the device. In the case of the head yaw angle of the person 202 being 60°, the bystander judgment module 102 determines that the person 202 does not face the device.
[0022] In some embodiments of the present disclosure, the angle information of the corresponding face information of the current object includes a gaze yaw angle, and the aforementioned step S1902 includes: determining that the current object faces the device in response to the gaze yaw angle being within an angle threshold range; and determining that the current object does not face the device in response to the gaze yaw angle not being within the angle threshold range.
[0023] Figure 6 is a detection diagram according to some embodiments of the present disclosure. Figure 7 is a key point diagram according to some embodiments of the present disclosure. Please refer to Figures 6-7 When the image 103 includes the image 600, the bystander judgment module 102 judges, based on the received face information, that the face of the person 601 is detected, and the face range of the person 601 is indicated by the face frame 6011. At this time, the bystander judgment module 102 performs the aforementioned step S1902. However, if the image 103 only includes the image 6015, the bystander judgment module 102 performs step S1903 to determine whether the current object faces the device based on the key point information of the current object. In Figure 7 In the case of the head yaw angle of the person 202 being 3°, the bystander judgment module 102 determines that the person 202 faces the device. In the case of the head yaw angle of the person 202 being 60°, the bystander judgment module 102 determines that the person 202 does not face the device.
[0024] Figure 20 is a bystander detection method flowchart according to some embodiments of the present disclosure. In Figure 20In some embodiments of the present disclosure, the person detection module 101 detects the center 701, the left shoulder 705, and the right shoulder 702 of the person to obtain the left shoulder point coordinate of the left shoulder point 6012, the right shoulder point coordinate of the right shoulder point 6014, and the center point coordinate of the center point 6013 as the key point information. The aforementioned step S1903 includes steps S2001-S2004. In step S2001, a first distance from the left shoulder point coordinate to the center point coordinate and a second distance from the right shoulder point coordinate to the center point coordinate are calculated. The first distance is divided by the distance of the current object relative to the device to obtain a first normalized distance, and the second distance is divided by the distance of the current object relative to the device to obtain a second normalized distance.
[0025] In step S2002, it is determined whether the absolute value of the difference between the first normalized distance and the second normalized distance is greater than a distance difference threshold (that is, it is determined whether the following inequality holds: first normalized distance-second normalized distance>distance difference threshold In step S2003, if not, step S2004 is performed. In step S2003, in response to the aforementioned inequality holding, it is determined that the current object does not face the device. In step S2004, in response to the aforementioned inequality not holding, it is determined that the current object faces the device.
[0026] Figure 9 is a distance measurement schematic diagram according to some embodiments of the present disclosure. In Figure 9 In some embodiments of the present disclosure, the person detection module 101 includes an infrared laser diode 902 and an infrared image sensor 903. The infrared laser diode 902 emits infrared light to the person 901 via a direction 904, and the infrared image sensor 903 receives the reflection via a direction 905. The person detection module 101 calculates the distance of the person 901 relative to the device based on the time difference between the emission and the reception of the reflection to serve as the distance information in the person information of the person 901. It is worth noting that the face distance estimation can also be achieved by other means, such as using a TOF sensor or a multi-sensor (based on a phase difference method) to obtain a depth map, or a method of estimating through a single-lens (Mono Camera) image.
[0027] Figure 15 is an electronic device system block diagram according to some embodiments of the present disclosure. In Figure 15 In some embodiments of the present disclosure, the electronic device 1500 includes the bystander detection system 100 and a display module 1501. The electronic device 1500 also includes a display screen, which is the device mentioned in the aforementioned embodiments. The display module 1501 controls the display screen. The bystander detection method further includes: in response to the bystander determination module 102 determining that at least one of the non-users belongs to the bystander category, sending a signal to make the device start performing the anti-peeping program. In some embodiments of the present disclosure, the bystander determination module 102 sends a signal to the display module 1501 to control the display screen, and the display module 1501 controls the display screen to display a warning message.Figure 4 Change Secure to display Figure 5 The word "Alert".
[0028] Figure 10 This is a block diagram of a neural network module depicted according to some embodiments of the present invention. Figure 21 This is a flowchart illustrating a bystander detection method according to some embodiments of the present invention. Please refer to... Figure 1 , Figure 10 , Figure 21 In this embodiment, the personnel detection module 101 includes a neural network module 1000. The neural network module 1000 is configured to receive image 103 and, in response to the presence of at least one person in the image, output multiple information tensors. The aforementioned step S1601 includes step S2101: the neural network module 1000 receives image 103 and, in response to the presence of at least one person in image 103, outputs multiple information tensors; and step S2102: the personnel detection module 101, in response to the presence of at least one person in image 103, outputs personnel information for each person based on the aforementioned information tensors.
[0029] The following provides further description of various implementations of the neural network module 1000. The neural network module 1000 includes an output feature tensor generation module 1001 and prediction modules 1002-1 to 1002-M, where M > 1. The output feature tensor generation module 1001 generates multiple output feature tensors of different sizes based on the image 103. Each of the prediction modules 1002-1 to 1002-M receives one of the aforementioned output feature tensors to generate a corresponding information tensor. The information tensor indicates face information, confidence score information, category information, angle information corresponding to the face information, and key point information. Based on all the information tensors generated by each of the prediction modules 1002-1 to 1002-M, the personnel detection module 101 outputs personnel information for each person based on the aforementioned information tensors in response to the presence of at least one person in the image 103.
[0030] Figure 22 This is a flowchart depicting a determination method according to some embodiments of the present invention. Please refer to... Figure 22The step S2101 includes steps S2201-S2202. In step S2201, the output feature tensor generation module 1001 generates a plurality of output feature tensors with different sizes based on the image 103. In step S2202, each of the prediction modules 1002-1-1002-M receives a corresponding one of the plurality of output feature tensors to generate a corresponding one of the information tensors, respectively. Each of the information tensors is configured to indicate face information, confidence score information, class information, angle information of a corresponding face information, and key point information. The face information includes a position, a height, and a width of a face bounding box of a detected person, and if there is a detected person without a face bounding box, the face information indicates that there is no face of the detected person. The angle information of the corresponding face information includes a head-pitch angle, a head-yaw angle, and a head-roll angle of the face. The key point information includes a left shoulder point coordinate, a right shoulder point coordinate, and a center point coordinate. It is worth mentioning that the angle information of the corresponding face information can be selected to include only the needed angles, such as the head-yaw angle, and the key point information can also include other key points based on different applications. Figure 7 other key points.
[0031] Figure 11 is a block diagram of an output feature tensor generation module according to some embodiments of the present disclosure. Please refer to Figures 10-11 Hereinafter, M=3 is taken as an example. The output feature tensor generation module 1001 includes a backbone module 10011 and a feature pyramid module 10012.
[0032] In some embodiments of the present disclosure, the backbone module 10011 includes backbone layers 100111-100114 with different sizes. The backbone module 10011 generates a plurality of feature tensors with different sizes and a first order based on the image 103 through the backbone layers 100111-100114. As Figure 11 depicted, the plurality of feature tensors are output tensors of the backbone layers 100112-100114. The first order is an arrangement order of the plurality of feature tensors from large to small according to their sizes. The feature pyramid module 10012 performs feature fusion on the feature tensors to obtain a plurality of output feature tensors.
[0033] Please refer to Figure 11 , Figure 22In some embodiments of the present application, the aforementioned step S2201 comprises the following steps: generating, by the backbone module 10011, a plurality of feature tensors of different sizes and having a first order based on the image 103 through the backbone layers 100111-100114, wherein the first order is an arrangement order of the plurality of feature tensors according to their sizes from large to small; and performing, by the feature pyramid module 10012, feature fusion on the aforementioned feature tensors of different sizes and having the first order generated by the backbone module 10011 through the backbone layers 100111-100114 based on the image 103 to obtain output feature tensors.
[0034] Please refer to Figure 11 The feature pyramid module 10012 comprises fusion modules 100121-1-100121-2. The feature pyramid module 10012 performs the following steps to perform feature fusion on the aforementioned feature tensors to obtain a plurality of output feature tensors.
[0035] First, the feature pyramid module 10012 sets the smallest feature tensor corresponding to the last one in the aforementioned first order as one of the temporary feature tensor set, so as to Figure 11 For example, the depicted embodiment, the smallest feature tensor is the output tensor of the backbone layer 100114, and the smallest feature tensor is stored in the temporary feature tensor 100122-3 as one of the temporary feature tensor set.
[0036] Then, the feature pyramid module 10012 performs an upsampling operation on the temporary feature tensor 100122-3 through the fusion module 100121-1 to obtain an upsampled temporary feature tensor 100122-3 of the same size as the output tensor of the backbone layer 100113, and the feature pyramid module 10012 performs feature fusion on the upsampled temporary feature tensor 100122-3 and the output tensor of the backbone layer 100113 through the fusion module 100121-1 to obtain a temporary feature tensor 100122-2 of the same size as the output tensor of the convolution layer of the backbone layer 100113. The feature pyramid module 10012 performs feature fusion on the upsampled temporary feature tensor 100122-2 and the output tensor of the convolution layer of the backbone layer 100112 through the fusion module 100121-2 to obtain a temporary feature tensor 100122-1 of the same size as the output tensor of the convolution layer of the backbone layer 100112. The feature pyramid module 10012 outputs the temporary feature tensors 100122-3, 100122-2, 100122-1 as the plurality of output feature tensors of the aforementioned feature pyramid module 10012.
[0037] Figure 12 is a block diagram of a fusion module according to some embodiments of the present application. In Figure 12In some embodiments, the structure of the fusion module 100121-1~100121-2 is depicted as fusion module 1200. The fusion module 1200 comprises an up-sampling module 1201, a point-wise convolution layer 1202, and a point-wise addition module 1203. The up-sampling module 1201 is configured to perform an up-sampling operation on the input of the up-sampling module 1201, wherein the up-sampling operation is to repeat the input of the up-sampling module 1201 in the height axis and the width axis directions by 2 times so that the size of the input of the up-sampling module 1201 is converted to be twice of the original size. The point-wise convolution layer 1202 is configured to perform a point-wise convolution operation. The point-wise addition module 1203 is configured to perform a point-wise addition operation on two input tensors received by the point-wise addition module 1203 to obtain an output tensor of the point-wise addition module 1203. It is worth mentioning that the up-sampling module 1201 can employ other up-sampling methods.
[0038] Figure 13 is a schematic diagram of a prediction module structure according to some embodiments of the present application. Figure 14 is a schematic diagram of an information tensor structure according to some embodiments of the present application. Please refer to Figures 13-14 In some embodiments, the structure of the prediction module 1002-1~1002-3 is depicted as prediction module 1300. The prediction module 1300 comprises t W p ×H p ×128 convolution layers: convolution layers 1301-1~1301-t, and a W p ×H p ×PA convolution layer: convolution layer 1302, wherein t is a positive integer; W p and H p are positive integers representing the dimensions of the width axis and the height axis of the convolution layers 1301-1~1301-t, A is a positive integer representing the number of anchors, and P is a positive integer. It is worth mentioning that the convolution layer labeled as W p ×H p ×128 represents that the convolution layer performs a convolution operation on an input tensor with 128 convolution kernels, and concatenates the tensors obtained by performing the convolution operation on the input tensor with the 128 convolution kernels in sequence to obtain an output tensor with the number of width axes being W p , the number of height axes being H p , and the number of channels being 128, and such an output tensor is a tensor with the dimensions of W p ×H p ×128.
[0039] The neural network module 1000 sets A bounding boxes with different sizes on the aforementioned multiple output feature tensors. The value of P is 4+1+number of all classes+3+6, where 4 represents the number of tensor elements needed to describe the position coordinates of a certain vertex of the bounding box and the width and height of the bounding box, 1 represents the number of tensor elements needed to describe the possibility of having a detection target in the bounding box and the accuracy of the bounding box, 3 represents the number of tensor elements needed to describe the head pitch angle, head yaw angle and head roll angle of the face, and 6 represents the number of tensor elements needed to describe the left shoulder point coordinates, right shoulder point coordinates and center point coordinates (two tensors are needed for each coordinate). p p The values of P, A, and t can be set by the user according to the needs. It is worth noting that since the sizes of the output feature tensors received by the prediction modules 1002-1 to 1002-M are different, the values of W p and H p of each of the prediction modules 1002-1 to 1002-M are different.
[0040] The prediction module 1300 receives any one of the aforementioned multiple output feature tensors, and when the output feature tensor passes through the convolution layers 1301-1 to 1301-t and the convolution layer 1302 of the class prediction module 1300, an information tensor 1401 can be obtained. The information tensor 1401 includes sub-information tensors 1401-1 to 1401-A, and each of the sub-information tensors 1401-1 to 1401-A corresponds to one of the aforementioned A bounding boxes. Each of the sub-information tensors 1401-1 to 1401-A includes W p ·H p P vectors. As Figure 14 depicted, each of the P vectors includes tensor elements 14021-1, 14021-2, 14022-1, 14022-2, 14023-1, 14023-2, 1403-1, 1403-2, 1403-3, 1404-1409, and the tensor elements 14021-1, 14021-2, 14022-1, 14022-2, 14023-1, 14023-2, 1403-1, 1403-2, 1403-3 respectively indicate the horizontal coordinate of the left shoulder point (such as the left shoulder point 6012), the vertical coordinate of the left shoulder point, the horizontal coordinate of the right shoulder point, the vertical coordinate of the right shoulder point, the horizontal coordinate of the center point, the vertical coordinate of the center point, the head pitch angle of the face, the head yaw angle and the head roll angle.
[0041] The tensor element 1404 includes a plurality of sub-tensor elements, each of the sub-tensor elements of the tensor element 1404 indicates a probability that the object in the detection frame belongs to a respective class, the tensor element 1405 indicates a confidence score, the confidence score represents a possibility that there is a detection target in the detection frame and an accuracy degree of the detection frame, the tensor element 1406 indicates a height of the detection frame, the tensor element 1407 indicates a width of the detection frame, and the tensor elements 1408 and 1409 indicate detection frame coordinates. The face information includes the detection frame coordinates, the height of the detection frame, and the width of the detection frame, the aforementioned probability that the object in the detection frame belongs to a respective class is the aforementioned class information, the aforementioned confidence score is the aforementioned confidence score information, the angle information corresponding to the face information includes a head pitch angle, a head yaw angle, and a head roll angle of the face, and the key point information includes a horizontal coordinate of a left shoulder point, a vertical coordinate of the left shoulder point, a horizontal coordinate of a right shoulder point, a vertical coordinate of the right shoulder point, a horizontal coordinate of a center point, and a vertical coordinate of the center point. The personnel detection module 101 integrates all the information tensors generated by each of the prediction modules 1002-1 to 1002-M, and can obtain personnel information of each person.
[0042] It is worth noting that the personnel detection module 101 integrates all the information tensors generated by each of the prediction modules 1002-1 to 1002-M, and can obtain the width and height of the face frame. In some embodiments of the present application, the bystander detection system 100 captures the image 103 by a lens arranged at a fixed position of the device, so the width and height of the face frame are inversely proportional to the distance of the face relative to the lens (and relative to the device), and thus the personnel detection module 101 can obtain the distance of the face relative to the lens based on the width or height of the face frame.
[0043] It is worth noting that, when training the neural network module 1000 in the training Figures 10-15 , as long as the data of the head pitch angle, the head yaw angle, and the head roll angle of the face, the coordinates of the left shoulder point, the coordinates of the right shoulder point, and the coordinates of the center point are added to the training set and trained by the training method of the object detection model, the trained neural network module 1000 can be obtained.
[0044] In the training Figures 10-15 , the prediction modules 1002-1 to 1002-M are referred to as network heads in the technical field of the present application. The prediction modules 1002-1 to 1002-M disclosed in the aforementioned embodiments can replace the network heads of the object detection models of other stages, so that the object detection models of other stages can output personnel information. The present application is not limited to the backbone module 10011 and the feature pyramid module 10012. Symbol explanation
[0045] 100: bystander detection system 101: person detection module 102: bystander judgment module 103, 200, 300, 600, 6015: image 201, 202, 501, 502, 503, 601, 901: person 2011, 2021, 5011, 5021, 5031, 6011: face frame 6012: left shoulder point 6013: center point 6014: right shoulder point 700-717: human body key point 801: x-axis 802: y-axis 803: z-axis 804: head 805: gaze direction 902: infrared laser diode 903: infrared image sensor 904, 905: direction 1000: neural network module 1001: output feature tensor generation module 1002-1-1002-M: prediction module M: positive integer greater than 1 10011: backbone module 10012: feature pyramid module 100111-100114: backbone layer 100121-1, 100121-2: fusion module 100122-1, 100122-2, 100122-3: temporary storage feature tensor 1200: fusion module 1201: up-sampling module 1202: point-wise convolution layer 1203: point-wise addition module 1300: prediction module 1301-1-1301-t, 1302: convolution layer T, W p , H p , A, P: positive integer 1401: information tensor 1401-1-1401-A: sub-information tensor 14021-1, 14021-2, 14022-1, 14022-2, 14023-1, 14023-2, 1403-1, 1403-2, 1403-3, 1404-1409: tensor elements 1500: electronic device 1501: display module S1601-S1603, S1701-S1706, S1801-S1808, S1901-S1903, S2001-S2004, S2101-S2102, S2201-S2202: steps
Claims
1. A bystander detection system, comprising: a person detection module configured to receive the video and, in response to the presence of at least one person in the video, obtain person information for each of the at least one person, wherein said personnel information comprises distance information relative to the device; and a bystander judging module configured to perform: (a) judging whether there is at least one non-user-in-range in said at least one person based on said distance information of said personnel information of each of said at least one person; and (b) in response to there being said at least one non-user-in-range, judging a security classification to which each of said at least one non-user belongs based on said personnel information of each of said at least one non-user, wherein said security classification comprises a bystander category. said security classification comprises a non-bystander category, and said step (b) comprises:
2. The bystander detection system of claim 1, wherein, (b 1) judging, for a current object in said at least one non-user, whether said current object is facing said device, judging said current object to belong to said bystander category in response to said current object facing said device, and judging said current object to belong to said non-bystander category in response to said current object not facing said device; and (b2) in response to there being an unselected object in said at least one non-user, selecting an unselected one of said at least one non-user as said current object and returning to step (b 1). said security classification comprises a passer-by category and a shared-user category, and said step (b) comprises: (b 1) judging, for a current object in said at least one non-user, whether a distance between said current object and a user is less than a preset distance, judging said current object to belong to said shared-user category in response to said distance between said current object and said user being less than said preset distance, judging whether said current object is facing said device in response to said distance between said current object and said user not being less than said preset distance, judging said current object to belong to said bystander category in response to said current object facing said device, and judging said current object to belong to said passer-by category in response to said current object not facing said device; and 3. The bystander detection system of claim 1, wherein, (b2) in response to there being an unselected object in said at least one non-user, selecting an unselected one of said at least one non-user as said current object and returning to step (b 1). said personnel information comprises face information, angle information corresponding to said face information, and key point information, and said step of judging whether said current object is facing said device comprises: (b 11) judging whether a face of said current object is detected based on said face information of said current object, (b 12) judging whether said current object is facing said device based on said angle information corresponding to said face information of said current object in response to said face of said current object being detected, and (b 13) judging whether said current object is facing said device based on said key point information of said current object in response to said face of said current object not being detected. 4. The bystander detection system of claim 2 or 3, wherein, 5. The bystander detection system of claim 4, wherein, The angle information of the current object corresponding to the face information includes a head yaw angle, and the step (b12) includes: in response to the head yaw angle being within an angle threshold range, determining that the current object faces the device; and in response to the head yaw angle not being within the angle threshold range, determining that the current object does not face the device.
6. The bystander detection system of claim 4, wherein, The angle information of the current object corresponding to the face information includes a gaze point yaw angle, and the step (b12) includes: in response to the gaze point yaw angle being within an angle threshold range, determining that the current object faces the device; and in response to the gaze point yaw angle not being within the angle threshold range, determining that the current object does not face the device.
7. The bystander detection system of claim 4, wherein, The key point information of the current object includes left shoulder point coordinates, right shoulder point coordinates, and center point coordinates; the step (b13) includes: (b131) calculating a first distance from the left shoulder point coordinates to the center point coordinates and a second distance from the right shoulder point coordinates to the center point coordinates; dividing the first distance by the distance of the current object relative to the device to obtain a first normalized distance, dividing the second distance by the distance of the current object relative to the device to obtain a second normalized distance, and determining whether the absolute value of the difference between the first normalized distance and the second normalized distance is greater than a distance difference threshold; and (b132) in response to the absolute value of the difference between the first normalized distance and the second normalized distance being not greater than the distance difference threshold, determining that the current object faces the device; and in response to the absolute value of the difference between the first normalized distance and the second normalized distance being greater than the distance difference threshold, determining that the current object does not face the device. The bystander determination module is configured to perform: in response to determining that one of the at least one non-user belongs to the bystander category, sending a signal to cause the device to start performing an anti-peeping program.
8. The bystander detection system of claim 1, wherein, The person detection module includes a neural network module configured to receive the image and output a plurality of information tensors in response to the presence of the at least one person in the image, and the person detection module is configured to output the person information of each of the at least one person based on the plurality of information tensors in response to the presence of the at least one person in the image.
9. The bystander detection system of claim 1, wherein, 10. A bystander detection method, comprising: The person information includes distance information relative to a device; The image is received by a person detection module, and in response to the presence of at least one person in the image, person information of each of the at least one person is obtained, wherein, and performed by a bystander determination module; (a) determining whether there is at least one non-user in the range based on the distance information of the person information of each of the at least one person; and (b) in response to the presence of the at least one non-user in the range, determining the security category to which each of the at least one non-user belongs based on the person information of each of the at least one non-user, wherein the security category includes a bystander category.