Onlooler detection system and onlooler detection method
Patent Information
- Application Number
- TW113117800
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-05-14
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2044-05-13
AI Technical Summary
Existing systems using depth maps to detect onlookers can generate false alarms when someone passes by a user without engaging in spying behavior, leading to inaccurate security judgments.
A bystander detection system and method that utilizes a personnel detection module to obtain distance information from images and a bystander determination module to classify non-users into security categories based on their proximity and orientation relative to a device.
Enhances the accuracy of bystander detection by comprehensively assessing personnel in images, reducing false alarms and improving security judgments.
Smart Images

Figure TWG2TB001910150_001 
Figure TWG2TB001910150_002 
Figure TWG2TB001910150_003
Abstract
Description
[Technical Field]
[0001] This disclosure relates to the technical field of information display security. In particular, it relates to a technology that uses the characteristics of non-users in images to determine the security classification of non-users. [Previous Technology]
[0002] In recent years, with the increasing awareness of information security, many systems and applications have introduced the function of using depth maps to detect onlookers in order to ensure the privacy and security of users. However, when someone passes by a user from behind without engaging in any spying behavior, false alarms may occur because the judgment is made solely based on depth information in the depth map. [Summary of the Invention]
[0003] In view of the above, some embodiments of the present invention provide a bystander detection system and a bystander detection method to improve the problems of the prior art.
[0004] Some embodiments of the present invention provide a bystander detection system, comprising: a personnel detection module configured to receive images and, in response to the presence of personnel in the images, obtaining personnel information for each personnel, wherein the personnel information includes distance information relative to a device; and a bystander determination module configured to perform: determining, based on the distance information of the personnel information of each personnel, whether at least one non-user is within range; and, in response to the presence of at least one non-user within range, determining, based on the personnel information of each non-user, the security category to which each non-user belongs, wherein the security category includes a bystander category.
[0005] Some embodiments of the present invention provide a bystander detection method, comprising: receiving an image by a personnel detection module, and in response to the presence of at least one person in the image, obtaining personnel information of each person, wherein the personnel information includes distance information relative to a device; and being performed by a bystander determination module; determining, based on the distance information of the personnel information of each person, whether there is at least one non-user within range; and in response to the presence of a non-user within range, determining the security category to which each non-user belongs based on the personnel information of each non-user, wherein the security category includes a bystander category.
[0006] Based on the above, the bystander detection system and bystander detection method provided in the embodiments of the present invention obtain a variety of information through vision, so as to more comprehensively evaluate the situation of personnel in the images obtained by the camera and increase the accuracy of the judgment.
Implementation Method
[0008] The foregoing descriptions and other technical contents, features, and effects of this invention will be clearly presented in the following detailed description of the embodiments with reference to the accompanying drawings. Any actions that do not affect the effects and objectives achieved by this invention should still fall within the scope of the technical contents disclosed in this invention.
[0009] Figure 1 is a block diagram illustrating a bystander detection system according to some embodiments of the present invention. Figures 2 and 3 are schematic diagrams illustrating the operation of a personnel detection module according to some embodiments of the present invention. Referring to Figures 1 and 3, the bystander detection system 100 includes a personnel detection module 101 and a bystander determination module 102. The personnel detection module 101 is configured to receive an image 103 (e.g., image 200 in Figure 2 or image 300 in Figure 3), and in response to the presence of at least one person in the image 103 (e.g., persons 201-202 in image 200 and image 300 in Figure 3), obtains personnel information for each person, wherein the personnel information includes distance information relative to a device. The aforementioned device is, for example, the display screen of an electronic device, and the distance information relative to the device included in the aforementioned personnel information includes the distance of the person in the image relative to the display surface of the electronic device's display screen. The bystander judgment module 102 is configured to receive personnel information detected by the personnel detection module 101 and make further judgments and processing.
[0010] In some embodiments of the present invention, the personnel detection module 101 first determines which person in the image 103 is the user. The personnel detection module 101 can determine that the person closest to the device is the user, or it can first find multiple people closest to the device and then find the one closest to the center and determine it as the user. The present invention does not limit the method of determining the user. Taking the image 200 in FIG2 and the image 300 in FIG3 as examples, the personnel detection module 101 determines that person 201 is the user. In the image 103, all persons who are not users are referred to as non-users.
[0011] The following describes in detail, with reference to the drawings, how the bystander detection method and the various modules of the bystander detection system 100 of some embodiments of the present invention work together.
[0012] Figure 16 is a flowchart illustrating a bystander detection method according to some embodiments of the present invention. Please refer to Figures 1-3 and Figure 16. In the embodiment of Figure 16, the bystander detection method includes steps S1601-S1603. In step S1601, the personnel detection module 101 receives an image 103 and, in response to the presence of personnel in the image, obtains personnel information for each person, wherein the personnel information includes distance information relative to a device. In step S1602, the bystander determination module 102 determines, based on the distance information of the personnel information detected by the personnel detection module 101, whether there is a non-user within the range. Taking the aforementioned device as an example of an electronic device's display screen, the aforementioned range is set to a preset distance in front of the electronic device's display screen.
[0013] In step S1603, in response to determining that at least one non-user is within the aforementioned range, the bystander judgment module 102 determines the security category to which each non-user belongs based on the personnel information of each non-user, wherein the security category includes the bystander category. The bystander judgment module 102's determination that a non-user belongs to the bystander category indicates that the bystander judgment module 102 determines that this non-user poses a danger of being spied on.
[0014] Figure 17 is a flowchart illustrating a bystander detection method according to some embodiments of the present invention. Please refer to Figures 2-3 and Figure 17. In the embodiment of Figure 17, the aforementioned security classification includes a non-bystander category in addition to the bystander category. The bystander determination module 102 determines that the non-user belongs to the non-bystander category, indicating that the non-user does not pose a danger of viewing the device. The aforementioned step S1603 includes steps S1701-S1706. In step S1701, the bystander determination module 102 determines whether the current object among the non-users is facing the aforementioned device. If yes, step S1702 is executed; otherwise, step S1703 is executed. In step S1702, in response to the current object facing the device, the current object is determined to belong to the bystander category. In step S1703, in response to the current object not facing the device, the current object is determined to belong to the non-bystander category. For example, in Figures 2 and 3, the bystander judgment module 102 determines that person 201 is a user and person 202 is a non-user within the range. In Figure 2, the bystander judgment module 102 determines that person 202 belongs to the bystander category. In Figure 3, the bystander judgment module 102 determines that person 202 belongs to the non-bystander category. In step S1704, the bystander judgment module 102 determines whether there is an unselected object among the non-users. If yes, step S1706 is executed; otherwise, step S1705 is executed. In step S1705, since the bystander judgment module 102 has determined the security category to which all non-users belong, it exits the program. In step S1706, in response to the existence of an unselected object among the non-users, the bystander judgment module 102 selects at least one non-user that has not been selected as the current object and returns to step S1701.
[0015] Figures 4 and 5 are schematic diagrams illustrating the operation of a personnel detection module according to some embodiments of the present invention. Figure 18 is a flowchart illustrating a bystander detection method according to some embodiments of the present invention. Please refer to Figures 4-5 and Figure 18. In Figure 18, the aforementioned security classification includes, in addition to the bystander category, a passerby category and a sharing user category. The bystander judgment module 102 determines that a non-user belongs to the passerby category, indicating that although the non-user is within the scope of eavesdropping, they have no intention of eavesdropping. The bystander judgment module 102 determines that a non-user belongs to the sharing user category, indicating that the non-user is a person who shares information content with the user. The aforementioned step S1603 includes steps S1801 to S1808. In step S1801, the bystander judgment module 102 determines, for at least one non-user, whether the distance between the current object and the user is less than a preset distance. If yes, step S1803 is executed; otherwise, step S1802 is executed. In step S1803, in response to the distance between the current object and the user being less than a preset distance, it is determined that the current object belongs to the shared user category. In step S1802, in response to the distance between the current object and the user being not less than the preset distance, it is further determined whether the current object is facing the device. If yes, step S1804 is executed; otherwise, step S1805 is executed.
[0016] In some embodiments of the present invention, the bystander judgment module 102 calculates the distance between the current object and the user based on the distance information of each person relative to the device. In Figure 4, the distance between person 502, who is judged to be a non-user, and the distance between person 501, who is judged to be a user, and the distance between them is 45cm. Therefore, the bystander judgment module 102 judges the distance between person 502 and person 501 to be 50cm - 45cm = 5cm. In Figure 5, the distance between person 503, who is judged to be a non-user, and the distance between them is 150cm. Therefore, the bystander judgment module 102 judges the distance between person 503 and person 501 to be 150cm - 45cm = 105cm.
[0017] In step S1804, in response to the current object facing a device, it is determined that the current object belongs to the bystander category. In step S1805, in response to the current object not facing a device, it is determined that the current object belongs to the passerby category. In step S1806, the bystander determination module 102 determines whether there is an unselected object among the non-users. If yes, step S1808 is executed; otherwise, step S1807 is executed. In step S1807, since the bystander determination module 102 has determined the security category to which all non-users belong, it exits the program. In step S1808, in response to the existence of an unselected object among the non-users, the bystander determination module 102 selects one of the non-users that has not been selected as the current object and returns to step S1801.
[0018] Figure 19 is a flowchart illustrating a bystander detection method according to some embodiments of the present invention. In Figure 19, the aforementioned personnel information includes face information, angle information corresponding to the face information, and key point information. The aforementioned face information includes the position, height, and width of the face frame of the detected personnel (face frames 2011, 2021 in Figures 2-3, and face frames 5011, 5021, 5031 in Figures 4-5). If a personnel is detected but there is no position, height, and width of the face frame, it indicates that no personnel's face has been detected, which will be further explained in later embodiments. The aforementioned step of determining whether the current object is facing the device includes the bystander determination module 102 executing steps S1901-S1903. In step S1901, based on the face information of the current object, it is determined whether the face of the current object is detected. If yes, step S1902 is executed; if no, step S1903 is executed. In step S1902, in response to detecting a face of the current object, it is determined whether the current object is facing the device based on the angle information of the corresponding face information of the current object. In step S1903, in response to not detecting a face of the current object, it is determined whether the current object is facing the device based on the key point information of the current object.
[0019] Figure 8 is a schematic diagram illustrating the angle information of corresponding facial information according to some embodiments of the present invention. Referring to Figure 8, the head of the face in image 103 is shown as head 804. The head pitch angle of the face in image 103 is the angle of rotation about axis 801 with the center of the face of head 804 facing the device as a reference, and its range is [,]. The head yaw angle of the face in image 103 is the angle of rotation about axis 802 with the center of the face of head 804 facing the device as a reference, and its range is [,]. The head roll angle of the face in image 103 is the angle of rotation about axis 803 with the center of the face of head 804 facing the device as a reference, and its range is [,]. When the center of the face of head 804 is aligned with the device, the head pitch angle, head yaw angle, and head roll angle are 0 degrees.
[0020] The pitch angle of the gaze point of the face in image 103 is the angle of rotation about axis 801 with the gaze direction 805 of the head 804 facing the device as a reference, and its range is [,). The yaw angle of the gaze point of the face in image 103 is the angle of rotation about axis 802 with the gaze direction 805 of the head 804 facing the device as a reference, and its range is [,). The roll angle of the gaze point of the face in image 103 is the angle of rotation about axis 803 with the gaze direction 805 of the head 804 facing the device as a reference, and its range is [,). When the gaze direction 805 of the head 804 is aligned with the device, the pitch angle, yaw angle, and roll angle of the gaze point of the face are 0 degrees.
[0021] In some embodiments of the present invention, the angle information of the corresponding face information of the current object includes the head yaw angle. The aforementioned step S1902 includes: determining that the current object is facing the device in response to the head yaw angle being within the angle threshold range; and determining that the current object is not facing the device in response to the head yaw angle not being within the angle threshold range. Taking Figures 2 and 3 as examples, in Figures 2 and 3, pitch represents the head pitch angle, roll represents the head roll angle, and yaw represents the head yaw angle. The angle threshold range is [missing information]. In Figure 2, since the head yaw angle of person 202 is [missing information], the bystander judgment module 102 determines that person 202 is facing the device. In Figure 3, since the head yaw angle of person 202 is [missing information], the bystander judgment module 102 determines that person 202 is not facing the device.
[0022] In some embodiments of the present invention, the angle information of the corresponding face information of the current object includes the gaze point yaw angle. The aforementioned step S1902 includes: in response to the gaze point yaw angle being within the angle threshold range, determining that the current object is facing the device; and in response to the gaze point yaw angle not being within the angle threshold range, determining that the current object is not facing the device.
[0023] Figure 6 is a detection schematic diagram according to some embodiments of the present invention. Figure 7 is a key point schematic diagram according to some embodiments of the present invention. Referring to Figures 6 and 7, when image 103 includes image 600, the bystander judgment module 102 determines that the face of person 601 is detected based on the received face information, and the face range of person 601 is marked as shown by face frame 6011. At this time, the bystander judgment module 102 performs the aforementioned step S1902. However, if image 103 only includes image 6015, the bystander judgment module 102 performs step S1903 to determine whether the current object is facing the device based on the key point information of the current object. In Figure 7, key points of the human body include the center 701, left shoulder 705, right shoulder 702, left elbow 706, right elbow 703, left wrist 707, right wrist 704, left hip 711, right hip 708, left knee 712, right knee 709, left ankle 713, right ankle 710, nose 700, left ear 717, right ear 716, left eye 715, and right eye 714.
[0024] Figure 20 is a flowchart illustrating a bystander detection method according to some embodiments of the present invention. In Figure 20, the personnel detection module 101 detects the center 701, left shoulder 705, and right shoulder 702 of a person, obtaining the left shoulder coordinates of left shoulder point 6012, the right shoulder coordinates of right shoulder point 6014, and the center coordinates of center point 6013 as key point information. The aforementioned step S1903 includes steps S2001 to S2004. In step S2001, a first distance from the left shoulder coordinates to the center coordinates and a second distance from the right shoulder coordinates to the center coordinates are calculated. The first distance is divided by the distance of the current object relative to the device to obtain a first normalized distance, and the second distance is divided by the distance of the current object relative to the device to obtain a second normalized distance.
[0025] In step S2002, it is determined whether the absolute value of the difference between the first normalized distance and the second normalized distance is greater than the distance difference threshold (that is, whether the following inequality holds). If yes, proceed to step S2003; if no, proceed to step S2004. In step S2003, in response to the aforementioned inequality holding, it is determined that the current object is not facing a device. In step S2004, in response to the aforementioned inequality not holding, it is determined that the current object is facing a device.
[0026] Figure 9 is a schematic diagram of distance measurement according to some embodiments of the present invention. In Figure 9, the personnel detection module 101 includes an infrared laser diode 902 and an infrared image sensor 903. The infrared laser diode 902 emits infrared light towards the personnel 901 via direction 904, and the infrared image sensor 903 receives the reflected light via direction 905. The personnel detection module 101 calculates the distance between the personnel 901 and the device by the time difference between the emission and the reception of the reflected light, and uses this distance information in the personnel information of the personnel 901. It is worth noting that face distance estimation can also be achieved in other ways, such as using a TOF sensor or multiple sensors (based on phase difference method) to obtain a depth map, or through a method of estimation using a single-lens (Mono Camera) image.
[0027] FIG15 is a block diagram of an electronic device system according to some embodiments of the present invention. In FIG15, the electronic device 1500 includes a bystander detection system 100 and a display module 1501. The electronic device 1500 also includes a display screen, which is the device mentioned in the foregoing embodiments. The display module 1501 controls the display screen. The bystander detection method further includes: in response to determining that at least one non-user belongs to the bystander category, the bystander determination module 102 sends a signal to cause the device to start executing an anti-spying procedure. In some embodiments of the present invention, the bystander determination module 102 sends a signal to the display module 1501 to control the display screen, changing the display of "Secure" in FIG4 to the display of "Alert" in FIG5.
[0028] Figure 10 is a block diagram of a neural network module according to some embodiments of the present invention. Figure 21 is a flowchart of a bystander detection method according to some embodiments of the present invention. Referring to Figures 1, 10, and 21, in this embodiment, the personnel detection module 101 includes a neural network module 1000. The neural network module 1000 is configured to receive an image 103 and, in response to the presence of at least one person in the image, outputs a plurality of information tensors. The aforementioned step S1601 includes step S2101: the neural network module 1000 receives the image 103 and, in response to the presence of at least one person in the image 103, outputs a plurality of information tensors; and step S2102: the personnel detection module 101, in response to the presence of at least one person in the image 103, outputs personnel information for each person based on the aforementioned information tensors.
[0029] The following provides further description of various embodiments of the neural network module 1000. The neural network module 1000 includes an output feature tensor generation module 1001 and prediction modules 1002-1 to 1002-M, where M > 1. The output feature tensor generation module 1001 generates multiple output feature tensors of different sizes based on the image 103. Each of the prediction modules 1002-1 to 1002-M receives one of the aforementioned output feature tensors to generate an information tensor accordingly. The information tensor indicates face information, confidence score information, category information, angle information corresponding to the face information, and key point information. Based on all the information tensors generated by each of the prediction modules 1002-1 to 1002-M, the personnel detection module 101 outputs personnel information for each person based on the aforementioned information tensors in response to the presence of at least one person in the image 103.
[0030] Figure 22 is a flowchart illustrating a judgment method according to some embodiments of the present invention. Referring to Figure 22, the aforementioned step S2101 includes steps S2201 to S2202. In step S2201, the output feature tensor generation module 1001 generates multiple output feature tensors of different sizes based on the image 103. In step S2202, each of the prediction modules 1002-1 to 1002-M receives one of the aforementioned multiple output feature tensors to generate the aforementioned information tensors respectively. Each of the aforementioned information tensors is configured to indicate face information, confidence score information, category information, angle information corresponding to the face information, and key point information. The aforementioned face information includes the position, height, and width of the face bounding box of the detected person. If a person is detected but there is no position, height, and width of the face bounding box, it indicates that no person's face has been detected. The angle information corresponding to the face information includes the head pitch angle, head yaw angle, and head roll angle of the face. The key point information includes the coordinates of the left shoulder point, the right shoulder point, and the center point. It is worth noting that the angle information corresponding to the face information can be selected to include only the required angles, such as the head tilt angle, and the key point information can also include other key points shown in Figure 7 based on different applications.
[0031] Figure 11 is a block diagram of the output feature tensor generation module according to some embodiments of the present invention. Please refer to Figures 10 and 11, and the following description uses M=3. The output feature tensor generation module 1001 includes a backbone module 10011 and a feature pyramid module 10012.
[0032] In some embodiments of the present invention, the backbone module 10011 includes backbone layers 100111 to 100114 of different sizes. The backbone module 10011 generates multiple feature tensors of different sizes and with a first order based on the image 103 using the backbone layers 100111 to 100114. As illustrated in FIG11, the aforementioned multiple feature tensors are the output tensors of the backbone layers 100112 to 100114. The first order is the arrangement of these feature tensors according to their size from largest to smallest. The feature pyramid module 10012 performs feature fusion on the aforementioned feature tensors to obtain multiple output feature tensors.
[0033] Please refer to Figures 11 and 22. In some embodiments of the present invention, the aforementioned step S2201 includes the following steps: the backbone module 10011 generates multiple feature tensors of different sizes and having a first order based on the image 103 through the backbone layers 100111~100114, wherein the first order is the arrangement order of these feature tensors according to their size from largest to smallest; and the feature pyramid module 10012 performs feature fusion on the aforementioned feature tensors of different sizes and having a first order generated by the backbone module 10011 through the backbone layers 100111~100114 based on the image 103 to obtain an output feature tensor.
[0034] Please refer to Figure 11. The feature pyramid module 10012 includes fusion modules 100121-1 to 100121-2. The feature pyramid module 10012 performs the following steps to fuse the aforementioned feature tensors to obtain multiple output feature tensors.
[0035] First, the feature pyramid module 10012 sets the smallest feature tensor corresponding to the last one in the aforementioned first sequence as one of the temporary feature tensor sets. Taking the embodiment shown in FIG11 as an example, the smallest feature tensor is the output tensor of the backbone layer 100114, and the smallest feature tensor is stored in the temporary feature tensors 100122-3 as one of the temporary feature tensor sets.
[0036] Next, the feature pyramid module 10012 performs an upsampling operation on the temporary feature tensor 100122-3 via the fusion module 100121-1 to obtain an upsampled temporary feature tensor 100122-3 of the same size as the output tensor of the backbone layer 100113. The feature pyramid module 10012 then performs feature fusion on the upsampled temporary feature tensor 100122-3 and the output tensor of the backbone layer 100113 via the fusion module 100121-1 to obtain a temporary feature tensor 100122-2 of the same size as the output tensor of the convolutional layer of the backbone layer 100113. The feature pyramid module 10012 then uses the fusion module 100121-2 to fuse the sampled temporary feature tensor 100122-2 with the output tensor of the convolutional layer of the backbone layer 100112 to obtain a temporary feature tensor 100122-1 of the same size as the output tensor of the convolutional layer of the backbone layer 100112. The feature pyramid module 10012 outputs temporary feature tensors 100122-3, 100122-2, and 100122-1 as multiple output feature tensors of the aforementioned feature pyramid module 10012.
[0037] Figure 12 is a block diagram of a fusion module according to some embodiments of the present invention. In Figure 12, the structures of fusion modules 100121-1 to 100121-2 are as shown in fusion module 1200. Fusion module 1200 includes an upsampling module 1201, a pointwise convolutional layer 1202, and a pointwise addition module 1203. The upsampling module 1201 is configured to perform an upsampling operation on the input of the upsampling module 1201. The upsampling operation involves repeating elements twice in the height and width axes of the input of the upsampling module 1201 to double the size of the input of the upsampling module 1201. The pointwise convolutional layer 1202 performs a pointwise convolution operation. The pointwise addition module 1203 is configured to perform pointwise addition on the two received input tensors to obtain the output tensor of the pointwise addition module 1203. It is worth noting that the upsampling module 1201 can employ other upsampling methods.
[0038] Figure 13 is a schematic diagram of the prediction module structure according to some embodiments of the present invention. Figure 14 is a schematic diagram of the information tensor structure according to some embodiments of the present invention. Referring to Figures 13 and 14, the structure of prediction modules 1002-1 to 1002-3 is as shown in prediction module 1300. Prediction module 1300 includes t convolutional layers: convolutional layers 1301-1 to 1301-t, and one convolutional layer: convolutional layer 1302, where t is a positive integer; t is a positive integer representing the dimensions of the width and height axes of convolutional layers 1301-1 to 1301-t, t is a positive integer representing the number of detection boxes (anchors), and t is a positive integer. It is worth noting that the aforementioned convolutional layer indicates that the convolutional layer performs convolution operations on the input tensor with 128 convolution kernels, and the tensors obtained by performing convolution operations on the input tensor with 128 convolution kernels are stacked sequentially to obtain an output tensor with 128 channels along the width axis, 128 channels along the height axis, and 128 channels along the channel axis. Such an output tensor is a tensor with dimension 1.
[0039] The neural network module 1000 sets detection boxes of different sizes on the aforementioned multiple output feature tensors. The value of is 4 + 1 + the number of all categories + 3 + 6, where 4 represents the number of tensor elements required to describe the position coordinates of a vertex of the detection box and the detection width and height; 1 represents the probability of a detected target within the detection box and the accuracy of the detection box with one tensor element; 3 represents the number of tensor elements required to describe the head pitch angle, head yaw angle and head roll angle of the face; and 6 represents the number of tensor elements required to describe the coordinates of the left shoulder point, right shoulder point and center point (two tensors are required for each coordinate). The values of , , , t can be set by the user according to their needs. It is worth noting that since the sizes of the output feature tensors received by the prediction modules 1002-1 to 1002-M are different, the value of for each of the prediction modules 1002-1 to 1002-M is different.
[0040] The prediction module 1300 receives any one of the aforementioned multiple output feature tensors. After the output feature tensor passes through the convolutional layers 1301-1 to 1301-t and convolutional layer 1302 of the category prediction module 1300, an information tensor 1401 is obtained. The information tensor 1401 contains sub-information tensors 1401-1 to 1401-A, each of which corresponds to a detection box in the aforementioned detection boxes. Each sub-information tensor 1401-1 to 1401-A contains a multi-dimensional vector. As shown in Figure 14, each dimensional vector contains tensor elements 14021-1, 14021-2, 14022-1, 14022-2, 14023-1, 14023-2, 1403-1, 1403-2, 1403-3, 1404~1409. Tensor elements 14021-1, 14021-2, 14022-1, 14022-2, 14023-1, 14023-2, 1403-1, 1403-2, and 1403-3 respectively indicate the x-coordinate of the left shoulder point (e.g., left shoulder point 6012), the y-coordinate of the left shoulder point, the x-coordinate of the right shoulder point, the y-coordinate of the right shoulder point, the x-coordinate and y-coordinate of the center point, the head pitch angle, the head yaw angle, and the head roll angle of the face.
[0041] Tensor element 1404 contains multiple sub-tensor elements. Each sub-tensor element of tensor element 1404 indicates the probability that an object in the detection box belongs to each category. Tensor element 1405 indicates the confidence score, which represents the probability that a target is detected within the detection box and the accuracy of the detection box. Tensor element 1406 indicates the height of the detection box. Tensor element 1407 indicates the width of the detection box. Tensor elements 1408 and 1409 indicate the coordinates of the detection box. Among them, the face information includes the coordinates of the detection box, the height of the detection box, and the width of the detection box. The probability that an object in the aforementioned detection box belongs to each category is the aforementioned category information. The aforementioned confidence score is the aforementioned confidence score information. The angle information corresponding to the face information includes the head pitch angle, head yaw angle, and head roll angle of the face. The key point information includes the x-coordinate of the left shoulder point, the y-coordinate of the left shoulder point, the x-coordinate of the right shoulder point, the y-coordinate of the right shoulder point, the x-coordinate of the center point, and the y-coordinate of the center point. Personnel detection module 101 integrates all the information tensors generated by each of prediction modules 1002-1 to 1002-M to obtain personnel information for each person.
[0042] It is worth noting that the personnel detection module 101 integrates all the information tensors generated by each of the prediction modules 1002-1 to 1002-M to obtain the width and height of the face frame. In some embodiments of the present invention, the bystander detection system 100 captures the image 103 with a lens set at a fixed position on the device. Therefore, the width and height of the face frame are inversely proportional to the distance of the face relative to the lens (and simultaneously relative to the device). Thus, the personnel detection module 101 can obtain the distance of the face relative to the lens based on the width or height of the face frame.
[0043] It is worth noting that when training the neural network module 1000 in Figures 10-15, the trained neural network module 1000 can be obtained by adding the data of the head pitch angle, head yaw angle and head roll angle, left shoulder point coordinates, right shoulder point coordinates and center point coordinates of the human face to the training set and then training it using the object detection model training method.
[0044] In training figures 10-15, prediction modules 1002-1 to 1002-M are referred to as network heads in the technical field of this invention. The prediction modules 1002-1 to 1002-M disclosed in the foregoing embodiments can replace the network heads of other stages of object detection models, enabling those other stages of object detection models to output personnel information. This invention is not limited to the aforementioned backbone module 10011 and feature pyramid module 10012. [Simplified Explanation of the Diagram]
[0007] Figure 1 is a block diagram of a bystander detection system according to some embodiments of the present invention. Figures 2-5 are schematic diagrams of the operation of a personnel detection module according to some embodiments of the present invention. Figure 6 is a detection schematic diagram according to some embodiments of the present invention. Figure 7 is a schematic diagram of key points according to some embodiments of the present invention. Figure 8 is a schematic diagram of angle information corresponding to facial information according to some embodiments of the present invention. Figure 9 is a schematic diagram of distance measurement according to some embodiments of the present invention. Figure 10 is a block diagram of a neural network module according to some embodiments of the present invention. Figure 11 is a block diagram of an output feature tensor generation module according to some embodiments of the present invention. Figure 12 is a block diagram of a fusion module according to some embodiments of the present invention. Figure 13 is a schematic diagram of a prediction module structure according to some embodiments of the present invention. Figure 14 is a schematic diagram of an information tensor structure according to some embodiments of the present invention. Figure 15 is a block diagram of an electronic device system according to some embodiments of the present invention. Figures 16 to 22 are flowcharts illustrating bystander detection methods according to some embodiments of the present invention.
Claims
1. A bystander detection system, comprising: a personnel detection module configured to receive an image, and in response to the presence of at least one person in the image, obtaining personnel information for each of the at least one person, wherein, The personnel information includes distance information relative to a device; and a bystander judgment module configured to perform: (a) determining, based on the distance information of the personnel information of each of the at least one personnel, whether there is at least one non-user within a range among the at least one personnel; (b) In response to the presence of the at least one non-user in the range, based on the personnel information of each of the at least one non-user, determine a security category to which each of the at least one non-user belongs, wherein the security category includes a bystander category.
2. The bystander detection system as described in claim 1, wherein, The security classification includes a non-bystander category, and the aforementioned step (b) includes: (b1) for a current object among the at least one non-user, determining whether the current object is facing the device; in response to the current object being facing the device, determining that the current object belongs to the bystander category; in response to the current object not being facing the device, determining that the current object belongs to the non-bystander category; and (b2) in response to the existence of an unselected object among the at least one non-user, selecting one of the at least one non-users that has not been selected as the current object and returning to step (b1).
3. The bystander detection system as described in claim 1, wherein, The safety classification includes a passerby category and a shared user category. The aforementioned step (b) includes: (b1) for a current object among the at least one non-user, determining whether the distance between the current object and a user is less than a preset distance; in response to the distance between the current object and the user being less than the preset distance, determining that the current object belongs to the shared user category; in response to the distance between the current object and the user being not less than the preset distance, determining whether the current object is facing the device; in response to the current object facing the device, determining that the current object belongs to the bystander category; in response to the current object not facing the device, determining that the current object belongs to the passerby category; and (b2) in response to the existence of an unselected object among the at least one non-user, selecting one of the at least one non-users that has not been selected as the current object and returning to step (b1).
4. The bystander detection system as described in claim 2 or 3, wherein, The person information includes facial information, angle information corresponding to the facial information, and key point information. The aforementioned step of determining whether the current object is facing the device includes: (b11) determining whether the face of the current object is detected based on the facial information of the current object; (b12) in response to detecting the face of the current object, determining whether the current object is facing the device based on the angle information corresponding to the facial information of the current object; and (b13) in response to not detecting the face of the current object, determining whether the current object is facing the device based on the key point information of the current object.
5. The bystander detection system as described in claim 4, wherein, The angle information corresponding to the face information of the current object includes a head yaw angle. The aforementioned step (b12) includes: in response to the head yaw angle being within an angle threshold range, determining that the current object is facing the device; and in response to the head yaw angle not being within the angle threshold range, determining that the current object is not facing the device.
6. The bystander detection system as described in claim 4, wherein, The angle information of the face information of the current object includes a gaze yaw angle. The aforementioned step (b12) includes: in response to the gaze yaw angle being within an angle threshold range, determining that the current object is facing the device; and in response to the gaze yaw angle not being within the angle threshold range, determining that the current object is not facing the device.
7. The bystander detection system as described in claim 4, wherein, The key point information of the current object includes a left shoulder point coordinate, a right shoulder point coordinate, and a center point coordinate; the aforementioned step (b13) includes: (b131) calculating a first distance from the left shoulder point coordinate to the center point coordinate and calculating a second distance from the right shoulder point coordinate to the center point coordinate; dividing the first distance by a distance of the current object relative to the device to obtain a first normalized distance, dividing the second distance by the distance of the current object relative to the device to obtain a second normalized distance, determining whether an absolute value of a difference between the first normalized distance and the second normalized distance is greater than a distance difference threshold; and (b132) in response to the absolute value of the difference between the first normalized distance and the second normalized distance not being greater than the distance difference threshold, determining that the current object is facing the device; And in response to the absolute value of the difference between the first normalized distance and the second normalized distance being greater than the distance difference threshold, it is determined that the current object is not facing the device.
8. The bystander detection system as described in claim 1, wherein, The bystander determination module is configured to perform the following: in response to determining that at least one of the non-users belongs to the bystander category, a signal is sent to cause the device to begin executing an anti-spying procedure.
9. The bystander detection system as described in claim 1, wherein, The personnel detection module includes a neural network module configured to receive the image and, in response to the presence of at least one person in the image, output multiple information tensors. The personnel detection module is configured to, in response to the presence of at least one person in the image, output personnel information for each of the at least one person based on the information tensors.
10. A bystander detection method, comprising: receiving an image by a personnel detection module, and in response to the presence of at least one person in the image, obtaining personnel information for each of the at least one person, wherein, The personnel information includes distance information relative to a device; and is executed by an observer judgment module; (a) based on the distance information of the personnel information of each of the at least one personnel, determine whether there is at least one non-user within the range of the at least one personnel; (b) In response to the presence of the at least one non-user in the range, based on the personnel information of each of the at least one non-user, determine a security category to which each of the at least one non-user belongs, wherein the security category includes a bystander category.
Citation Information
Patent Citations
Integration of head mounted displays with public display devices
TW201514847A
Lid controller hub
TW202227999A
Bystander-centric privacy controls for recording devices
TW202305661A