Pedestrian flow monitoring method and system based on privacy protection
By calculating the average optical flow map in traditional camera videos to generate an initial mask, and combining the clustering and noise reduction technology of minimal cameras, the problems of privacy leakage and low accuracy in traditional traffic monitoring are solved, and efficient and accurate privacy protection traffic monitoring is achieved.
Patent Information
- Application Number
- CN202510079853.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional traffic monitoring technology is prone to the risk of privacy leakage, and the mask training process and noise of minimal cameras have a great impact, resulting in a decrease in monitoring accuracy.
By acquiring videos with a traditional camera, multiple average optical flow patterns are calculated to generate an initial mask, and the mask is obtained through training, and attached to the detector of the minimalist camera; the detector values are clustered in advance under different lighting conditions, and noise reduction is used by clustering to improve monitoring accuracy.
It effectively reduces the mask training time and sample set number, reduces the impact of noise, and improves the accuracy of traffic monitoring and privacy protection capabilities.
Smart Images

Figure CN119942457A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a method and system for monitoring human traffic based on privacy protection. Background Art
[0002] The statistics of passenger flow play a great role in commercial activities and public health. For example, by analyzing the passenger flow in different time periods and different areas, merchants can more accurately grasp the consumption behavior and preferences of customers, so as to formulate more effective business strategies and increase sales and customer satisfaction. By real-time monitoring of the density of passenger flow in roads, stations, airports and other places, the transportation department can adjust traffic signals, optimize public transportation routes, and predict congestion in time, thereby improving traffic efficiency and reducing congestion and delays. In large public places, by monitoring the density of passenger flow, abnormal crowding can be detected in time to prevent stampedes; in places such as border ports and airports, by monitoring the flow of people, security inspections can be strengthened. Passenger flow monitoring provides important decision-making support for all walks of life, helping to improve operational efficiency, improve service quality, and ensure public safety.
[0003] However, traditional crowd flow monitoring technology often involves the collection and analysis of individual information, which will cause serious privacy leakage risks. Traditional monitoring methods, such as camera-based video surveillance systems, can provide relatively accurate crowd flow data, but they also record a large amount of sensitive information such as facial images and behavioral trajectories. Once these data are abused or leaked, they will seriously infringe on personal privacy and may even lead to adverse consequences such as identity theft, stalking and harassment. Compared with traditional cameras, minimalist cameras have only a few detectors and will not capture detailed information such as faces. They can effectively protect user privacy. Minimalist cameras need to be used in conjunction with masks. Existing masks require a long period of training. In addition, since minimalist cameras have very few detectors, they are easily affected by noise, which affects the accuracy of crowd flow monitoring. Summary of the invention
[0004] In order to speed up the generation of a minimalist camera mask and reduce noise, a first aspect of the present invention provides a method for monitoring human flow based on privacy protection, the method comprising: A traditional camera is used to obtain a video of a pedestrian flow area to be monitored under different lighting conditions, a plurality of average optical flow maps of the pedestrian flow area video are calculated, a plurality of initial masks are generated based on the plurality of average optical flow maps, masks are obtained according to the initial mask training, and the masks are attached to different detectors of a minimalist camera; Use a simple camera to shoot the pedestrian flow area under different lighting conditions in advance to obtain the value of each detector, and cluster the value of each detector; A minimalist camera is used to shoot the pedestrian flow area to obtain the current value of each detector, and the cluster closest to the current distance is found from the cluster corresponding to the detector. The current value is enhanced using the cluster closest to the current distance, and all the current values enhanced by the minimalist camera are input into the neural network to obtain the pedestrian flow value.
[0005] Preferably, the step of calculating a plurality of average optical flow maps of the video of the pedestrian traffic area is specifically as follows: Obtain the number of detectors in the minimalist camera, and divide the video into groups with the number of detectors according to the lighting conditions; For each group of videos, the average optical flow map of the videos in the group is calculated.
[0006] Preferably, the generating of multiple initial masks based on the multiple average optical flow maps is specifically: Normalize each average optical flow map separately, and convert the average optical flow map to the mask image size by downsampling; The value of the average optical flow map is used as the projection rate of the mask image to obtain the initial mask.
[0007] Preferably, the step of obtaining the mask according to the initial mask training is specifically as follows: Marking frames according to the number of people in the frames of the video of the pedestrian flow area; setting the mask to the initial mask; Convert the frame to the size of the mask, multiply the transmittance of each point of the frame by the point at the corresponding position of the mask and add them up to obtain the detector value, input all the detector values into the neural network to obtain the number of people, and update the mask and the neural network parameters according to the obtained number of people and the marked number of people; The mask and the neural network parameters are continuously updated until preset conditions are met.
[0008] Preferably, the enhancing the current value by using the cluster with the closest distance is specifically as follows: Add the current value to the closest cluster, sort the values in the cluster according to the distance from the cluster center, record the position of the value to be enhanced in the sorting, and enhance the sorted value using wavelet denoising; The enhanced value is obtained from the position of the enhanced sequence.
[0009] In a second aspect of the present invention, a human flow monitoring system based on privacy protection is provided, the system comprising: A mask acquisition module, used to use a traditional camera to acquire a video of a pedestrian flow area to be monitored under different lighting conditions, calculate a plurality of average optical flow maps of the pedestrian flow area video, generate a plurality of initial masks based on the plurality of average optical flow maps, obtain masks according to the initial mask training, and attach the masks to different detectors of the minimalist camera; The clustering module is used to obtain the value of each detector by photographing the pedestrian flow area with a minimalist camera under different lighting conditions in advance, and cluster the value of each detector; The pedestrian flow monitoring module is used to use a minimalist camera to shoot the pedestrian flow area to obtain the current value of each detector, find the cluster closest to the current distance from the cluster corresponding to the detector, use the cluster closest to the current distance to enhance the current value, and input all the current values enhanced by the minimalist camera into the neural network to obtain the pedestrian flow value.
[0010] Preferably, the step of calculating a plurality of average optical flow maps of the video of the pedestrian traffic area is specifically as follows: Obtain the number of detectors in the minimalist camera, and divide the video into groups with the number of detectors according to the lighting conditions; For each group of videos, the average optical flow map of the videos in the group is calculated.
[0011] Preferably, the generating of multiple initial masks based on the multiple average optical flow maps is specifically: Normalize each average optical flow map separately, and convert the average optical flow map to the mask image size by downsampling; The value of the average optical flow map is used as the projection rate of the mask image to obtain the initial mask.
[0012] Preferably, the step of obtaining the mask according to the initial mask training is specifically as follows: Marking frames according to the number of people in the frames of the video of the pedestrian flow area; setting the mask to the initial mask; Convert the frame to the size of the mask, multiply the transmittance of each point of the frame by the point at the corresponding position of the mask and add them up to obtain the detector value, input all the detector values into the neural network to obtain the number of people, and update the mask and the neural network parameters according to the obtained number of people and the marked number of people; The mask and the neural network parameters are continuously updated until preset conditions are met.
[0013] Preferably, the enhancing the current value by using the cluster with the closest distance is specifically as follows: Add the current value to the closest cluster, sort the values in the cluster according to the distance from the cluster center, record the position of the value to be enhanced in the sorting, and enhance the sorted value using wavelet denoising; The enhanced value is obtained from the position of the enhanced sequence.
[0014] In a third aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method provided in any one of the embodiments of the first aspect.
[0015] In a fourth aspect of the present invention, a computing device is provided, comprising a memory and a processor. The memory stores executable code, and when the processor executes the executable code, the method provided in any one of the embodiments of the first aspect is implemented.
[0016] The minimalist camera has a small number of detectors and can protect user privacy. However, the mask training process and noise have a great impact on the minimalist camera's monitoring of human flow. The present invention generates multiple initial masks according to the optical flow of the video, and then trains to obtain the masks, which can reduce the mask training time and the number of sample sets. In addition, the values of the same detector of the minimalist camera under different lighting conditions are clustered, and the cluster is used to reduce the noise of the current value, thereby improving the accuracy of the detector and the accuracy of human flow. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a flow chart of Embodiment 1; Figure 2 This is a schematic diagram of the principle of a minimalist camera detector and mask; Figure 3 Flowchart for mask training; Figure 4 Schematic diagram of the mask and neural network structure. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0019] In the following description, the terms "first\second" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0021] Embodiment 1, as Figure 1 A method for monitoring human traffic based on privacy protection is provided, the method comprising: S110, using a traditional camera to obtain a video of a pedestrian flow area to be monitored under different lighting conditions, calculating a plurality of average optical flow maps of the video of the pedestrian flow area, generating a plurality of initial masks based on the plurality of average optical flow maps, obtaining masks according to initial mask training, and attaching the masks to different detectors of a minimalist camera; In traditional digital cameras, the sensor is composed of thousands or even millions of small photosensitive elements, or detectors, such as CMOS sensors or CCD sensors. Each photosensitive element corresponds to a pixel in the image. The color information of each pixel is usually expressed by the different intensities of the three colors red, green, and blue (RGB). This information is converted into digital data and finally forms a digital image file.
[0022] Minimalist Cameras (mincams) have only a few photosensitive elements, for example, 1-64. A mask is used in front of each photosensitive element. The mask is a small picture that can be printed on film or other materials. Different parts of the mask have different transmittances, ranging from 0 to 1. If it is 0, it means that it is completely opaque, and if it is 1, it means that it is completely transparent. The value obtained by the photosensitive element is the sum of the product of the projection of the external image on the mask and the transmittance, such as Figure 2 This is because the minimalist camera has very few photosensitive elements and will not take very clear pictures, thus protecting privacy.
[0023] In pedestrian flow monitoring, you only need to get the value of the photosensitive element of the minimalist camera, and then input it into the trained neural network to get the pedestrian flow value. Among them, the key lies in the generation of the mask, that is, the generation of the mask image. Since the mask is generated by the mask image, it is uniformly expressed as a mask in the following introduction. Before using the minimalist camera, a video shot with a traditional camera is used to generate a mask. In order to be able to use a small number of samples and speed up the training speed, the optical flow of the video shot with a traditional camera is used to generate the initial mask. In one embodiment, the calculation of multiple average optical flow maps of the video of the pedestrian flow area is specifically: The number of detectors in the minimalist camera is obtained, and the video is divided into groups of the number of detectors according to the lighting conditions; for each group of videos, the average optical flow map of the videos in the group is calculated. Since different lighting has a great influence on the value of the detector, especially after adding a mask in front of the detector, the video is divided into groups of the number of detectors according to the lighting conditions. For example, if the minimalist camera has 6 CMOS, the video is divided into 6 groups according to the lighting conditions, and the lighting conditions of each group are different. Specifically, the videos are divided into 6 groups according to the average brightness of each video and the average brightness of all videos. For example, there are 30 videos, each of which is 5 minutes long and is shot at different times of the day or monitoring time periods. The average brightness of each video in these 30 videos is calculated and divided into 6 groups. For each group of videos, the optical flow map of the frames in the video is calculated, and then the optical flow map of this video is calculated. The average value of the optical flow maps of all videos in this group of videos is used as the average optical flow map of the videos in the group. For example, if a video group has 100 optical flow maps, the average value of the same position of these 100 optical flow maps is calculated to obtain the value of the corresponding position of the average optical flow map. If the person moves more, the average value of the optical flow is larger, and vice versa.
[0024] After obtaining the average optical flow map of each group of videos, the average optical flow map corresponding to this group of videos is normalized, and the average optical flow map is converted to the size of the mask image by downsampling, and the value of the optical flow map is used as the transmittance of the mask image, so that the initial mask is obtained. If there are 6 detectors, 6 initial masks will be obtained.
[0025] The mask is an image. After obtaining the initial mask, the mask needs to be further trained to obtain the final mask that covers or masks the detector, such as Figure 3 As shown, specifically: S111, marking frames according to the number of people in the frames of the video of the pedestrian flow area; setting the mask to the initial mask; S112, converting the frame to the size of the mask, multiplying the transmittance of each point of the frame and the point at the corresponding position of the mask and then accumulating them to obtain the detector value, inputting all the detector values into the neural network to obtain the number of people, and updating the mask and the neural network parameters according to the obtained number of people and the marked number of people; S113, continuously updating the mask and the neural network parameters until preset conditions are met.
[0026] Label each frame of the video manually or by using other methods such as existing mature people counting algorithms, and record the actual number of people in each frame. At the beginning of training, use the initial mask as the mask, and then convert the video frame to the mask size. For example, if the video frame is 1024×1024 and the mask image size is 128×128, then convert the video frame to the same size as the mask, 128×128, by downsampling or projection. At this time, the mask and video frame are the same size. After multiplying the mask and the video frame at the same position, accumulate all the multiplication results as the detector value, and input all the detector values into the neural network to obtain the number of people, such as Figure 4 As shown, the estimated number of people output by the neural network is compared with the real number of people marked before, the error is calculated, and the back propagation algorithm is used to adjust the parameters of the neural network and the transmittance of the mask according to the error. More specifically, the mask and frame are flattened into one-dimensional vectors respectively, where the values in the one-dimensional vector of the mask are learnable, and the inner product of the two, or the transpose of the one-dimensional vector of the frame and the one-dimensional vector of the mask, is multiplied to obtain the value of the detector. When the training is completed, the one-dimensional vector of the mask is restored to the mask. The training of the mask and the neural network is realized by continuously updating the mask and the neural network, and the training is stopped when the preset conditions are met. Among them, the preset conditions are that a certain number of training iterations is reached, the accuracy of the estimated number of people reaches a certain threshold, or the error change becomes very small. It should be noted that during the training process, the value of the mask needs to be controlled between [0,1]. After obtaining the mask image, a mask is generated by printing or other methods, and the mask is covered or masked in front of the detector.
[0027] S120, using a minimalist camera to photograph a pedestrian flow area under different lighting conditions in advance to obtain a value of each detector, and clustering the value of each detector; Use a minimalist camera to capture videos or images of pedestrian areas under different lighting conditions, such as daytime, nighttime, cloudy, and sunny days, and record the output value of each detector under different lighting and different number of people. Use clustering algorithms such as K-means and DBSCAN to cluster the values of each detector, and classify similar detector values into the same category. Different categories represent different lighting or number of people. For example, when the light is dim, the output value of a detector may be concentrated in a specific range, and these values can be grouped into one category through clustering. Clustering indicators include, but are not limited to, lighting conditions and number of people, that is, the value of each detector is clustered according to lighting conditions and / or number of people.
[0028] S130, using a minimalist camera to shoot the pedestrian flow area to obtain the current value of each detector, finding the cluster closest to the current distance from the cluster corresponding to the detector, using the cluster closest to the current value to enhance the current value, and inputting all the current values enhanced by the minimalist camera into the neural network to obtain the pedestrian flow value.
[0029] When it is necessary to use a minimalist camera to monitor the flow of people in a pedestrian area, obtain the current value of each detector. For example, if the minimalist camera has 6 detectors, each of the 6 detectors will have a current value, and each detector will have a cluster in S120. Determine which cluster the current value is in. If the current value of the second detector is in the third cluster of the second detector, add the current value to the closest cluster, sort the values in the cluster according to the distance from the cluster center, and record the position of the value to be enhanced in the sorting, and enhance the sorted value by wavelet denoising; obtain the enhanced value from the position of the enhanced sequence. Classify the current detector reading into the cluster that is most similar to it, and select the closest cluster by calculating the distance between the current value and the center points of all clusters, such as the Euclidean distance, and temporarily add the current value to the cluster. Sort all values in the new cluster containing the current value in ascending order according to their distance from the center point of the cluster. Record the index position of the current value in the sorted sequence. The subsequent wavelet denoising operation will change the size of all values in the sequence, but will not change their relative positions. The purpose of recording the index position information is to accurately find and extract the enhanced value corresponding to the original current value after denoising. Apply wavelet denoising technology to the sorted value sequence. Wavelet denoising decomposes the signal into components of different frequencies and removes the high-frequency noise components, thereby achieving the purpose of smoothing the signal. The wavelet denoising process will make the adjacent values smoother, so that the current value is more in line with the overall distribution trend of the cluster. In the sequence after wavelet denoising, according to the previously recorded position information, take out the value at the corresponding position, which is the enhanced current detector value. After obtaining the enhanced value, all detector values are used as the input of the neural network to obtain the number of people information. Here, the neural network has been trained in the mask training, and the enhanced value can be directly used as input.
[0030] Embodiment 2 provides a human flow monitoring system based on privacy protection, the system comprising: A mask acquisition module, used to use a traditional camera to acquire a video of a pedestrian flow area to be monitored under different lighting conditions, calculate a plurality of average optical flow maps of the pedestrian flow area video, generate a plurality of initial masks based on the plurality of average optical flow maps, obtain masks according to the initial mask training, and attach the masks to different detectors of the minimalist camera; The clustering module is used to obtain the value of each detector by photographing the pedestrian flow area with a minimalist camera under different lighting conditions in advance, and cluster the value of each detector; The pedestrian flow monitoring module is used to use a minimalist camera to shoot the pedestrian flow area to obtain the current value of each detector, find the cluster closest to the current distance from the cluster corresponding to the detector, use the cluster closest to the current distance to enhance the current value, and input all the current values enhanced by the minimalist camera into the neural network to obtain the pedestrian flow value.
[0031] Preferably, the step of calculating a plurality of average optical flow maps of the video of the pedestrian traffic area is specifically as follows: Obtain the number of detectors in the minimalist camera, and divide the video into groups with the number of detectors according to the lighting conditions; For each group of videos, the average optical flow map of the videos in the group is calculated.
[0032] Preferably, the generating of multiple initial masks based on the multiple average optical flow maps is specifically: Normalize each average optical flow map separately, and convert the average optical flow map to the mask image size by downsampling; The value of the average optical flow map is used as the projection rate of the mask image to obtain the initial mask.
[0033] Preferably, the step of obtaining the mask according to the initial mask training is specifically as follows: Marking frames according to the number of people in the frames of the video of the pedestrian flow area; setting the mask to the initial mask; Convert the frame to the size of the mask, multiply the transmittance of each point of the frame by the point at the corresponding position of the mask and add them up to obtain the detector value, input all the detector values into the neural network to obtain the number of people, and update the mask and the neural network parameters according to the obtained number of people and the marked number of people; The mask and the neural network parameters are continuously updated until preset conditions are met.
[0034] Preferably, the enhancing the current value by using the cluster with the closest distance is specifically as follows: Add the current value to the closest cluster, sort the values in the cluster according to the distance from the cluster center, record the position of the value to be enhanced in the sorting, and enhance the sorted value using wavelet denoising; The enhanced value is obtained from the position of the enhanced sequence.
[0035] Embodiment 3 provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method provided in the above embodiment 1.
[0036] Embodiment 4 provides a computing device, including a memory and a processor. The memory stores executable code, and when the processor executes the executable code, the method provided in the above embodiment 1 is implemented.
[0037] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the multiple embodiments disclosed in this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0038] The specific implementation methods described above further illustrate in detail the purposes, technical solutions and beneficial effects of the multiple embodiments disclosed in this specification. It should be understood that the above description is only the specific implementation methods of the multiple embodiments disclosed in this specification, and is not used to limit the protection scope of the multiple embodiments disclosed in this specification. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solutions of the multiple embodiments disclosed in this specification should be included in the protection scope of the multiple embodiments disclosed in this specification.
Claims
1. A method for monitoring human flow based on privacy protection, characterized in that: The method comprises: A traditional camera is used to obtain a video of a pedestrian flow area to be monitored under different lighting conditions, a plurality of average optical flow maps of the pedestrian flow area video are calculated, a plurality of initial masks are generated based on the plurality of average optical flow maps, masks are obtained according to the initial mask training, and the masks are attached to different detectors of a minimalist camera; Use a simple camera to shoot the pedestrian flow area under different lighting conditions in advance to obtain the value of each detector, and cluster the value of each detector; A minimalist camera is used to shoot the pedestrian flow area to obtain the current value of each detector, and the cluster closest to the current distance is found from the cluster corresponding to the detector. The current value is enhanced using the cluster closest to the current distance, and all the current values enhanced by the minimalist camera are input into the neural network to obtain the pedestrian flow value.
2. The method according to claim 1, characterized in that The calculation of multiple average optical flow maps of the video of the pedestrian traffic area is specifically as follows: Obtain the number of detectors in the minimalist camera, and divide the video into groups with the number of detectors according to the illumination; For each group of videos, the average optical flow map of the videos in the group is calculated.
3. The method according to claim 1, characterized in that The generating of multiple initial masks based on the multiple average optical flow maps is specifically: Normalize each average optical flow map separately, and convert the average optical flow map to the mask image size by downsampling; The value of the average optical flow map is used as the projection rate of the mask image to obtain the initial mask.
4. The method according to claim 1, characterized in that The mask is obtained according to the initial mask training, specifically: Marking frames according to the number of people in the frames of the video of the pedestrian flow area; setting the mask to the initial mask; Convert the frame to the size of the mask, multiply the transmittance of each point of the frame by the point at the corresponding position of the mask and add them up to obtain the detector value, input all the detector values into the neural network to obtain the number of people, and update the mask and the neural network parameters according to the obtained number of people and the marked number of people; The mask and the neural network parameters are continuously updated until preset conditions are met.
5. The method according to claim 1, characterized in that The method of enhancing the current value by using the cluster closest to the current value is specifically as follows: Add the current value to the closest cluster, sort the values in the cluster according to the distance from the cluster center, record the position of the value to be enhanced in the sorting, and enhance the sorted value using wavelet denoising; The enhanced value is obtained from the position of the enhanced sequence.
6. A human flow monitoring system based on privacy protection, characterized in that: The system comprises: A mask acquisition module, used to use a traditional camera to acquire a video of a pedestrian flow area to be monitored under different lighting conditions, calculate a plurality of average optical flow maps of the pedestrian flow area video, generate a plurality of initial masks based on the plurality of average optical flow maps, obtain masks according to the initial mask training, and attach the masks to different detectors of the minimalist camera; The clustering module is used to obtain the value of each detector by photographing the pedestrian flow area with a minimalist camera under different lighting conditions in advance, and cluster the value of each detector; The pedestrian flow monitoring module is used to use a minimalist camera to shoot the pedestrian flow area to obtain the current value of each detector, find the cluster closest to the current distance from the cluster corresponding to the detector, use the cluster closest to the current distance to enhance the current value, and input all the current values enhanced by the minimalist camera into the neural network to obtain the pedestrian flow value.
7. The system according to claim 6, characterized in that The calculation of multiple average optical flow maps of the video of the pedestrian traffic area is specifically as follows: Obtain the number of detectors in the minimalist camera, and divide the video into groups with the number of detectors according to the illumination; For each group of videos, the average optical flow map of the videos in the group is calculated.
8. The system according to claim 6, characterized in that The generating of multiple initial masks based on the multiple average optical flow maps is specifically: Normalize each average optical flow map separately, and convert the average optical flow map to the mask image size by downsampling; The value of the average optical flow map is used as the projection rate of the mask image to obtain the initial mask.
9. The system according to claim 6, characterized in that The mask is obtained according to the initial mask training, specifically: Marking frames according to the number of people in the frames of the video of the pedestrian flow area; setting the mask to the initial mask; Convert the frame to the size of the mask, multiply the transmittance of each point of the frame by the point at the corresponding position of the mask and add them up to obtain the detector value, input all the detector values into the neural network to obtain the number of people, and update the mask and the neural network parameters according to the obtained number of people and the marked number of people; The mask and the neural network parameters are continuously updated until preset conditions are met.
10. The system according to claim 6, characterized in that The method of enhancing the current value by using the cluster closest to the current value is specifically as follows: Add the current value to the closest cluster, sort the values in the cluster according to the distance from the cluster center, record the position of the value to be enhanced in the sorting, and enhance the sorted value using wavelet denoising; The enhanced value is obtained from the position of the enhanced sequence.