Crowd density identification method and system based on dynamic optimization diffusion model
Through the combination of layer-by-layer diffusion simulation, enhanced learning correction and potential spatial dynamic optimization, the problem of accuracy and resource limitation in complex environments of population density recognition is solved, and efficient and flexible identification results are achieved.
Patent Information
- Application Number
- CN202510489714.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing population density recognition methods are not accurate in complex environments and are difficult to operate efficiently in resource-constrained environments, mainly due to image quality attenuation and excessive demand for computing resources.
Layer-by-layer diffusion simulation is used to simulate image quality attenuation, high-quality images are restored through enhanced learning correction, and model parameters are dynamically optimized in the latent space, and population density recognition is used using dynamic optimization diffusion model.
It improves the accuracy and flexibility of crowd density identification, reduces the demand for computing resources, is suitable for variable and complex monitoring scenarios, and maintains efficient operation.
Smart Images

Figure CN120356154A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and machine learning, and in particular to a crowd density recognition method and system based on a dynamic optimization diffusion model. Background Art
[0002] In the field of crowd density recognition, existing recognition methods mainly rely on traditional image processing and machine learning, which usually have high requirements on the quality of input images. In practical applications, such as urban monitoring or public security monitoring systems, images often degrade in quality due to environmental factors such as lighting changes, occlusions, or camera movement. These changes can seriously affect the accuracy and reliability of recognition. In addition, existing methods often have high requirements for computing resources when processing large amounts of real-time video data, which is particularly prominent in resource-constrained environments.
[0003] The main problems faced by existing methods include:
[0004] 1. Under complex environmental conditions (such as unstable lighting, high dynamic background, etc.), image quality degradation will lead to inaccurate crowd density estimation.
[0005] 2. Existing systems have difficulty operating efficiently in resource-constrained environments because they usually require a lot of computing resources to process and analyze image data. Summary of the invention
[0006] The object of the present invention is to provide a crowd density recognition method and system based on a dynamic optimization diffusion model, so as to solve the above-mentioned problems existing in the prior art.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] A crowd density recognition method based on a dynamic optimization diffusion model comprises the following steps:
[0009] S1, layer-by-layer diffusion simulation: by gradually increasing the noise of the image, simulating the gradual decay process of the image data, and obtaining a noisy image;
[0010] S2, enhanced learning correction: using the denoising process and image quality preservation strategy to recover high-quality crowd density maps from noisy images;
[0011] S3. Dynamic optimization of latent space: Based on the crowd density map, the parameters of the crowd density recognition model are dynamically adjusted in the latent space to obtain the crowd density recognition model under the optimal parameter state.
[0012] Preferably, step S1 is specifically as follows: gradually introduce noise into the image data to simulate the decline of image quality in a real environment, so as to train the dynamic optimization diffusion model to adapt to the impact of interference factors that may be encountered in the actual monitoring scenario on the quality of image data; the calculation formula is,
[0013]
[0014] where, I t and I t-1 are the noisy images at the t-th and (t - 1)-th time steps respectively; N(0, 1) is a random variable subject to a standard normal distribution with a mean of 0 and a standard deviation of 1, which is used to simulate the random noise added in each step; k t is the diffusion coefficient at time step t; dt is the time increment.
[0015] Preferably, in the layer-by-layer diffusion simulation, starting from the initial clear image, gradually introduce random noise, and the noise in each step is calculated according to the diffusion coefficient, and the diffusion coefficient is dynamically adjusted according to the environmental changes in the actual monitoring scenario.
[0016] Preferably, step S2 is specifically as follows: determine the total loss function based on the denoising function and the regularization function, and use the total loss function to recover a high-quality denoised crowd density map from the noisy image; the calculation formula is,
[0017]
[0018] where, L is the total loss function; D(I t ) is the denoising function; ‖I t - D(I t )‖ 2 is the mean square error between the crowd density map and the noisy image; λ is the regularization coefficient; R(I t ) is the regularization function; T is the total number of time steps.
[0019] Preferably, the denoising function is obtained by training a deep learning model through a corresponding training dataset; the denoising function can learn to identify and eliminate the noise components in the image while keeping the original features and quality of the image undamaged.
[0020] Preferably, the regularization function is obtained by training a deep learning model through a corresponding training dataset; the regularization function can be automatically adjusted based on the content complexity of the image, and it provides additional constraints to maintain the important details and quality of the noisy image.
[0021] Preferably, step S3 is specifically as follows: use the total reward function in the latent space to optimize the parameters of the crowd density recognition model in the latent space to make it adapt to various monitoring environments and crowd density changes; the calculation formula is,
[0022]
[0023] wherein, R is the total reward function of the latent space; x j is the j-th feature in the latent space, that is, the data abstracted and extracted from the crowd density map; θ is the parameter of the crowd density recognition model; f is the reward function, which is used to adjust the behavior of the crowd density recognition model in the latent space to maximize the accuracy and efficiency of model recognition; N is the total number of features in the latent space.
[0024] Preferably, the dynamic optimization of the latent space utilizes the feature learning theory in machine learning to extract the deep features of the crowd density map through an autoencoder or a generative adversarial network. These features are converted into a multi-dimensional latent space, where each dimension captures certain key aspects of the crowd density map; in the latent space, the model does not directly operate on the crowd density map, but on higher-level abstract features, thereby enabling the model to learn and optimize more flexibly and effectively.
[0025] Preferably, after step S3, it further includes
[0026] S4. Crowd density estimation: Use the dynamically optimized diffusion model in the optimal parameter state to perform crowd density recognition and obtain the final crowd density estimation result.
[0027] The purpose of the present invention also lies in providing a crowd density recognition system based on a dynamically optimized diffusion model. The recognition system can implement the above-mentioned method. The recognition system includes
[0028] Layer-by-layer diffusion simulation module: By gradually increasing the noise of the image, simulate the process of gradual decay of image data to obtain a noisy image;
[0029] Enhanced learning correction module: Use the denoising process and image quality preservation strategy to recover a high-quality crowd density map from the noisy image;
[0030] Latent space dynamic optimization module: Based on the crowd density map, dynamically adjust the parameters of the crowd density recognition model in the latent space to obtain the crowd density recognition model in the optimal parameter state.
[0031] The beneficial effects of the present invention are as follows: 1. Through layer-by-layer diffusion simulation, the present invention simulates various image quality attenuation factors that may be encountered in a real environment, such as light changes and occlusions. This simulation enables the recognition system to better adapt to these changes in practical applications, thereby maintaining high accuracy. In the experiment, compared with the prior art, the present invention improves the accuracy of crowd density recognition by about 10% to 15% in an environment with unstable light and occlusions. 2. The enhanced learning correction and potential space dynamic optimization steps significantly reduce the demand for computing resources by optimizing the model processing flow and parameter adjustment. This makes the present invention applicable to environments with limited computing power, such as efficient operation can also be achieved on low-cost hardware. Compared with the prior art, when processing the same amount of data, the present invention reduces the computing resource consumption by about 20%. 3. The potential space dynamic optimization enables the system to dynamically adjust its behavior according to different monitoring scenarios, and this flexibility is not available in the prior art. The present invention can automatically adjust parameters according to environmental changes and is applicable to a variety of monitoring scenarios from indoor to outdoor and from day to night, greatly expanding the application scope. 4. In the test, the present invention evaluates the stability and accuracy of the model by conducting a series of crowd density recognition experiments under different light and occlusion conditions. The test results show that the recognition accuracy of the present invention remains above 85% under extreme light conditions, while the accuracy of the prior art is only about 70% under the same conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a flowchart of the recognition method in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0034] In this embodiment, a method for crowd density recognition based on dynamic optimization diffusion simulation is provided, aiming to improve the existing crowd monitoring system through a series of innovative technical steps and enhance its accuracy and reliability in complex environments. First, various factors that may cause image quality degradation in the real monitoring environment, such as light changes and occlusions, are simulated through layer-by-layer diffusion, and image noise is gradually introduced to train the model to adapt to these adverse conditions. Then, enhanced learning correction restores a high-quality crowd density map from the noisy image through a complex denoising process and image quality preservation strategy, ensuring data accuracy and image usability. Finally, latent space dynamic optimization optimizes the model's response to different monitoring scenarios by dynamically adjusting model parameters in the latent space, thereby significantly improving the model's adaptability and recognition accuracy. Overall, the method of the present invention significantly improves the performance of crowd density recognition technology in variable and complex real-world environments, is applicable to fields such as public safety and urban monitoring, reduces the dependence on high-performance computing resources, and enables the deployment of efficient monitoring systems in resource-constrained environments. As Figure 1 shown, the recognition method specifically includes the following parts,
[0035] I. Layer-by-layer diffusion simulation
[0036] The layer-by-layer diffusion simulation provides a method for simulating image quality degradation in the real environment. By gradually introducing noise, it effectively trains the model to adapt to various image degradation situations that may be encountered in actual monitoring scenarios, such as light changes and occlusions. This simulation provides an ideal training basis for subsequent enhanced learning correction, ensuring that the denoising function can operate effectively under the complex conditions of the real world.
[0037] The layer-by-layer diffusion simulation is specifically as follows: By simulating the impact of various interference factors (such as light changes, occlusions, dynamic background lights) that may be encountered in the real monitoring environment on the quality of image data. This process gradually increases the noise in the image, enabling the model to more robustly handle image quality degradation problems in actual applications. The relevant formula is,
[0038]
[0039] where, I t and I t-1 are the noisy images at the t-th and (t - 1)-th time steps respectively; N(0, 1) is a random variable following a standard normal distribution with a mean of 0 and a standard deviation of 1, used to simulate the random noise added in each step; k t is the diffusion coefficient at time step t, which is a key parameter used to control the intensity of noise increase, and its value can be dynamically adjusted according to environmental changes in the actual monitoring scenario; dt is the time increment, which defines the time interval for each step and determines the rate of noise addition.
[0040] The main idea of layer-by-layer diffusion simulation is derived from the diffusion processes in physics and chemistry, similar to heat diffusion or the random movement of particles in a fluid. In this method, this concept is applied to image processing to simulate the process of gradual degradation of image quality in the real world. In actual implementation, starting from the initial clear image, random noise is gradually introduced, and the noise at each step is calculated according to the diffusion coefficient k t which can be adjusted according to the actual environmental conditions (such as the visual reduction at night or in fog). The dynamic adjustment of k t is set based on environmental monitoring data (such as light sensors, weather information, etc.). For example, under poor lighting conditions, the value of k t will be set higher to simulate stronger image noise. This method enables the recognition system to take into account various complex factors during training, thus having better adaptability and robustness to these non-ideal conditions in practical applications.
[0041] In practical applications, layer-by-layer diffusion simulation can greatly improve the performance of the crowd density recognition system under various environmental conditions. For example, in a large music festival or sports event, the monitoring system often faces complex environmental conditions, such as rapid changes in lighting and fast movement of the crowd. Traditional crowd density recognition systems may fail due to sudden changes in image quality. By using layer-by-layer diffusion simulation, the system can handle these sudden changes more effectively because it has been trained enough under simulated harsh conditions. In addition, this technology is also applicable to the fields of traffic monitoring and urban security, especially under adverse weather conditions such as haze or heavy rain, which will significantly affect the image quality captured by the camera. By simulating these conditions during the model training process, the applicability and accuracy of the system can be significantly improved. In summary, layer-by-layer diffusion simulation not only provides an effective means to simulate the actual monitoring environment, but also provides strong support for the actual deployment of crowd density recognition technology, enabling it to remain efficient and accurate in a changing and complex real environment.
[0042] II. Enhanced Learning Correction
[0043] Enhanced learning correction improves the ability to recover high-quality images from the noisy images generated by layer-by-layer diffusion simulation through fine denoising and image quality maintenance. This not only improves the image quality, but more importantly, improves the accuracy of crowd density estimation. This step directly utilizes the data generated by layer-by-layer diffusion simulation, demonstrating how to reduce noise while maintaining image details.
[0044] Enhanced learning correction specifically involves: recovering a high-quality crowd density map from a noisy image through an advanced model training strategy. This process lies in a carefully designed loss function that not only emphasizes noise removal but also takes into account the maintenance of image quality to ensure that the recovered image accurately estimates the crowd density. The relevant formula is,
[0045]
[0046] where L is the total loss function, which is used as the optimization objective during the training process; D(I t ) is the denoising function that processes the input noisy image I t and outputs the denoised image. This function is learned and can identify and eliminate the noise components in the image; ‖I t - D(I t )‖ 2 is the mean square error between the crowd density map and the noisy image, which is used to evaluate the quality of the denoising effect; λ is the regularization coefficient, a hyperparameter used to adjust the balance between the denoising effect and the preservation of image details. A higher value of λ strengthens the preservation of image details, while a lower value emphasizes denoising; R(I t ) is the regularization function that provides additional constraints to maintain the important details and quality of the image. This is to prevent over-denoising that may erase the detailed information crucial for crowd density estimation in the image; T is the total number of time steps.
[0047] In practical implementation, the denoising function D(I t ) is usually implemented through a deep learning model, such as a convolutional neural network, which can learn complex image features and noise patterns. Through the training dataset, this network learns to identify and eliminate noise while trying to keep the original features and quality of the image intact. The regularization function R(I t ) is also learned and can be automatically adjusted based on the content complexity of the image. For example, it strengthens the protection of image details in crowded areas while being less strict in background areas.
[0048] In the practical application of crowd density recognition, especially in the fields of public safety and urban surveillance, the image quality is affected by various factors, such as weather conditions, light changes, etc. In these cases, the enhanced learning correction step ensures that the system can accurately estimate the crowd density even when the image quality is not ideal. For example, in a large outdoor event, the images captured by the surveillance cameras may become blurred due to the alternation of day and night or sudden weather changes. At this time, the enhanced learning correction dynamically adjusts the denoising and detail-preserving strategies of the model, enabling the system to work stably under various environmental conditions and provide accurate crowd statistics. In addition, the application of this technology is also very suitable for use in traffic surveillance systems, which can help accurately identify and count the number of vehicles and pedestrians in complex traffic flows and maintain a high recognition accuracy even in low-visibility environments such as rainy days or foggy days. To sum up, the enhanced learning correction not only improves the accuracy of the crowd density recognition system but also greatly enhances the applicability and robustness of the system in various environments, making it an indispensable part of modern surveillance systems.
[0049] III. Dynamic Optimization of the Latent Space
[0050] The dynamic optimization of the latent space optimizes the overall system response by adjusting the model's performance in the latent space. This step can dynamically adjust the model parameters according to the outputs of the previous two steps to optimize the detection accuracy of crowd density. It utilizes the data features processed by the previous two steps to further enhance the model's adaptability and accuracy to complex scenarios.
[0051] Specifically, the dynamic optimization of the latent space is achieved by redesigning the model's learning process in a dynamically optimized latent space to achieve a high degree of adaptability and recognition accuracy for complex crowd scenarios. This step uses highly customized functions to optimize the parameters in the latent space to adapt to various surveillance environments and changes in crowd density. The relevant formula is,
[0052]
[0053] where R is the total reward function of the latent space, which is used to evaluate and optimize the performance of the entire model; x j is the j-th feature in the latent space. These features are data abstracted and extracted from the crowd density map, which are used to capture the key information and complexity of the image; θ are the parameters of the crowd density recognition model, which control the behavior of the reward function f. These parameters are optimized through the training process; f is the reward function, which is customized according to the specific requirements of crowd density recognition and is used to adjust the behavior of the crowd density recognition model in the latent space to maximize the accuracy and efficiency of model recognition; N is the total number of features in the latent space.
[0054] Latent space dynamic optimization utilizes the latent feature learning theory in machine learning to extract the deep features of images through deep learning networks such as autoencoders or generative adversarial networks. These features are transformed into a multi-dimensional latent space, where each dimension captures certain key aspects of the input data. In this latent space, the model operates not directly on the original image data, but on these higher-level abstract features, enabling the model to learn and optimize more flexibly and effectively. The reward function f is designed to be adjustable and usually contains multiple levels and parameters to adapt to different environments and requirements. For example, for crowd density recognition in scenarios with large lighting variations, the reward function may particularly emphasize features that are insensitive to lighting changes, while in scenarios with highly crowded crowds, it may place more emphasis on the ability to distinguish different crowds.
[0055] In practical applications, latent space dynamic optimization provides great flexibility and adaptability for crowd density recognition systems. For example, in a large multi-functional commercial center, the density and mobility of the crowd vary greatly throughout the day and are also affected by different activities and time periods. Through latent space dynamic optimization, the recognition system can dynamically adjust its parameters according to the current specific situation to ensure accurate estimation of crowd density regardless of the environment. In addition, this technology is particularly effective in dealing with large-scale public events such as music festivals or sports events. In these events, the crowd density changes rapidly and is often accompanied by significant environmental noise and visual occlusion. Through the dynamic optimization of the latent space, the system can adjust in real time, quickly adapt to these changes, and improve the accuracy and response speed of recognition. In summary, latent space dynamic optimization not only enhances the adaptability of crowd density recognition technology but also significantly improves its application performance in complex environments, making it an indispensable part of modern intelligent monitoring systems.
[0056] IV. Crowd Density Estimation
[0057] Use the dynamic optimization diffusion model in the optimal parameter state for crowd density recognition to obtain the final crowd density estimation result.
[0058] In this embodiment, a crowd density recognition system based on a dynamic optimization diffusion model is also provided. The recognition system can implement the above-mentioned method. The recognition system includes,
[0059] (1) Layer-by-layer diffusion simulation module: By gradually increasing the noise of the image, simulate the process of gradual decay of image data to obtain a noisy image;
[0060] (2) Enhanced learning correction module: Use the denoising process and image quality preservation strategy to recover a high-quality crowd density map from the noisy image;
[0061] (3) Latent Space Dynamic Optimization Module: Based on the crowd density map, dynamically adjust the parameters of the crowd density recognition model in the latent space to obtain the crowd density recognition model in the optimal parameter state.
[0062] The present invention combines and applies these three core steps of layer-by-layer diffusion simulation, enhanced learning correction, and latent space dynamic optimization, producing an effect that is overall better than the sum of its parts. The layer-by-layer diffusion simulation provides a strong foundation, enabling the model to learn how to maintain stability in a changing environment; the enhanced learning correction further improves the image processing quality, ensuring the accuracy and usability of the data; finally, the latent space dynamic optimization ensures rapid and accurate response to various scenario changes in practical applications by finely tuning the model behavior.
[0063] The substantial feature of the combination of these three steps lies in its high adaptability and accuracy, which can significantly improve the performance of crowd density recognition in complex environments. In key fields such as public safety and urban surveillance, this invention provides significant progress. It can not only handle complex and changing environmental conditions but also provide high-precision data support in real time, greatly improving the practicality and reliability of the surveillance system. In addition, this method reduces the dependence on high-performance computing resources, enabling the deployment of an efficient surveillance system in resource-constrained environments. By using the above three core steps, the present invention not only solves the limitations of individual technologies but also significantly improves the performance and practical value of the entire method through their organic combination.
[0064] In summary, the present invention improves the accuracy and reliability of the crowd density recognition system under complex environmental conditions, especially for the case where the image quality deteriorates due to environmental factors such as light changes and occlusions. By introducing a method based on dynamic optimization diffusion simulation, the present invention effectively adapts to and processes these image quality changes, thus ensuring accurate estimation of crowd density even in dynamic or adverse visual environments. At the same time, this method also reduces the dependence on high-performance computing resources, enabling the crowd density recognition system to operate more efficiently in resource-constrained environments.
[0065] By adopting the above technical solutions disclosed in the present invention, the following beneficial effects are obtained:
[0066] The present invention provides a method and system for crowd density recognition based on a dynamic optimization diffusion model. Through layer-by-layer diffusion simulation, the present invention simulates various image quality attenuation factors that may be encountered in a real environment, such as light changes and occlusions. This simulation enables the recognition system to better adapt to these changes in practical applications, thereby maintaining high accuracy. In experiments, the present invention improves the accuracy of crowd density recognition by approximately 10% to 15% compared with the prior art in environments with unstable lighting and occlusions. The enhanced learning correction and latent space dynamic optimization steps significantly reduce the demand for computing resources by optimizing the model processing flow and parameter adjustment. This makes the present invention applicable to environments with limited computing power, such as enabling efficient operation on low-cost hardware. Compared with the prior art, when processing the same amount of data, the computing resource consumption of the present invention is reduced by approximately 20%. The latent space dynamic optimization enables the system to dynamically adjust its behavior according to different monitoring scenarios, and this flexibility is not available in the prior art. The present invention can automatically adjust parameters according to environmental changes and is applicable to a variety of monitoring scenarios from indoor to outdoor and from day to night, greatly expanding the application scope. In tests, the present invention evaluates the stability and accuracy of the model by conducting a series of crowd density recognition experiments under different lighting and occlusion conditions. The test results show that the recognition accuracy of the present invention remains above 85% under extreme lighting conditions, while the accuracy of the prior art is only about 70% under the same conditions.
[0067] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for crowd density recognition based on a dynamic optimization diffusion model, characterized in that: It includes the following steps: S1. Layer-by-layer diffusion simulation: By gradually increasing the noise of the image, simulate the process of gradual decay of image data to obtain a noisy image. S2. Enhanced learning correction: Utilize the denoising process and image quality preservation strategy to recover a high-quality population density map from the noisy image. S3. Latent space dynamic optimization: Based on the population density map, dynamically adjust the parameters of the population density recognition model in the latent space to obtain the population density recognition model in the optimal parameter state.
2. The method for identifying crowd density based on a dynamic optimization diffusion model according to claim 1, wherein: Specifically, in step S1, noise is gradually introduced into the image data to simulate the decline of image quality in the real environment, so as to train the dynamic optimization diffusion model to adapt to the influence of interference factors that may be encountered in the actual monitoring scenario on the quality of image data. The calculation formula is where, I t and I t-1 are the noisy images at the t-th and (t-1)-th time steps respectively; N(0,1) is a random variable following a standard normal distribution with a mean of 0 and a standard deviation of 1, which is used to simulate the random noise added in each step; k t is the diffusion coefficient at time step t; dt is the time increment.
3. The method for crowd density recognition based on a dynamic optimization diffusion model according to claim 1, wherein: In the layer-by-layer diffusion simulation, starting from the initial clear image, random noise is gradually introduced. The noise at each step is calculated according to the diffusion coefficient, and the diffusion coefficient is dynamically adjusted according to the environmental changes in the actual monitoring scenario.
4. The method for identifying crowd density based on a dynamic optimization diffusion model according to claim 1, wherein: Specifically, in step S2, determine the total loss function based on the denoising function and the regularization function, and use the total loss function to recover a high-quality denoised population density map from the noisy image. The calculation formula is Among them, L is the total loss function; D(I t ) is the denoising function; ‖I t - D(I t )‖ 2 is the mean square error between the crowd density map and the noisy image; λ is the regularization coefficient; R(I t ) is the regularization function; T is the total number of time steps.
5. The method for crowd density recognition based on a dynamically optimized diffusion model according to claim 1, characterized in that: The denoising function is obtained by training a deep learning model through the corresponding training data set. The denoising function can learn to identify and eliminate the noise components in the image while keeping the original features and quality of the image undamaged.
6. The method for crowd density recognition based on a dynamically optimized diffusion model according to claim 1, wherein: The regularization function is obtained by training a deep learning model through the corresponding training data set; the regularization function can be automatically adjusted based on the content complexity of the image, and it provides additional constraints to maintain the important details and quality of the noisy image.
7. The method for crowd density recognition based on a dynamic optimization diffusion model according to claim 1, wherein: Specifically, in step S3, use the total reward function in the latent space to optimize the parameters of the population density recognition model in the latent space to make it adapt to various monitoring environments and population density changes; the calculation formula is where R is the total reward function of the latent space; x j is the j-th feature in the latent space, that is, the data abstracted and extracted from the crowd density map; θ is the parameter of the crowd density recognition model; f is the reward function used to adjust the behavior of the crowd density recognition model in the latent space to maximize the accuracy and efficiency of model recognition; N is the total number of features in the latent space.
8. The method for crowd density recognition based on a dynamic optimization diffusion model according to claim 1, characterized in that: Latent space dynamic optimization utilizes the feature learning theory in machine learning to extract the deep features of the population density map through an autoencoder or a generative adversarial network. These features are transformed into a multi-dimensional latent space, where each dimension captures certain key aspects of the population density map; in the latent space, the model does not directly operate on the population density map, but on the higher-level abstract features, so that the model can learn and optimize more flexibly and effectively.
9. The method for crowd density recognition based on a dynamic optimization diffusion model according to claim 1, characterized in that: After step S3, it also includes S4. Population density estimation: Use the dynamic optimization diffusion model in the optimal parameter state to perform population density recognition to obtain the final population density estimation result.
10. A crowd density recognition system based on a dynamic optimization diffusion model, characterized in that: The recognition system can implement the method described in any one of claims 1 to 9 above. The recognition system includes A layer-by-layer diffusion simulation module: By gradually increasing the noise of the image, simulate the process of gradual decay of image data to obtain a noisy image. An enhanced learning correction module: Utilize the denoising process and image quality preservation strategy to recover a high-quality population density map from the noisy image. A latent space dynamic optimization module: Based on the population density map, dynamically adjust the parameters of the population density recognition model in the latent space to obtain the population density recognition model in the optimal parameter state.
Citation Information
Patent Citations
Crowd counting method based on generative adversarial network
CN108764085A
Crowd density detection method based on deep learning
CN111639668A
Dense crowd counting method for subway station scene
CN112818944A
Video image crowd counting method based on multi-scale attention mechanism
CN115631454A
Crowd density estimation method based on deep neural network
CN117351414A