A method for detecting crowd abnormal situation and single person motion state

By using multi-source heterogeneous data fusion technology and deep neural networks and multi-view, multi-modal analysis methods, the problem of detecting abnormal situations in crowds and individual movement states has been solved, achieving high-precision detection and early warning functions, which can be applied to crowd monitoring and emergency management.

CN116758469BActive Publication Date: 2026-06-02THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP
Filing Date
2023-05-24
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively integrate multi-source heterogeneous data, especially in highly complex and uncertain environments. They cannot accurately detect abnormal situations in crowds or individual movement patterns, and are particularly unsuitable for integrating multi-source data that includes both our own and the monitored target.

Method used

A multi-source heterogeneous data fusion method is adopted, including the processing of image data, overall environmental data and text information data. ResNet50 and DarkNet53 deep neural networks are used for feature extraction and target detection. PANet neural network and long short-term memory network are combined for data analysis. Multi-view and multi-modal fusion technology is used to detect abnormal situations in crowds and individual movement states.

Benefits of technology

It achieves high-precision detection of abnormal situations in crowds and individual movement status, providing technical support for crowd monitoring, personnel/material scheduling, and early warning of emergencies, and has good commercial application prospects and social benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116758469B_ABST
    Figure CN116758469B_ABST
Patent Text Reader

Abstract

The application discloses a kind of crowd abnormal situation and single person motion state detection method, the people and vehicle trajectory and flow data generated by unmanned aerial vehicle monitoring equipment are comprehensively integrated to single person carrying mobile perception device, by effectively analyzing and mining these data, to prevent possible danger, to prevent possible danger, so that single person outdoor execution task can be safely and orderly carried out.The application is applied to single person motion state evaluation and crowd integrated situation awareness system construction, provides technical support for crowd monitoring, personnel / materials scheduling, flow control, emergency warning and the like.Has good commercial application prospect and social benefits.The application mainly includes the following contents:1) crowd abnormal situation analysis and research based on multi-source heterogeneous data fusion.2) single person comprehensive capability evaluation and analysis based on multi-sensing data fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a data fusion information processing method, and in particular to a method for detecting abnormal situations in a crowd and the movement status of an individual. Background Technology

[0002] In the information age, personal information is of paramount importance. With the development of equipment, especially the improved performance of multi-source heterogeneous sensors, a large amount of multi-source heterogeneous personal information data is generated. The crowd and environmental target characteristic information mined from this data can be used to discover and identify individual targets within a crowd, and also to correlate and mine crowd data to understand target activity patterns, thereby predicting their trends and intentions to construct accurate crowd situation maps. High-level characteristics of crowd and environmental targets can improve the accuracy of target pattern mining and prediction. By fusing multi-source data to verify, match, and infer crowd target data, accurate prediction and judgment of crowd target activity trends and intentions can be achieved. Crowd target multi-source data fusion technology is an important information technology that countries around the world are vying to develop.

[0003] Multi-source data fusion has a wide range of applications, and progress has been made in various fields such as urban big data fusion, medical big data fusion, and flight trajectory big data fusion. The fusion of multi-source population data based on personal information also urgently needs research and development. Although the aforementioned multi-source data fusion methods have achieved good results in some scenarios and within certain scopes, these are all applications of data fusion in single domains. Currently, there is no specific multi-source data fusion method for collaborative work based on single-person movement, where the data source simultaneously includes both our own and the monitored target's data. Due to the high complexity and uncertainty of the latter's environment, multiple sensor nodes and related sensing and monitoring equipment such as drones move at high speeds in three dimensions, making rapid hierarchical dynamic resource scheduling difficult. Therefore, the aforementioned methods are not well-suited for application. Summary of the Invention

[0004] Purpose of the invention: The technical problem to be solved by the present invention is to provide a method for detecting abnormal situations in a crowd and the movement status of a single person, in order to address the shortcomings of the existing technology.

[0005] To address the aforementioned technical problems, this invention discloses a method for detecting abnormal situations in a crowd and the movement state of a single person, comprising the following steps:

[0006] Step 1: Collect multi-source heterogeneous data in a crowd environment. The multi-source heterogeneous data includes: image data, overall environmental data, and text information data.

[0007] Step 2, analysis and research on abnormal population situations based on multi-source heterogeneous data fusion, includes the following steps:

[0008] Step 2-1: For image data, establish a multi-target recognition model for crowd environments and perform multi-target recognition in crowd environments;

[0009] The multi-target recognition model for crowd environments consists of a crowd target quantity estimation module and a crowd target detection output module.

[0010] The target population estimation module includes the following steps:

[0011] First, the multi-view module is initialized based on the acquired image data. This module consists of a feature extraction layer, a view encoding layer, a view fusion layer, and a final classification layer. The invention uses the feature layer to extract features from the image data, the view encoding layer to encode features from different viewpoints, the view fusion layer to fuse the encoded features from different viewpoints, and finally, a fully connected layer to classify the fused features. The multi-view module initialization includes three steps: First, a crowd information classification model is established based on the multi-view image training dataset. This model includes a basic quintuple, which comprises: a set describing the crowd environment, a set of individual targets within the crowd, and a set of people. The first step involves setting the locations of individual targets within the crowd, the locations of people carrying equipment within the crowd, and the locations of equipment carried by individuals within the crowd. These are used to perform fine-grained classification of the image data collected on-site. The second step involves independently initializing the multi-view image data. The third step involves constructing an image standardization module to correct the crowd environment image data. After the multi-view module initialization task is completed, the data from the manually collected crowd information database is fused with the multi-view module to form multi-view environmental information. A ResNet50 deep neural network model is used to extract human features and perform fine-grained classification and deduplication on the multi-view environmental information, thus completing the construction of the crowd target quantity estimation module.

[0012] The crowd target detection output module includes the following steps:

[0013] First, crowd information is extracted from a manually collected crowd information database. Then, crowd image information preprocessing and crowd abnormality monitoring model loading are performed based on this database. Next, a Darknet53 deep convolutional neural network is used for target feature extraction, and an image feature pyramid method is employed to obtain the final target detection result. Finally, feature layer decoding regression classification is used. Based on the image feature pyramid, a deep neural network is used to extract image features at multiple scales. A single neural network is used to extract high-level features from the original image, and downsampling is performed layer by layer to obtain feature layers at different scales, thus completing the output of the crowd target detection.

[0014] Step 2-2: For individual environmental data, establish a target recognition system based on multi-source heterogeneous data, using an infrared fusion detection method. The specific method is as follows:

[0015] Using a manually collected infrared information database, infrared image data is enhanced and infrared human features are extracted. The PANet neural network model is then used for feature stitching. A confidence threshold filtering method is used to screen the confidence of the target recognition results and exclude detection results with low confidence.

[0016] Step 2-3: Based on the target recognition results described in Step 2-1 and Step 2-2, establish a multi-view crowd detail analysis model and perform multi-view crowd detail analysis.

[0017] Steps 2-4 involve conducting target situation analysis based on multimodal data for the text information data, specifically including:

[0018] For text information data, multimodal data analysis is adopted, which includes image information and text information; a text decoding model is established to extract key information from the text information database; target feature information extracted from the multi-view module and text information are fused; and the fused information is detected and output using confidence threshold filtering to complete text fusion detection, i.e. target situation analysis.

[0019] Step 3: Collect multi-sensor data for a single target in the crowd. The multi-sensor data includes: vital signs data, equipment data, and individual environmental data.

[0020] Step 4, Single-person motion state detection based on multi-sensor data fusion, includes the following steps:

[0021] Step 4-1: Based on vital sign data, establish a physical fitness assessment model. Specific methods include:

[0022] Basic physiological data of an individual is collected and preprocessed. The preprocessing includes using the mean replacement method to handle abnormal data for physiological features with missing values ​​and performing L2 norm feature standardization preprocessing. Based on the preprocessed collected data, a long short-term memory neural network model is trained as a physical fitness assessment model, and the individual's physical status result is output, i.e., the individual physical fitness assessment result.

[0023] The mean substitution method is as follows:

[0024] Step 1: Sort the physiological characteristic data in ascending order from smallest to largest;

[0025] Step 2: Calculate the position: The position of calibration parameter Q1 is (n+1)×0.25; the position of calibration parameter Q2 is (n+1)×0.5; the position of calibration parameter Q3 is (n+1)×0.75, where n is the total number of data;

[0026] Step 3: Calculate the quantile values:

[0027] If the result of the second step is an integer, then directly take the value corresponding to that position;

[0028] If the result of the second step is a decimal, then the value corresponding to the second position is multiplied by (1 - the decimal part) and the value corresponding to the third position is multiplied by the decimal part.

[0029] Step 4: Calculate the interquartile range (IQR):

[0030] IQR = Q3 - Q1

[0031] Step 5: Determine if the physiological characteristic data are outliers:

[0032] If the physiological characteristic data exceeds the range of [(Q1-1.5*IQR)~(Q3+1.5*IQR)], it is judged as an outlier;

[0033] Step 6: Calculate the mean of physiological characteristic data;

[0034] Step 7: Replace the outliers filtered out in Step 5 with the mean value;

[0035] The L2 norm feature standardization preprocessing described above is as follows:

[0036] Physiological feature data vector x(x1,x2,…,x) n The L2 norm of is defined as:

[0037]

[0038] Where, x n Let x represent the nth physiological characteristic data. To normalize x to the L2 norm, we need to establish a mapping from x to x' such that the L2 norm of x' is 1. This is achieved by dividing each dimension of vector x by ||x||² to obtain a new vector X².

[0039]

[0040] Step 4-2: Based on the equipment data, establish a single-person equipment status assessment model. The specific method is as follows:

[0041] Equipment data for a single individual is collected, and maximum / minimum normalization feature processing is used to scale different feature values ​​to the same range to obtain equipment feature data. Principal component analysis is used to extract features from the equipment feature data by mapping high-dimensional data to a low-dimensional space to remove noise and redundant features, thus completing dimensionality reduction and feature extraction. A convolutional neural network model is trained based on the processed data to serve as a single-person equipment status assessment model, which is used to assess the single-person equipment status and obtain the single-person equipment status assessment result.

[0042] Step 4-3: Based on environmental data, establish a population environmental status assessment model. Specific methods include:

[0043] Environmental data of the population's environment is collected and preprocessed using zero-mean standardization to ensure that the mean is 0 and the standard deviation is 1. The preprocessed data is then scaled to the same scale. Linear discriminant analysis is used to extract features from the scaled data, projecting the high-dimensional data into a low-dimensional space to maximize the distance between different categories and minimize the distance within the same category, thus achieving dimensionality reduction and classification. Based on the above-processed data, a recurrent neural network model is trained as a population environmental status assessment model, outputting the population environmental assessment results.

[0044] Step 4-4: Based on the three models mentioned above, establish an individual comprehensive ability assessment model and conduct an individual comprehensive ability assessment, specifically including:

[0045] The individual's comprehensive ability score is calculated by integrating the results of individual physical fitness assessment, individual equipment status assessment, and crowd environment assessment by combining the coefficient of variation method and the weighted rank sum ratio method.

[0046] The coefficient of variation method is used to determine the weights of the evaluation indicators, and the specific method is as follows:

[0047] The coefficients of variation for the three indicators—psychological stress, athletic ability, and equipment status—are calculated using the following method:

[0048] CV l =S l / X l (l=1,2,3)

[0049] In the formula, CV l S is the coefficient of variation of the l-th evaluation indicator. l X is the standard deviation of the l-th evaluation indicator. l It is the average of the l-th evaluation indicator;

[0050] Then calculate the weight W of each indicator. l The method is as follows:

[0051]

[0052] In the formula, W l is the weight of the l-th evaluation indicator, and m represents the number of evaluation indicators;

[0053] The weighted rank sum ratio method is as follows:

[0054] WRSR=∑RW / m

[0055]

[0056] Wherein, WRSR represents the weighted rank sum ratio, RSR represents the rank sum ratio, ∑R represents the rank sum value of the evaluation object index, W is the weight of each evaluation index, m is the number of evaluation indexes, k represents, and l represents; firstly, high-quality indicators are ranked from small to large, low-quality indicators are ranked from large to small, and indicators with the same value are ranked by average.

[0057] Beneficial effects:

[0058] The method proposed in this invention has been applied to the assessment of individual movement status and the construction of comprehensive crowd situational awareness systems, providing technical support for crowd monitoring, personnel / material scheduling, traffic control, and early warning of emergencies. It has excellent commercial application prospects and social benefits. Attached Figure Description

[0059] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.

[0060] Figure 1 This is a framework diagram for analyzing and researching abnormal situations in populations based on the fusion of multi-source heterogeneous data.

[0061] Figure 2 A framework diagram for assessing an individual's comprehensive capabilities through multi-sensory data fusion.

[0062] Figure 3 This is a system architecture diagram for fusion of multi-source heterogeneous data to address abnormal situations in the population.

[0063] Figure 4 This is a schematic diagram of a multi-target recognition model for crowd-oriented environments.

[0064] Figure 5 This is a system architecture diagram for a single person's comprehensive capabilities based on multi-sensory data fusion.

[0065] Figure 6 This is a structural diagram of the individual competency assessment subsystem.

[0066] Figure 7 This is a structural diagram of the equipment status assessment subsystem.

[0067] Figure 8 This is a structural diagram of the environmental state sub-assessment system.

[0068] Figure 9 This is a diagram of a comprehensive ability assessment model for an individual. Detailed Implementation

[0069] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0070] This invention focuses on addressing the diversity and complexity of perceptual data collected in crowd environments, including images, text, environmental data, and individual status data. Based on efficient data fusion technology using intelligent algorithms, it integrates the overall characteristics of multi-view data, mines the correlations of multimodal data, and conducts in-depth analysis of individual information and our own individual information based on intelligent data fusion technology. The scheme for analyzing and researching abnormal situations in crowds based on multi-source heterogeneous data fusion is as follows: Figure 1 As shown, the individual comprehensive ability assessment scheme based on multi-sensory data fusion is as follows: Figure 2 As shown.

[0071] The invention is mainly divided into two parts: analysis and research on abnormal situations in the population based on multi-source heterogeneous data fusion, and comprehensive ability assessment of an individual based on multi-sensory data fusion.

[0072] 1) Analysis and research on abnormal population situations based on multi-source heterogeneous data fusion

[0073] like Figure 3 As shown, the present invention consists of four parts: (1) construction of a multi-target recognition model for crowd environment; (2) construction of a target recognition model based on multi-source heterogeneous data; (3) crowd detail analysis based on multiple perspectives; and (4) target situation analysis method based on multi-modal data.

[0074] First, the first part is to establish a multi-target recognition model for crowd environments based on image data, and to perform multi-target recognition in crowd environments. This multi-target recognition model for crowd environments mainly consists of a crowd target quantity estimation module and a crowd target detection output module.

[0075] The implementation process of the target population size estimation module includes the following key steps:

[0076] First, the multi-view module is initialized based on the collected multi-angle crowd environment image data. This task includes three steps:

[0077] The first step is to establish a crowd information classification model based on a multi-view image training dataset, used for fine-grained classification of image data collected on-site; specifically, the model established in this invention includes a basic quintuple:

[0078] {E,T,TP,PP,P}

[0079] E represents a set of descriptions of the human environment:

[0080] E={Temperature,Humidity,Visibility,Others}

[0081] Temperature represents the ambient temperature of the crowd, Humidity represents the ambient humidity of the crowd, Visibility represents the ambient visibility of the crowd, and Others represents other environmental conditions in the crowd's environment.

[0082] T represents the set of individual goals within the population:

[0083] T(Target) = {T i |i=1,2...NT}

[0084] Where NT represents the number of individual targets in the population;

[0085] TP represents the set of individual target locations within a population:

[0086] TP(Target) = {TP} i |i=1,2...NT}

[0087] PP represents the set of location information for personnel carrying equipment within a group:

[0088] PP(Person) = {PP} i |i=1,2...NT}

[0089] Where NT represents the number of people carrying equipment in the crowd;

[0090] P represents the set of location information for the equipment carried by individuals within the population:

[0091] P(Equipment) = {PE} i |i=1,2...NT}

[0092] Where NT represents the number of devices carried by individuals in the population;

[0093] The second step is to use microservice technology to independently initialize the multi-view data so that complete multi-perspective crowd environment data can be built later.

[0094] The third step involves constructing an image standardization module to correct crowd environment image data. First, this invention collects crowd environment image data from image datasets of different sources and performs preprocessing. Then, it performs color space conversion, transforming the image data from the RGB color space to the LAB color space, thereby reducing color deviation and illumination variations in the image. Finally, it uses morphological filters to process the image and performs scale transformation to complete the image data correction task.

[0095] After the multi-view module initialization task is completed, this invention integrates data from a manually collected crowd information database with the multi-view module to form multi-view environmental information. [1] Then, the ResNet50 deep neural network model was used. [2] The system extracts human features and performs fine-grained classification and deduplication on multi-view environmental information to build a module for estimating the number of people in a crowd.

[0096] The implementation of the crowd target detection output module first requires extracting crowd information from a manually collected crowd information database. Based on this database, it then performs crowd image information preprocessing and loads a crowd abnormality monitoring model. The feature extraction task of this module uses a Darknet53 deep convolutional neural network, which is part of the YOLOv3 target detection algorithm. In the feature extraction task, targets may appear in the image at different scales and sizes; therefore, this invention employs an image feature pyramid technique to address this issue. Specifically, at the bottom of the pyramid, the initial input image is processed to its smallest size, which is also its highest resolution. Then, the feature map of this layer is extracted through a convolutional neural network. Simultaneously, the image size is reduced, and the feature map of this layer is extracted again through the convolutional neural network, repeating this process until the top of the pyramid is reached. Targets can be detected at each scale of the pyramid. By integrating and filtering these detection results, the final target detection result can be obtained. Ultimately, this invention employs feature layer decoding regression classification technology. Based on the image feature pyramid, it uses a deep neural network to extract image features at multiple scales, and uses a single neural network to extract high-level features of the original image. It then downsamples layer by layer to obtain feature layers at different scales, so that the model of this invention can simultaneously perceive target features at different scales, thereby achieving the output of crowd target detection.

[0097] In summary, such as Figure 4As shown, the multi-target recognition model for crowd environments designed in this invention mainly includes two stages: feature extraction and metric learning. The feature extraction stage is divided into three parts: single-view feature extraction, multi-view feature extraction, and global view feature extraction. Metric learning performs similarity metric learning on the AC image features obtained in the feature extraction stage to complete multi-class classification and multi-class quantity estimation of individual targets in the crowd. Single-view feature extraction inputs images of the crowd into the feature extraction network. After low-order general feature extraction through convolutional layers and pooling layers, ROI pooling is performed to extract high-order specific features from the crowd images, resulting in single-view crowd image features. Multi-view feature extraction involves initial feature extraction from multiple input crowd images through convolutional layers, followed by horizontal pooling to complete local and global feature extraction. When performing global feature extraction, the number of channels in the feature image needs to be appropriately reduced to obtain the global features of the crowd images. Then, global features and horizontal local features are fused to obtain multi-view features of the crowd input image. For overall view feature extraction, the global individual images of the crowd are input into a feature extraction network. After low-order general feature extraction through convolutional and pooling layers, ROI pooling is performed, and high-order specific feature extraction is conducted to obtain overall view features. After completing these three feature extractions, the data is fed into a metric learning network for multiple image similarity learning iterations. The metric-learned image features are then processed by a multi-class classifier and multi-class quantity estimation to complete the task of building a multi-target recognition model for crowd images.

[0098] Then, the second part is to establish target recognition based on multi-source heterogeneous data for the overall environmental data. This step is mainly achieved by infrared fusion detection technology.

[0099] This invention utilizes a manually collected infrared information database to enhance infrared image data, thereby achieving infrared human feature extraction. Furthermore, it employs the PANet neural network model for feature concatenation, addressing information loss and scale mismatch issues between different levels of features in the feature pyramid network to generate a global feature representation for better target identification and localization. In crowd target detection, this invention uses a focus-based crowd feature fusion mechanism. By calculating the weights of different parts of the input data, it achieves weighted information processing, allowing the model to focus more attention on data relevant to the target task, thus improving model performance. Finally, this invention employs a confidence threshold filtering technique to screen the confidence levels in the detection results, eliminating detection results with low confidence levels, thereby improving detection accuracy and model robustness.

[0100] The third part involves establishing a multi-view crowd detail analysis model based on the target recognition results achieved in the previous two parts, and conducting multi-view crowd detail analysis. First, this invention utilizes information obtained from the target recognition results, such as the number of people, location, and behavior, combined with multi-view data sources for comprehensive analysis. Multi-view data sources can include image data, sensor data, and social media information from portable devices, to obtain richer crowd details. Second, for the overall environmental data, this invention employs infrared image data and image enhancement technology to enhance and correct single-source data, thereby obtaining more accurate crowd detail analysis.

[0101] Finally, the fourth part aims to propose a target situation analysis method based on multimodal data. Specifically, for text information data, this invention employs multimodal data analysis, including image and text information. To fully utilize text information, this invention establishes a text decoding model to extract key information from the text information database. Based on this, this invention implements the loading of a multimodal fusion model, fusing target feature information and text information. Finally, confidence threshold filtering is used to detect and output the fused information, thereby achieving the purpose of text fusion detection.

[0102] This invention relates to a text decoding model that extracts key information from a text database. This model can transform text information into a computer-processable representation, such as a vector or matrix. The text decoding model consists of two parts: an encoder and a decoder. The encoder's task is to encode the input feature representation into a fixed-length vector; this invention uses a convolutional neural network structure for this process. When processing input features, the encoder typically considers information such as different time steps and hierarchical structures to better understand the semantics of the input. The decoder's task is to decode the fixed-length vector output by the encoder into a piece of natural language text. The decoder employs a Transformer structure, considering the words generated in the previous step, the features of the current time step, and previously generated text during the generation process to produce more fluent and accurate natural language text. The entire text decoding model is an end-to-end model; its input is a set of text input features, and its output is a piece of natural language text. The model training process typically uses methods such as maximum likelihood estimation to optimize model parameters by minimizing the difference between the model-generated text and the real text. This representation can be used by other models, such as multimodal fusion models.

[0103] This invention relates to a multimodal fusion model, which fuses information from different modalities to obtain more complete and accurate information. In this invention, multimodal information includes text and image information. The fusion model employs techniques such as deep neural networks to fuse this information. The multimodal fusion model structure consists of feature extractors for multiple modalities, a fusion layer, and a classifier. Each modal feature extractor uses a different deep neural network to extract features from the input data. The fusion layer combines features from different modalities using methods such as weighted summation, concatenation, and dot product. Finally, the classifier classifies the input based on the feature fusion result. The multimodal fusion model is divided into two methods: early fusion and late fusion. Early fusion refers to extracting features from the input data of multiple modalities before fusing them. Late fusion involves adding a fusion layer between the feature extractors and classifiers of each modality, fusing features from different modalities before the classifier. The function of the multimodal fusion model is to fuse target feature information and text information to achieve text fusion detection output. By fusing this information, the accuracy and efficiency of target situation analysis can be improved.

[0104] 2) Design and Implementation of a Single-Person Comprehensive Ability Assessment Method Based on Multi-Sensory Data Fusion

[0105] like Figure 4 As shown, the present invention consists of four parts: (1) physical fitness assessment model construction; (2) method implementation of single-person equipment status assessment model based on equipment data; (3) population environment status assessment model construction based on multi-source data; and (4) design and implementation of single-person comprehensive ability scoring algorithm.

[0106] The first part of this invention aims to establish a physical fitness assessment model based on vital sign data. Firstly, this invention collects basic physiological data such as individual height, weight, body fat percentage, heart rate, blood pressure, and respiratory rate. For physiological features with missing values, a mean replacement method is used for abnormal data processing. Specifically, for features containing missing values, this invention calculates the mean of the feature in the existing data and replaces the missing feature value. For the collected data, this invention performs L2 norm feature standardization preprocessing to normalize the data features according to different scaling ratios, thereby ensuring that the feature values ​​remain consistent in magnitude and avoiding the impact of differences between data features on model training. Since individual body sensor data exhibits temporal relationships, this invention uses a Long Short-Term Memory (LSTM) network, which also has temporal memory capabilities, to learn changes in an individual's physical state, analyze the individual's current vital signs, and assess individual physical fitness in five aspects: stress, endurance, speed, strength, and flexibility. Figure 5 The diagram shows the learning and training process of the physical fitness assessment model.

[0107] This invention employs a mean replacement method for outlier processing of physiological features with missing values. Specifically, for features containing missing values, this invention calculates the mean of the feature in the existing data and replaces the missing feature value. Because the sampling frequency of individual physiological data is high, it increases data instability. Furthermore, body sensors are easily affected by factors such as the distance between the sensor and the body and the external environment when collecting data such as heart rate and temperature. Therefore, most of the collected data will contain outliers. To compensate for prediction errors caused by outliers, this invention uses the quartile method to filter out potential outliers and then uses the mean replacement method to process them, thus better preserving the true nature of the data.

[0108] quartiles are very effective in exploring data distribution, as they compensate for the weakness of the mean, which is easily affected by outliers. The steps of the quartile mean substitution method are as follows:

[0109] Step 1: Sort the data in ascending order from smallest to largest.

[0110] Step 2: Calculate the positions (n ​​is the total number of data): Position of Q1 = (n+1) × 0.25; Position of Q2 = (n+1) × 0.5; Position of Q3 = (n+1) × 0.75.

[0111] Step 3: Calculate the quantile values:

[0112] 1) If the result of the second step is an integer, simply take the value corresponding to that position.

[0113] 2) If the result of the second step is a decimal—for example, 2.25: the value corresponding to the second position * (1-0.25) + the value corresponding to the third position * 0.25.

[0114] Step 4: Calculate the interquartile range (IQR):

[0115] IQR = Q3 - Q1

[0116] Step 5: Determine if the data is outlier

[0117] If the data exceeds the range of [(Q1-1.5*IQR)~(Q3+1.5*IQR)], it is considered an outlier.

[0118] Step 6: Calculate the mean of the data

[0119]

[0120] Step 7: Use Replace the outlier x selected in step 5

[0121] Next, for the data after anomaly handling, this invention performs L2 norm feature standardization preprocessing to normalize the data features according to different scaling ratios, thereby keeping the feature values ​​consistent in magnitude and avoiding the impact of differences between data features on model training.

[0122] The L2 norm feature normalization preprocessing method used in this invention is as follows:

[0123] The L2 norm of a vector x(x1,x2,…,xn) is defined as

[0124] To normalize x to the L2 norm, we need to establish a mapping from x to x' such that the L2 norm of x' is 1. Specifically, this involves dividing each dimension of vector X, x1, x2, ..., xn, by ||x||2 to obtain a new vector.

[0125]

[0126] After L2 norm normalization, the Euclidean distance and cosine similarity of individual human characteristics data are equivalent.

[0127] Based on the preprocessed collected data, this invention trains a long short-term memory neural network model. This model can output individual physical status results, presented as a five-dimensional radar chart of individual physical fitness and a textual assessment of individual physical fitness.

[0128] like Figure 7 As shown, the second part of this invention aims to provide a method for establishing a single-person equipment status assessment model based on equipment data; to this end, this invention first collects multi-source data such as the single person's communication status, peripheral device status, computer, tablet, mobile phone and drone equipment.

[0129] For this data, this invention employs extremum normalization feature processing, scaling different feature values ​​to the same range to avoid dimensional issues between different feature values ​​and ensuring a relatively balanced contribution of each feature to the model. The extremum normalization transformation formula used in this invention is as follows:

[0130]

[0131] Next, this invention uses principal component analysis to extract features from the feature data. This method maps high-dimensional data to a low-dimensional space, retaining the main information in the data and removing noise and redundant features, thereby achieving dimensionality reduction and feature extraction. This invention also performs aggregation operations on equipment status based on a weighted summation approach. Its basic steps are:

[0132] (1) Determine the specific stage at which each maturity capability attribute element is located, and determine the relative importance among them;

[0133] (2) Determine the maturity level of the capability attribute based on the weighted summation rule, and determine the relative importance among the capabilities at each maturity level;

[0134]

[0135] (3) The effectiveness maturity level of the equipment is evaluated based on the weighted summation rule.

[0136]

[0137] Among them, the maturity levels of physical domain action capability, information domain utilization capability, cognitive domain comprehension capability, and social domain management capability are Cap i (i = 1, 2, 3, 4), and their weight set is W = {w i (i=1,2,3,4)}, the maturity levels of each maturity capability attribute element are Cap ij (j = 1, 2, 3, ..., k) i The weight set between them is W. i ={w ij}, (k1, k2, k3, k4) = (7, 5, 4, 4).

[0138] Ultimately, as Figure 6 As shown, the present invention trains a convolutional neural network model based on the processed data, which is used to evaluate the equipment status of a single person;

[0139] Then, the third part is as follows Figure 8 As shown, this invention addresses environmental data and provides a population environmental status assessment model based on multi-source data. First, this invention collects multi-source environmental data such as weather, temperature, humidity, air pressure, wind speed, and altitude in the environment where the population is located.

[0140] To address the aforementioned data, this invention employs zero-mean standardization to preprocess the feature data, ensuring a mean of 0 and a standard deviation of 1. This scales the data features to a uniform size, avoiding inconsistencies in the units of measurement between features and improving the model's training speed and generalization ability. Before performing zero-mean standardization, this invention first obtains the following data:

[0141] 1) Mean (μ) of the population data

[0142] 2) The standard deviation (σ) of the population data, which must be on the same order of magnitude as the population in 1).

[0143] 3) Individual observation value (x)

[0144] By substituting the above three values ​​into the zero-normalized formula, we get:

[0145]

[0146] Thus, this invention can convert different data to the same magnitude, achieving standardization.

[0147] Next, this invention employs linear discriminant analysis to extract features from the data, projecting high-dimensional data into a low-dimensional space to maximize the distance between different categories and minimize the distance within the same category, thereby achieving data categorization and classification. Based on the preprocessed collected data, this invention trains a recurrent neural network model to output population environment assessment results.

[0148] Finally, as Figure 9 As shown, based on the above three models, this invention calculates an individual's comprehensive ability score by fusing individual physical fitness score, individual equipment status score, and crowd environment status score using the coefficient of variation method and the weighted rank sum ratio method.

[0149] The model algorithm mainly consists of two parts: the coefficient of variation method for calculating the weights of the evaluation results of individual physical fitness, group environment status, and individual equipment status, and the weighted rank sum ratio method for fusing and evaluating individual comprehensive ability. The specific implementation method of the algorithm is as follows:

[0150] 1) Determining the weights of evaluation indicators using the coefficient of variation method

[0151] First, calculate the coefficients of variation for the three indicators: psychological stress, athletic ability, and equipment status. The formula is:

[0152] CV j =S j / X j (j=1,2,3)

[0153] In the formula, CV j S is the coefficient of variation of the j-th evaluation indicator. j X is the standard deviation of the j-th evaluation indicator. j It is the average of the j-th evaluation indicator.

[0154] Then calculate the weight W of each indicator. j The formula is

[0155]

[0156] In the formula, W j It is the weight of the j-th evaluation indicator.

[0157] 2) Weighted rank sum ratio

[0158] The Rank-sum ratio (RSR) method is based on the idea of ​​obtaining a dimensionless statistical quantity RSR in an n-row (n evaluation objects) and m-column (m evaluation indicators or levels) matrix through rank transformation, and then ranking the evaluation objects according to their RSR values.

[0159] Formula expression:

[0160] WRSR=∑RW / n

[0161]

[0162] Where ∑R represents the rank sum of the evaluation indicators, W is the weight of each evaluation indicator, and n is the number of evaluation indicators. First, high-quality indicators are ranked from smallest to largest, low-quality indicators are ranked from largest to smallest, and indicators with the same value are ranked by average.

[0163] 3) Sorting by tier

[0164] Determine the rank and frequency f of each WRSR, calculate the cumulative frequency, and then calculate the corresponding probability unit. Based on the probability unit and the regression equation calculation results, rank the individual's comprehensive abilities using the results and the optimal grading principle.

[0165] This invention provides a concept and method for detecting abnormal situations in crowds and the movement status of a single person. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment of the invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.

Claims

1. A method for detecting abnormal situations in a crowd and the movement state of a single person, characterized in that, Includes the following steps: Step 1: Collect multi-source heterogeneous data in a crowd environment. The multi-source heterogeneous data includes: image data, overall environmental data, and text information data. Step 2: Analysis and research on abnormal population situations based on multi-source heterogeneous data fusion; Step 3: Collect multi-sensor data for a single target in the crowd. The multi-sensor data includes: vital signs data, equipment data, and individual environmental data. Step 4: Single-person motion state detection based on multi-sensor data fusion; Step 2, which involves the analysis and research of abnormal population conditions based on multi-source heterogeneous data fusion, includes the following steps: Step 2-1: For image data, establish a multi-target recognition model for crowd environments and perform multi-target recognition in crowd environments; Step 2-2: Establish target identification based on multi-source heterogeneous data for individual environmental data; Step 2-3: Based on the target recognition results described in Step 2-1 and Step 2-2, establish a multi-view crowd detail analysis model and perform multi-view crowd detail analysis. Steps 2-4 involve conducting target situation analysis based on multimodal data for the text information data, specifically including: For text information data, multimodal data analysis is adopted, which includes image information and text information; a text decoding model is established to extract key information from the text information database; target feature information extracted from the multi-view module and text information are fused; confidence threshold filtering is used to detect and output the fused information, thus completing text fusion detection, i.e. target situation analysis; Step 4, the single-person motion state detection based on multi-sensor data fusion, includes the following steps: Step 4-1: Establish a physical fitness assessment model based on vital sign data; Step 4-2: Based on the equipment data, establish a single-person equipment status assessment model; Step 4-3: Based on environmental data, establish a population environmental status assessment model; Step 4-4: Based on the above three models, establish an individual comprehensive ability assessment model and conduct an individual comprehensive ability assessment.

2. The method for detecting abnormal situations in a crowd and the movement state of a single person according to claim 1, characterized in that, The multi-target recognition model for crowd environments described in step 2-1 consists of a crowd target quantity estimation module and a crowd target detection output module; The target population estimation module includes the following steps: First, the multi-view module is initialized based on the acquired image data. The multi-view module consists of a feature extraction layer, a view encoding layer, a view fusion layer, and a final classification layer. The feature layer extracts features from the image data, the view encoding layer encodes features from different viewpoints, the view fusion layer fuses the encoded features from different viewpoints, and finally, a fully connected layer classifies the fused features. The multi-view module initialization includes three steps: First, a crowd information classification model is established based on the multi-view image training dataset. This model includes a basic quintuple, which comprises: a set of crowd environment descriptions, a set of individual targets within the crowd, and a set of... The first step involves a set of individual target locations, a set of location information for people carrying equipment within the crowd, and a set of location information for equipment carried by individuals within the crowd. These are used to perform fine-grained classification of the image data collected on-site. The second step involves independently initializing the multi-view image data. The third step involves constructing an image standardization module to correct the crowd environment image data. After the multi-view module initialization task is completed, the data from the manually collected crowd information database is fused with the multi-view module to form multi-view environment information. A ResNet50 deep neural network model is used to extract human features and perform fine-grained classification and deduplication on the multi-view environment information, thus completing the construction of the crowd target quantity estimation module. The crowd target detection output module includes the following steps: First, crowd information is extracted from a manually collected crowd information database. Then, crowd image information preprocessing and crowd abnormality monitoring model loading are performed based on this database. Next, a Darknet53 deep convolutional neural network is used for target feature extraction, and an image feature pyramid method is employed to obtain the final target detection result. Finally, feature layer decoding regression classification is used. Based on the image feature pyramid, a deep neural network is used to extract image features at multiple scales. A single neural network is used to extract high-level features from the original image, and downsampling is performed layer by layer to obtain feature layers at different scales, thus completing the output of the crowd target detection.

3. The method for detecting abnormal situations in a crowd and the movement state of a single person according to claim 2, characterized in that, Step 2-2 describes establishing target recognition based on multi-source heterogeneous data, using an infrared fusion detection method. The specific method is as follows: Using a manually collected infrared information database, infrared image data is enhanced and infrared human features are extracted, and the PANet neural network model is used for feature stitching. A confidence threshold filtering method is used to screen the confidence levels in the target recognition results and exclude detection results with low confidence levels.

4. The method for detecting abnormal situations in a crowd and the movement state of a single person according to claim 3, characterized in that, The specific methods for establishing the physical fitness assessment model described in step 4-1 include: Basic physiological data of an individual is collected and preprocessed. The preprocessing includes using the mean replacement method to handle abnormal data for physiological features with missing values ​​and performing L2 norm feature standardization preprocessing. Based on the preprocessed collected data, a long short-term memory neural network model is trained as a physical fitness assessment model, and the individual's physical status result is output, i.e., the individual physical fitness assessment result. The mean substitution method is as follows: Step 1: Sort the physiological characteristic data in ascending order from smallest to largest; Step 2: Calculate the position: The position of calibration parameter Q1 is (n+1) × 0.25; the position of calibration parameter Q2 is (n+1) × 0.5; the position of calibration parameter Q3 is (n+1) × 0.75, where n is the total number of data points; Step 3: Calculate the quantile values: If the result of the second step is an integer, then directly take the value corresponding to that position; If the result of the second step is a decimal, then the value corresponding to the second position is multiplied by (1 - the decimal part) and the value corresponding to the third position is multiplied by the decimal part. Step 4: Calculate the interquartile range (IQR): IQR = Q3 - Q1 Step 5: Determine if the physiological characteristic data are outliers: If the physiological characteristic data exceeds the range of [(Q1 - 1.5 * IQR) ~ (Q3 + 1.5 * IQR)], it is judged as an outlier; Step 6: Calculate the mean of physiological characteristic data; Step 7: Replace the outliers filtered out in Step 5 with the mean value; The L2 norm feature standardization preprocessing described above is as follows: Physiological feature data vector The L2 norm is defined as: ; in, Let x represent the nth physiological characteristic data; to normalize x to the L2 norm, we need to establish a mapping from x to x' such that the L2 norm of x' is 1. Divide each dimension of the data by A new vector is obtained ,Right now 。 5. The method for detecting abnormal situations in a crowd and the movement state of a single person according to claim 4, characterized in that, The specific method for establishing the single-person equipment status assessment model described in step 4-2 is as follows: Equipment data for a single individual is collected, and maximum / minimum normalization feature processing is used to scale different feature values ​​to the same range to obtain equipment feature data. Principal component analysis is used to extract features from the equipment feature data by mapping high-dimensional data to a low-dimensional space to remove noise and redundant features, thus completing dimensionality reduction and feature extraction. A convolutional neural network model is trained based on the processed data to serve as a single-person equipment status assessment model, which is used to assess the single-person equipment status and obtain the single-person equipment status assessment result.

6. The method for detecting abnormal situations in a crowd and the movement state of a single person according to claim 5, characterized in that, The specific methods for establishing the population environmental status assessment model described in step 4-3 include: Environmental data of the population's environment is collected and preprocessed using zero-mean standardization to ensure that the mean is 0 and the standard deviation is 1. The preprocessed data is then scaled to the same scale. Linear discriminant analysis is used to extract features from the scaled data, projecting the high-dimensional data into a low-dimensional space to maximize the distance between different categories and minimize the distance within the same category, thus achieving dimensionality reduction and classification. Based on the above-processed data, a recurrent neural network model is trained as a population environmental status assessment model, outputting the population environmental assessment results.

7. The method for detecting abnormal situations in a crowd and the movement state of a single person according to claim 6, characterized in that, Step 4-4, which describes conducting a comprehensive individual ability assessment, specifically includes: The individual's comprehensive ability score is calculated by integrating the results of individual physical fitness assessment, individual equipment status assessment, and crowd environment assessment by combining the coefficient of variation method and the weighted rank sum ratio method. The coefficient of variation method is used to determine the weights of the evaluation indicators, and the specific method is as follows: The coefficients of variation for the three indicators—psychological stress, athletic ability, and equipment status—are calculated using the following method: ; In the formula, It is the coefficient of variation of the l-th evaluation indicator. It is the standard deviation of the l-th evaluation indicator. It is the average of the l-th evaluation indicator; Then calculate the weight of each indicator. The method is as follows: ; In the formula, is the weight of the l-th evaluation indicator, and m represents the number of evaluation indicators; The weighted rank sum ratio method is as follows: ; Where WRSR represents the weighted rank-sum ratio, and RSR represents the rank-sum ratio. Let W represent the rank sum of the evaluation indicators, W be the weight of each evaluation indicator, m be the number of evaluation indicators, k be the kth evaluation indicator, and l be the lth evaluation indicator. First, high-quality indicators are ranked from smallest to largest, low-quality indicators are ranked from largest to smallest, and indicators with the same value are ranked by average.