Human body behavior recognition method based on machine learning

By constructing an LSTM network model optimized based on crayfish optimization algorithm and using inverse distance weighted feature extraction method, the problem of low accuracy of human behavior recognition is solved, and higher recognition accuracy and model training efficiency are achieved.

CN120164260AActive Publication Date: 2025-06-17LUSHAN COLLEGE OF GUANGXI UNIV OF SCI & TECH

Patent Information

Application Number
CN202510329360.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-17
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

In the prior art, the accuracy of human behavior recognition is low and is affected by a variety of factors.

Method used

Using a machine learning-based human behavior recognition method, a LSTM network model optimized based on crayfish optimization algorithm is constructed by obtaining target video data and historical data for preprocessing, and a global feature representation is extracted using inverse distance weighting, and finally input it into the optimal behavior recognition model for identification.

Benefits of technology

It improves the accuracy and reliability of human behavior recognition, enhances the training efficiency and robustness of the LSTM network model, and solves the problem of inaccurate behavior recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164260A_ABST
    Figure CN120164260A_ABST
Patent Text Reader

Abstract

The invention provides a human body behavior recognition method based on machine learning, and relates to the technical field of human body behavior recognition, and the method comprises the steps: obtaining target video data, and collecting video data containing different human body behaviors and corresponding behavior types as historical data; preprocessing the target video data and the historical data to obtain target data and training data; constructing an LSTM network model optimized based on a crayfish optimization algorithm, and inputting training data for training to obtain an optimal behavior recognition model; and inputting the target data into the optimal recognition model to obtain a behavior category. According to the method, the LSTM network model optimized based on the crayfish optimization algorithm is constructed, the training data is input for training, the optimal recognition model is obtained, feature extraction is performed on the target video data, global feature representation is obtained through an inverse weighting mode, behavior information in the video is comprehensively reflected, and the recognition efficiency is improved. And the accuracy and reliability of the optimal behavior recognition model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human behavior recognition, and in particular to a human behavior recognition method based on machine learning. Background Art

[0002] Human behavior recognition refers to the process of automatically detecting, analyzing, and understanding human actions, postures, and behavior patterns in video images or sensor data by using technologies such as computer vision, image processing, pattern recognition, and artificial intelligence. Its basic principle is to analyze and extract human action, posture, and behavior features, and use mathematical models and machine learning algorithms for pattern matching and classification recognition. Through human behavior recognition technology, the detection and alarm of suspicious behaviors can be realized. By analyzing human behaviors through surveillance cameras, unusual behavior patterns such as theft and harassment can be recognized and reported to the police in a timely manner, thereby improving public safety. Human behavior recognition can achieve a touchless human-computer interaction method, providing a more natural and intelligent interaction experience. Through gesture recognition, the operation of devices such as TVs and smart homes can be controlled by gestures. Through human behavior recognition technology, the health monitoring of special groups such as the elderly and children can be carried out. By analyzing the posture and activity trajectory of the human body, it can be judged whether the elderly have fallen, and early warning and rescue can be carried out in a timely manner. Human behavior recognition can realize the management of personnel identity and behavior. With the continuous development of artificial intelligence and big data technologies, human behavior recognition methods will have a broader development space.

[0003] However, in actual applications, affected by many factors, the accuracy of human behavior recognition is relatively low. Summary of the Invention

[0004] The present invention provides a human behavior recognition method based on machine learning to solve the defect that behavior recognition in the prior art is not accurate enough.

[0005] On the one hand, the present invention provides a human behavior recognition method based on machine learning, including: S1: Obtain target video data, and collect video data containing different human behaviors and corresponding behavior categories as historical data.

[0006] S2: Preprocess the target video data and the historical data respectively to obtain target data and training data.

[0007] S3: Construct an LSTM network model optimized by a crayfish optimization algorithm, and input the training data for training to obtain an optimal behavior recognition model.

[0008] S4: Input the target data into the optimal recognition model to obtain the behavior category.

[0009] A human behavior recognition method based on machine learning provided by the present invention. In step S2, the specific steps of preprocessing include: S21: Extract video frames from the video data at a preset time interval, process the video frames to obtain enhanced video frames, and use a convolutional neural network to extract features from the enhanced video frames to obtain a feature sequence.

[0010] S22: Perform normalization processing and smoothing processing on the feature sequence to obtain an enhanced feature sequence.

[0011] S23: Calculate the distances between the enhanced feature sequences, and use the inverse distance weighting method based on the distances to obtain weighted features.

[0012] S24: Summarize the weighted features through the method of average value aggregation to obtain a global feature representation.

[0013] A human behavior recognition method based on machine learning provided by the present invention. In step S21, the specific steps of processing the video frames include: S211: Use the method of hue transformation to enhance the video frames.

[0014] S212: Use Gaussian blur to filter the noise of the video frames.

[0015] A human behavior recognition method based on machine learning provided by the present invention. In step S22, the specific steps of the smoothing processing include: S221: Remove the outliers in the feature sequence, and further supplement the missing values in the feature sequence by the interpolation method.

[0016] S222: Use the moving average method to smooth the feature sequence.

[0017] A human behavior recognition method based on machine learning provided by the present invention. In step S23, the expression formula of the inverse distance weighting is:

[0018] Where, is the value of the unknown point x to be estimated, is the data value of the th known point, is the distance between the unknown point x and the th known point, is the exponent of the weight.

[0019] A human behavior recognition method based on machine learning provided by the present invention. In step S3, the specific steps of constructing the LSTM network model include: S31: Setting a network architecture including an input layer, an LSTM layer, and an output layer. The input layer is used to receive the target data. The LSTM layer is used to capture the time dependency of the target data and identify the dynamic changes of the behavior. The output layer is used to convert the output of the LSTM layer into a specific behavior category.

[0020] S32: two Dropout layers are set in the network architecture. The two Dropout layers are respectively set between the input layer and the LSTM layer, and between the LSTM layer and the output layer. The two Dropout layers are used to prevent the model from overfitting.

[0021] According to a human behavior recognition method based on machine learning provided by the present invention, in step S3, the specific steps of optimizing the crayfish optimization algorithm include: S33: Use mean square error as the fitness function of the LSTM network model.

[0022] S34: Initialize the population and randomly generate individual crayfish positions. Each crayfish individual represents a LSTM network model hyperparameter group.

[0023] S35: Calculate the individual fitness values ​​of the initialized population and randomly generate the ambient temperature value.

[0024] S36: An iteration mechanism is set. Each time a new population is formed, an ambient temperature is randomly selected, and population one or population two is obtained according to the iteration mechanism.

[0025] S37: Calculate the fitness value of each individual in the population one or the population two. If the output fitness value is higher than the preset fitness threshold, output the optimal hyperparameter. Otherwise, continue to iterate until the fitness value is higher than the preset fitness threshold or the maximum number of iterations is reached to obtain the optimal hyperparameter.

[0026] According to a human behavior recognition method based on machine learning provided by the present invention, in step S3, in step S36, the iteration mechanism specifically includes: S361: When the ambient temperature is higher than the preset ambient temperature threshold, the individuals of the population find the location of the summer cave based on the current global optimal position and the individual historical optimal position.

[0027] S362: Individuals of the population move to the location of the summer cave. If the locations of the individuals are the same, by comparing the individual fitness values, the individuals with low fitness will randomly select the location of another crayfish to adjust their positions to compete for the summer cave. When the positions of all individuals are updated, population one is obtained.

[0028] S363: When the ambient temperature is lower than or equal to the preset ambient temperature threshold, the optimal position of the individual in the current population is used as the food position.

[0029] S364: Determine the size of the food according to the individual fitness values ​​of the current population, and select a foraging method to forage according to the determination result, thereby obtaining population two.

[0030] According to a human behavior recognition method based on machine learning provided by the present invention, in step S361, the summer cave location expression formula is:

[0031] in, is the location of the cave, is the individual's best historical position, is the previous global optimal position.

[0032] According to a human behavior recognition method based on machine learning provided by the present invention, in step S364, the expression formula for determining the size of food is:

[0033] in, It is the identification value of food size. is the food factor, which is a constant, is a random number, between, The current individual's fitness value, The fitness value of food, when When , the food is judged to be small.

[0034] The human behavior recognition method based on machine learning provided by the present invention obtains an optimal behavior recognition model by constructing an LSTM network model optimized based on a crayfish optimization algorithm and inputting training data for training, performs feature extraction on target video data, obtains global feature representation by an inverse weighted manner, obtains target data, and inputs the target data into the optimal behavior recognition model to obtain behavior categories, thereby solving the problem of inaccurate behavior recognition. The global feature representation can more comprehensively reflect the behavior information in the video, and improve the accuracy and reliability of the optimal behavior recognition model. The crayfish optimization algorithm improves the time required for LSTM network model training, improves overall operational efficiency, and robustness, thereby further improving recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0036] Figure 1 It is a schematic flowchart of the human behavior recognition method based on machine learning provided by an embodiment of the present invention. Detailed implementation manners

[0037] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0038] Please refer to Figure 1 , the human behavior recognition method based on machine learning provided by an embodiment of the present invention, the method includes: S1: Obtain target video data, and collect video data containing different human behaviors and corresponding behavior categories as historical data.

[0039] In this embodiment, when collecting video data of different human behaviors, the content covered is quite extensive, mainly including some common limb movements. These common movements may include, but are not limited to, simple movements such as running, jumping, and waving. In addition, more complex limb movements, such as dancing, playing basketball, and swimming, are also included. Interaction actions between people and objects, such as making a phone call, eating, and opening a door, can also be considered. In the actual application process, relevant limb movement video data can also be customized according to the needs of a specific field to ensure the professionalism and applicability of the data. For example, if it is in the medical field, relevant actions of rehabilitation exercises can be focused on. If it is in the sports field, video collection can be carried out around the relevant behaviors of a specific sport. When collecting data on different human behaviors, it is crucial to ensure the diversity of the video data. Diversity is not only reflected in the limb movements themselves, but also should include different scene settings, lighting conditions, shooting angles, and backgrounds. The video can be shot in outdoor and indoor environments. The lighting can include natural light and artificial light sources. Different backgrounds can increase the diversity of the data set. The collected video data also needs to be labeled to ensure that each video has one or more behavior category labels for subsequent model training and evaluation. These labeling tasks can be completed through manual labeling or by using some existing efficient labeling tools to improve work efficiency.

[0040] S2: Preprocess the target video data and the historical data respectively to obtain target data and training data.

[0041] S21: Extract video frames from the video data at a preset time interval, process the video frames to obtain enhanced video frames, and use a convolutional neural network to extract features from the enhanced video frames to obtain a feature sequence.

[0042] In this embodiment, it is possible to set to extract 30 frames per second fixedly, and obtain a series of frame sequences from the entire video. For example, starting from the first second of the video, one frame can be extracted every 33 milliseconds, and a set of consecutive frames can be formed in this way. The number of frames extracted per second can be adjusted according to the complexity of the video content and the analysis requirements. For example, for a video with fast movement, a higher FPS may be required to capture more details. For a static or slowly changing video, a lower FPS may be sufficient. Arrange the extracted frames in chronological order to form a set of consecutive frame sequences. This set of frame sequences is processed for subsequent feature extraction steps. The convolutional neural network model will extract features from each frame image to generate corresponding feature vectors. These feature vectors contain key information in the image, such as: edges, textures, object shapes, etc. Arrange the feature vectors of each frame in chronological order to obtain a feature sequence.

[0043] S211: Use the method of hue transformation to enhance the video frames.

[0044] In this embodiment, through hue transformation, the contrast of the region of interest in the video frames can be enhanced, making it easier for the model to recognize important features and improving the accuracy of subsequent model recognition.

[0045] S212: Use Gaussian blur to filter the noise of the video frames.

[0046] In this embodiment, it is easier to capture important image information during feature extraction and enhance the sensitivity of the subsequent model to features.

[0047] S22: Perform normalization processing and smoothing processing on the feature sequence to obtain an enhanced feature sequence.

[0048] In this embodiment, the expression formula for normalization is:

[0049] where, is the normalized feature value, is the original feature value, is the minimum value of the feature sequence, is the maximum value of the feature sequence.

[0050] Through smoothing, unnecessary differences between adjacent samples are reduced, thereby improving the separability between different types of features and contributing to the improvement of the performance of classification algorithms.

[0051] S221: Remove the outliers from the feature sequence and further supplement the missing values in the feature sequence by interpolation.

[0052] In this embodiment, the outliers in the feature sequence originate from errors in the data collection process, malfunctions of measurement devices, or the inherent variability of the data itself. Outliers appear as extreme high or low values. To effectively identify and process outliers, by setting a threshold, values outside this range can be regarded as outliers and processed. However, the effectiveness of this method depends on the reasonable selection of the threshold, which usually needs to be determined based on the specific distribution and characteristics of the data. Another more statistical method is to use the median absolute deviation (MAD) to detect outliers. MAD is a robust statistic that is not affected by extreme values and can therefore more accurately reflect the central tendency and dispersion of the data. By calculating the absolute deviation between each weighted feature value and its median and comparing it with MAD, outliers that significantly deviate from the normal range can be identified. In addition, there are many other outlier detection techniques available, such as the box plot method, the Z-score method, density-based clustering methods, etc. Each method has its unique advantages and applicable scenarios, so in practical applications, the most suitable method can be selected according to the characteristics of the data and the requirements of the task. Once the outliers are identified, various strategies can be adopted to process them.

[0053] S222: Smooth the feature sequence using the moving average method.

[0054] S23: Calculate the distance between the enhanced feature sequences and obtain weighted features in the way of inverse distance weighting based on the distance.

[0055] In this embodiment, the Euclidean distance can be used to calculate the distance between the enhanced feature sequences. The expression formula of the Euclidean distance is:

[0056] where represents the Euclidean distance between feature vector A and feature vector B, is the value of vector A in the th dimension, is the value of vector B in the th dimension.

[0057] Another method of weight allocation focuses on the consideration of similarity between feature vectors. Feature vectors with higher similarity are assigned greater weights. Similar feature vectors often carry more consistent or relevant information, so they should occupy a more important position in analysis or model construction. To achieve this goal, various similarity metrics can be adopted, and cosine similarity is a commonly used option. By converting the calculated cosine similarity values into corresponding weight values, the relative importance between feature vectors can be quantified. In addition, Manhattan distance is also an effective means to measure the distance between feature vectors, especially in scenarios where high-dimensional data is processed and the features are sparse. Manhattan distance evaluates the distance between two vectors by calculating the sum of the absolute differences in each dimension. This calculation method makes it more sensitive to sparse features and can capture subtle differences that may be overlooked under traditional Euclidean distance. In certain specific cases, Manhattan distance may be more accurate than Euclidean distance in reflecting the actual relationship between feature vectors, thus providing a more reasonable basis for weight allocation.

[0058] In step S23, the expression formula for inverse distance weighting is:

[0059] where is the value of the unknown point x to be estimated, is the data value of the th known point, is the distance between the unknown point x and the th known point, is the exponent of the weight.

[0060] In this embodiment, another expression formula for inverse distance weighting can also be used:

[0061] where is the weight of the th feature vector, is the distance between the th feature vector and its adjacent feature vectors, is a very small constant used to prevent division by zero when calculating the weight.

[0062] When the feature distance between two frames is small, it means that they have high similarity in terms of visual content, motion patterns, or semantic information, etc. Such frames often belong to continuous parts of the same scene, action, or event, and are crucial for understanding the overall structure and content of the video. Therefore, by increasing the weights of these frame features, the feature representation of the video can be more focused on the continuous and consistent information, improving the accuracy and efficiency of video analysis or recognition.

[0063] Conversely, when the feature distance between two frames is large, they may represent significant change points in the video, such as scene transitions, object appearances or disappearances, etc. However, when dealing with the continuous changes in video content, these significant change points are usually not the main focus. Therefore, reducing the weights of these frame features can reduce noise and interference, making the feature representation smoother and more coherent. This weight assignment strategy not only helps improve the accuracy of video content analysis but also optimizes the use of computing resources. By focusing on temporally adjacent and content-similar frames, unnecessary data processing and computational overhead can be reduced, improving the real-time performance and efficiency of video processing.

[0064] In addition, this strategy can also be combined with other video processing techniques, such as motion estimation, object tracking, and event detection, etc., to build a more comprehensive and powerful video analysis system. By comprehensively considering the temporal continuity, content similarity, and feature weights between frames, we can achieve more accurate and efficient video content understanding and analysis.

[0065] S24: Summarize the weighted features through the method of average aggregation to obtain a global feature representation.

[0066] In this embodiment, the global feature representation is obtained by adding the continuous values of all weighted features and taking the average. In this process, each weighted feature is considered equally, and its value is summarized to form a comprehensive feature representation.

[0067] S3: Construct an LSTM network model optimized by the crayfish optimization algorithm, and input the training data for training to obtain an optimal behavior recognition model.

[0068] In this embodiment, the crayfish optimization algorithm is an efficient search and optimization algorithm that can search in a wide hyperparameter space and find the best parameter combination, thereby improving the performance of the LSTM network model. Using the optimization algorithm can enhance the model's learning ability, training speed, and final recognition effect, enabling the model to have stronger memory ability and accuracy when processing time series data.

[0069] In step S3, the specific steps for constructing the LSTM network model include: S31: Set a network architecture including an input layer, an LSTM layer, and an output layer. The input layer is used to receive the target data. The LSTM layer is used to capture the temporal dependencies of the target data and identify the dynamic changes in behavior. The output layer is used to convert the output of the LSTM layer into specific behavior categories.

[0070] In this embodiment, the input layer converts the input data into a format that the model can process efficiently, namely a three-dimensional tensor. This three-dimensional tensor contains three key dimensions: the number of samples, the number of time steps, and the number of features. The number of samples represents the number of independent records in the dataset, and each sample represents an independent observation or event. The number of time steps reflects the length of the sequential data, that is, the number of time points included in each sample, which is crucial for capturing the temporal dynamics in the data. The number of features refers to the number of variables observed at each time point, and these variables together constitute a comprehensive description of the data. Through the input layer, the data is passed to the LSTM layer, which is a core component of the model architecture. The LSTM layer, that is, the long short-term memory network layer, is designed specifically to capture long-term dependencies in sequential data. By utilizing special memory cells and gating mechanisms, it can effectively avoid the problems of vanishing gradients and exploding gradients when processing long sequence data, thereby more accurately capturing the temporal dynamics and long-term trends in the data. Finally, the output layer is responsible for converting the output of the LSTM layer into the final recognition result. The output layer can adopt various strategies and methods, such as classification, regression, or sequence-to-sequence conversion, etc., depending on the application scenario and objective of the model. The result of the output layer will directly reflect the model's understanding and analysis of the input data.

[0071] S32: Set two Dropout layers in the network architecture. The two Dropout layers are respectively set between the input layer and the LSTM layer, and between the LSTM layer and the output layer. The two Dropout layers are used to prevent the model from overfitting.

[0072] In this embodiment, the dropout rate of the Dropout layer can be set to 0.5, which means that in each training iteration, each neuron has a 50% probability of being randomly discarded. By randomly discarding neurons, the Dropout layer can force the model to learn more robust feature representations, thereby reducing the over-dependence on the training data. The model needs to adapt to the situation where different neurons are discarded, so it can show stronger generalization ability on unseen data. Although the Dropout layer adds randomness during the training process, it actually increases the convergence speed of the model and optimizes the final performance. In practical applications, it may be necessary to flexibly adjust the dropout rate according to the performance of the model and the characteristics of the dataset.

[0073] In step S3, the specific steps for optimizing with the crayfish optimization algorithm include: S33: Use the mean squared error as the fitness function of the LSTM network model.

[0074] In this embodiment, the expression formula for the mean squared error is:

[0075] in, is the mean square error, is the actual value, that is, The actual output in samples is is the predicted value, i.e., the model The output predicted on samples.

[0076] S34: Initialize the population and randomly generate individual crayfish positions. Each crayfish individual represents a LSTM network model hyperparameter group.

[0077] S35: Calculate the individual fitness values ​​of the initialized population and randomly generate the ambient temperature value.

[0078] S36: An iteration mechanism is set. Each time a new population is formed, an ambient temperature is randomly selected, and population one or population two is obtained according to the iteration mechanism.

[0079] In step S36, in step S36, the iteration mechanism specifically includes: S361: When the ambient temperature is higher than the preset ambient temperature threshold, the individuals of the population find the location of the summer cave based on the current global optimal position and the individual historical optimal position.

[0080] In step S361, the summer cave location expression formula is:

[0081] in, is the location of the cave, is the individual's best historical position, is the previous global optimal position.

[0082] S362: Individuals of the population move to the location of the summer cave. If the locations of the individuals are the same, by comparing the individual fitness values, the individuals with low fitness will randomly select the location of another crayfish to adjust their positions to compete for the summer cave. When the positions of all individuals are updated, population one is obtained.

[0083] S363: When the ambient temperature is lower than or equal to the preset ambient temperature threshold, the optimal position of the individual in the current population is used as the food position.

[0084] S364: Determine the size of the food according to the individual fitness values ​​of the current population, and select a foraging method to forage according to the determination result, thereby obtaining population two.

[0085] In step S364, the formula for determining the size of the food is:

[0086] in, It is the identification value of food size. is a food factor, which is a constant, is a random number within the range of the fitness value of the current individual, the fitness value of the food. When it is the case, it is determined that the food is larger. When it is the case, it is determined that the food is smaller.

[0087] S37: Calculate the fitness value of each individual in the first population or the second population. If the output fitness value is higher than the preset fitness threshold, output the optimal hyperparameters; otherwise, continue the iteration until the fitness value is higher than the preset fitness threshold or the maximum number of iterations is reached to obtain the optimal hyperparameters.

[0088] S4: Input the target data into the optimal recognition model to obtain the behavior category.

[0089] In this embodiment, through the accurate recognition model, the human behavior category presented in the target data can be quickly and accurately identified. The accuracy of the recognition result provides indispensable information support for many subsequent applications. In the intelligent monitoring system, the application of human behavior recognition is particularly crucial. It can help the monitoring system achieve automated monitoring, accurately capture and alarm abnormal behaviors, thereby effectively improving public safety and prevention capabilities. In the field of sports training, human behavior recognition also plays an important role. Coaches and athletes can use it to conduct a detailed analysis of the athletes' movements, accurately locate technical shortcomings, and then formulate more targeted training plans to continuously improve sports performance. In addition, the application prospects of human behavior recognition are far from limited to this. It has extensive potential applications in many fields such as human-computer interaction, virtual reality, and game entertainment.

[0090] In this embodiment, the present invention constructs an LSTM network model optimized by the crayfish optimization algorithm, inputs training data for training to obtain the optimal behavior recognition model, extracts features from the target video data, obtains the global feature representation through the inverse weighting method to obtain the target data, and inputs the target data into the optimal behavior recognition model to obtain the behavior category, solving the problem of inaccurate behavior recognition. The global feature representation can more comprehensively reflect the behavior information in the video, improving the accuracy and reliability of the optimal behavior recognition model. The crayfish optimization algorithm reduces the time required for training the LSTM network model, improves the overall operation efficiency, and the robustness, thereby further improving the accuracy of recognition.

[0091] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0092] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A human behavior recognition method based on machine learning, characterized in that: include: S1: Obtain target video data, collect video data containing different human behaviors and corresponding behavior categories as historical data; S2: Preprocess the target video data and the historical data respectively to obtain target data and training data; S3: construct an LSTM network model optimized based on the crayfish optimization algorithm, and input the training data for training to obtain an optimal behavior recognition model; S4: Input the target data into the optimal recognition model to obtain the behavior category.

2. The human behavior recognition method based on machine learning according to claim 1, characterized in that: In step S2, the specific steps of preprocessing include: S21: extracting video frames from the video data at preset time intervals, processing the video frames to obtain enhanced video frames, and using a convolutional neural network to extract features from the enhanced video frames to obtain a feature sequence; S22: performing normalization and smoothing processing on the feature sequence to obtain an enhanced feature sequence; S23: Calculate the distance between the enhanced feature sequences, and obtain weighted features using an inverse distance weighting method according to the distance; S24: Summarize the weighted features by the average value aggregation method to obtain a global feature representation.

3. The human behavior recognition method based on machine learning according to claim 2 is characterized in that: In step S21, the specific steps of processing the video frame include: S211: Enhance the video frame using a tone transformation method; S212: Use Gaussian blur to perform noise filtering on the video frame.

4. The human behavior recognition method based on machine learning according to claim 2 is characterized in that: In step S22, the specific steps of the smoothing process include: S221: removing abnormal values ​​in the feature sequence, and supplementing the missing values ​​in the feature sequence by interpolation; S222: Use a moving average method to smooth the feature sequence.

5. The human behavior recognition method based on machine learning according to claim 2 is characterized in that: In step S23, the inverse distance weighted expression formula is: in, is the value of the unknown point x to be estimated, It is The data value of the known point, is the unknown point x and the The distance between known points, is the exponential of the weight.

6. The human behavior recognition method based on machine learning according to claim 1, characterized in that: In step S3, the specific steps of building the LSTM network model include: S31: Setting a network architecture including an input layer, an LSTM layer, and an output layer; the input layer is used to receive the target data; the LSTM layer is used to capture the time dependency of the target data and identify the dynamic changes of the behavior; the output layer is used to convert the output of the LSTM layer into a specific behavior category; S32: Two Dropout layers are set in the network architecture; the two Dropout layers are respectively set between the input layer and the LSTM layer, and between the LSTM layer and the output layer; the two Dropout layers are used to prevent the model from overfitting.

7. The human behavior recognition method based on machine learning according to claim 1, characterized in that: In step S3, the specific steps of optimizing the crayfish optimization algorithm include: S33: Use mean square error as the fitness function of the LSTM network model; S34: Initialize the population and randomly generate individual crayfish positions. Each crayfish individual represents a LSTM network model hyperparameter group. S35: Calculate the individual fitness values ​​of the initialized population and randomly generate the environmental temperature value; S36: setting an iteration mechanism, each time a new population is formed, randomly setting an environmental temperature, and obtaining population one or population two according to the iteration mechanism; S37: Calculate the fitness value of each individual in the population one or the population two. If the output fitness value is higher than the preset fitness threshold, output the optimal hyperparameter. Otherwise, continue to iterate until the fitness value is higher than the preset fitness threshold or the maximum number of iterations is reached to obtain the optimal hyperparameter.

8. The human behavior recognition method based on machine learning according to claim 7 is characterized in that: In step S36, the iteration mechanism specifically includes: S361: When the ambient temperature is higher than the preset ambient temperature threshold, the individuals of the population find the location of the summer cave according to the current global optimal position and the individual historical optimal position; S362: Individuals of the population move to the location of the summer cave. If the locations of the individuals are the same, by comparing the individual fitness values, the individual with the lower fitness will randomly select the location of another crayfish to adjust its position to compete for the summer cave. When the positions of all individuals are updated, population one is obtained. S363: When the ambient temperature is lower than or equal to the preset ambient temperature threshold, using the optimal position of the individual in the current population as the food position; S364: Determine the size of the food according to the individual fitness values ​​of the current population, and select a foraging method to forage according to the determination result, thereby obtaining population two.

9. The human behavior recognition method based on machine learning according to claim 7, characterized in that: In step S361, the summer cave location expression formula is: in, is the location of the cave, is the individual's best historical position, is the previous global optimal position.

10. The human behavior recognition method based on machine learning according to claim 8, characterized in that: In step S364, the formula for determining the size of the food is: in, It is the identification value of food size. is the food factor, which is a constant, is a random number, between, The current individual's fitness value, The fitness value of food, when When , the food is judged to be small.

Citation Information

Patent Citations

  • Building indoor point cloud segmentation method based on local feature enhanced Point Net + + network

    CN115115839A

  • Computer network security analysis method and system based on big data

    CN118555117A

  • Rehabilitation training data processing method and device

    CN119580943A

  • Wind power prediction method and system for optimizing deep transformer network

    US20220197233A1

Cited By

  • Self-adaptive multi-mode fusion-based behavior identification method for intelligent robot with body

    CN120808036A