Monitoring video processing method and device, equipment, storage medium and program product
By performing frame-based and feature extraction of surveillance videos, using the spider bee algorithm to search for global optimal individuals of the population as the initial clustering center, the problem of traditional K-mean algorithm being sensitive to the selection of initial clustering centers is solved, and more efficient risk point identification and security management are achieved.
Patent Information
- Application Number
- CN202510427707.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-04
AI Technical Summary
The traditional K-mean clustering algorithm is sensitive to the selection of initial clustering centers, which leads to low clustering accuracy of monitoring video data and makes it difficult to accurately identify risk points.
The monitoring video is framed and feature information is extracted to form a population of the spider bee algorithm. The global optimal individual of the population is searched through the spider bee algorithm as the initial clustering center. The search path is dynamically adjusted with the two-way search strategy to improve clustering accuracy.
It improves the clustering accuracy of monitoring video data and the accuracy of identifying risk points, and improves the efficiency of security management.
Smart Images

Figure CN120259948A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical fields of video processing and fintech, and in particular, to a method, device, equipment, storage medium and program product for processing surveillance videos. Background Art
[0002] Video surveillance systems have been widely deployed in places such as banks and financial institutions. These video surveillance systems can record activities in key areas in real time and play an important role in daily security management and risk control. Due to the continuous operation of the video surveillance system, a large amount of surveillance video data has been generated. It is necessary to analyze this surveillance video data to quickly identify risk points and respond.
[0003] In order to identify preset risk point situations from videos, a clustering algorithm can be used to perform clustering analysis on surveillance video data. However, the traditional K-means clustering algorithm is very sensitive to the selection of the initial clustering center. If the initial clustering center is not selected well, it may fall into a local optimal solution, resulting in poor clustering effects. When facing large-scale surveillance video data, the randomness of the selection of the initial clustering center will amplify this problem, resulting in low clustering accuracy and difficulty in accurately monitoring and locating risk points. Summary of the Invention
[0004] Embodiments of the present invention provide a method, device, equipment, storage medium and program product for processing surveillance videos, which can improve clustering accuracy and accurately monitor and locate risk points.
[0005] In a first aspect, the method for processing surveillance videos provided by embodiments of the present invention includes:
[0006] Performing frame splitting on the surveillance video to obtain a plurality of video frames;
[0007] Obtaining the feature information of each video frame and vectorizing the feature information of each video frame to obtain the feature vector of each video frame;
[0008] Forming a population of the spider wasp algorithm according to the feature vectors of each video frame; wherein, the population includes a plurality of individuals, each individual includes a set number of video frames, the set of feature vectors of all video frames of each individual serves as the solution space position of the individual, and the solution space positions of all individuals in the population constitute the solution space of the population;
[0009] Using the spider wasp algorithm to start the individual search of the current round; wherein, the individual search process of the current round includes:
[0010] Setting a plurality of search agent individuals; each search agent individual includes a set number of search video frames, and the set of feature vectors of all search video frames of each search agent individual serves as the solution space position of the search agent individual;
[0011] Each search agent individual starts from its current solution space position and moves the solution space position according to the set path algorithm by adjusting the numerical values of the feature vectors. During the movement, the clustering fitness of the individuals with matching solution space positions is calculated. The clustering fitness is the clustering effect value of clustering the surveillance video when the feature vectors of each video frame of the individual are used as the clustering centers respectively.
[0012] When the local search stop condition is met, according to the clustering fitness of the matched individuals, the optimal individual and the worst individual in the current round of local search are determined. Among them, the solution space positions of the optimal individual and the worst individual are used in the next round of search process to apply to the set path algorithm to adjust the movement path of the search agent individual.
[0013] Return to start the individual search in the next round. Until the global search stop condition is reached, the finally determined global optimal individual is used as the initial clustering center of the surveillance video.
[0014] In a second aspect, the surveillance video processing device provided by the embodiments of the present invention includes:
[0015] A frame division module for dividing the surveillance video into frames to obtain a plurality of video frames.
[0016] A vectorization module for obtaining the feature information of each video frame and vectorizing the feature information of each video frame to obtain the feature vector of each video frame.
[0017] A generation module for forming a population of the spider wasp algorithm according to the feature vectors of each video frame. The population includes a plurality of individuals. Each individual includes a set number of video frames. The set of feature vectors of all video frames of each individual is used as the solution space position of the individual. The solution space positions of all individuals in the population constitute the solution space of the population.
[0018] A search module for starting the individual search in the current round by using the spider wasp algorithm. The individual search process in the current round includes:
[0019] Set a plurality of search agent individuals. Each search agent individual includes a set number of search video frames. The set of feature vectors of all search video frames of each search agent individual is used as the solution space position of the search agent individual.
[0020] Each search agent individual starts from its current solution space position and moves the solution space position according to the set path algorithm by adjusting the numerical values of the feature vectors. During the movement, the clustering fitness of the individuals with matching solution space positions is calculated. The clustering fitness is the clustering effect value of clustering the surveillance video when the feature vectors of each video frame of the individual are used as the clustering centers respectively.
[0021] When the local search stop condition is met, determine the optimal individual and the worst individual for the local search in the current round according to the clustering fitness of the matched individuals; among them, the solution space positions of the optimal individual and the worst individual are used in the next round of search process to apply to the set path algorithm to adjust the movement path of the search agent individuals.
[0022] Return to start the individual search in the next round, and until the global search stop condition is reached, use the finally determined global optimal individual as the initial clustering center of the surveillance video.
[0023] In a third aspect, the electronic device provided by an embodiment of the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the surveillance video processing method according to any embodiment of the present invention.
[0024] In a fourth aspect, the computer-readable storage medium provided by an embodiment of the present invention stores a computer program, and when the program is executed by a processor, it implements the surveillance video processing method according to any embodiment of the present invention.
[0025] In a fifth aspect, the computer program product provided by an embodiment of the present invention includes a computer program, and when the computer program is executed by a processor, it implements the surveillance video processing method according to any embodiment of the present invention.
[0026] In the embodiments of the present invention, the surveillance video is framed, the feature information of each video frame is extracted and vectorized. Through feature vectorization, the video frames are converted into quantifiable data, which is convenient for subsequent clustering analysis. This method can process massive video data more efficiently and reduce the computational complexity. The spider wasp algorithm is used to construct a population. Each individual in the population consists of a group of video frames, and this group of video frames forms the solution space position of the corresponding individual. The solution space positions of all individuals in the population constitute the solution space of the population. The spider wasp algorithm has strong global search ability. Using the spider wasp algorithm to search for the globally optimal individual in the population and taking the final globally optimal individual as the initial clustering center of the surveillance video can effectively avoid the sensitivity problem of the traditional K-means algorithm to the selection of the initial clustering center and improve the clustering accuracy. Multiple search agent individuals are set in the population. Each search agent individual starts from the current solution space position, moves by adjusting the numerical values of the feature vectors, and calculates the clustering fitness of the individuals matching the solution space position, providing interpretability for the execution of the algorithm. In each round of search, each search agent individual shares the search results, and each search agent individual dynamically adjusts its own movement path according to the solution space positions of the optimal individual and the worst individual, that is, dynamically adjusts the search path by combining the bidirectional search strategy, which can better adapt to the changes in the surveillance video data. The multiple search agent individuals cooperate to improve the accuracy and speed of the search. When the global search stop condition is met, the finally determined globally optimal individual is used as the initial clustering center of the surveillance video. By determining the globally optimal individual, a high-quality initial clustering center is provided for the subsequent clustering analysis. Using the initial clustering center to cluster all video frames can finally identify the risk points in the surveillance video more quickly and accurately, and improve the safety management efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0028] Figure 1 is a flowchart of a method for processing surveillance video provided by an embodiment of the present invention;
[0029] Figure 2 is a schematic diagram of the process of finding the globally optimal individual provided by an embodiment of the present invention;
[0030] Figure 3 is an example diagram of a method for processing surveillance video provided by an embodiment of the present invention;
[0031] Figure 4It is a schematic structural diagram of a monitoring video processing device provided by an embodiment of the present invention;
[0032] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0033] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] Figure 1 It is a schematic flowchart of a monitoring video processing method provided by an embodiment of the present invention. The monitoring video processing method provided by the embodiment of the present invention is applicable to scenarios where a large amount of monitoring video data is analyzed and processed. The monitoring video processing method can be executed by the monitoring video processing device provided by the embodiment of the present invention, and the device can be implemented in a software and / or hardware manner. In a specific embodiment, the device can be integrated in an electronic device, and the electronic device can be a computer, a server, etc. The following embodiments will be described by taking the integration of the monitoring video processing device in an electronic device as an example. Refer to Figure 1 , the monitoring video processing method of this embodiment may include the following steps:
[0036] Step 110, perform frame splitting on the monitoring video to obtain a plurality of video frames.
[0037] The surveillance video can be continuous video data recorded in real time by surveillance cameras in places such as banks and financial institutions. These video data can contain all activities within a specific place and a specific time period. Frame division is the process of decomposing continuous surveillance video data into a series of static image frames, and each frame represents the picture of the video at a certain point in time. A video frame is a single static image extracted from the surveillance video, and one video frame corresponds to the picture of the surveillance video at a specific moment.
[0038] By dividing the surveillance video into frames, all continuous surveillance video data are decomposed into static image frames at individual time points, providing basic data units for subsequent feature extraction and analysis processing. This process enables each frame of the subsequent video to be processed and analyzed separately, which helps to analyze and identify activities and potential risk points in the video in a more fine-grained manner.
[0039] Step 120: Obtain the feature information of each video frame and vectorize the feature information of each video frame to obtain the feature vector of each video frame.
[0040] Feature information is important information extracted from a video frame to describe the content of the frame. These features can include information such as color, texture, shape, pixel values, etc. The specific type or dimension of the features depends on the application requirements. Vectorization is the process of converting the extracted feature information into a specific form of array or vector to make it suitable for numerical processing and calculation. A feature vector is the result of vectorizing the feature information of a video frame. A feature vector is a numerical representation form of feature information, usually presented in the form of a multi-dimensional vector, which is convenient for processing and analysis in algorithms.
[0041] Important feature information describing the content of each video frame can be extracted, and then these feature information are converted into vectors in numerical form, that is, feature vectors. Through vectorization, the feature information of each video frame is represented as a numerical array convenient for processing. Feature vectors can be used for subsequent algorithm analysis and optimization steps. The purpose of doing this is to simplify the complex information of video frames into a unified numerical form for effective processing and clustering analysis by electronic devices.
[0042] Step 130: Form a population of the spider wasp algorithm according to the feature vectors of each video frame; where the population includes multiple individuals, each individual includes a set number of video frames, the set of feature vectors of all video frames of each individual is used as the solution space position of the individual, and the solution space positions of all individuals in the population constitute the solution space of the population.
[0043] The spider wasp algorithm is a bionic optimization algorithm that imitates the behaviors of spiders and wasps to perform global search and local search for solving complex optimization problems. In bionic algorithms, a population is the basic unit for the algorithm to run, which consists of multiple individuals, and each individual represents a possible solution to the problem. A solution refers to a candidate answer searched in the solution space, corresponding to the possible result of the optimization problem. For example, in a function optimization problem, a solution is a combination of variables that satisfies the constraints. In a clustering problem, a solution is a set of positions of the initial clustering centers.
[0044] Specifically, in this embodiment, each individual in the spider wasp algorithm population is composed of a set number of video frames, that is, each individual includes a group of video frames, and the feature vectors corresponding to this group of video frames represent a set of possible clustering centers, and the set number is the number of clustering centers. For example, each video frame has 128 - dimensional features, and each clustering scheme includes 3 clustering centers, then each individual can be represented by the feature vectors of 128 - dimensional features of 3 video frames. For example, an individual can be represented as (the feature vector of 128 - dimensional features of video frame 1, the feature vector of 128 - dimensional features of video frame 2, the feature vector of 128 - dimensional features of video frame 3).
[0045] The solution - space position of an individual refers to the specific position of the individual in the solution space, which is determined by the set of feature vectors of all video frames included in the individual. The solution - space position can be used to measure the quality of the individual. The clustering fitness reflects the clustering effect of this set of clustering centers. The solution space of the population refers to the set composed of the solution - space positions of all individuals in the population, which covers all possible combinations of clustering centers. The solution space of the population is the scope for the algorithm to search.
[0046] Specifically, population parameters can be set. The population parameters can include the number of iterations, the population size, and the optimization boundary. According to the set population parameters and the feature vectors of each video frame, a population of the spider wasp algorithm is formed. Among them, the number of iterations refers to the number of times the algorithm repeats during the optimization process. The population size refers to the number of individuals included in the population. A larger population can cover a wider search space, but it will also increase the computational amount. The optimization boundary is the boundary of the solution space. In a surveillance video, the optimization boundary corresponds to the minimum and maximum values of the feature values, for example, the value range of brightness or color intensity. By setting these parameters, the behavior and performance of the algorithm can be adjusted and controlled to make it more efficient in finding the globally optimal individual.
[0047] Step 140: Use the spider wasp algorithm to initiate the individual search for the current round. Among them, the individual search process for the current round includes: setting multiple search agent individuals; each search agent individual includes a set number of search video frames, and the set of feature vectors of all the search video frames of each search agent individual serves as the solution space position of this search agent individual; each search agent individual starts from its own current solution space position and moves the solution space position according to the set path algorithm by adjusting the feature vector values, and calculates the clustering fitness of the individuals matching the solution space position during the movement; the clustering fitness is the clustering effect value when the feature vectors of each video frame of the individual are used as the clustering centers respectively for clustering the surveillance video; when the local search stop condition is met, determine the optimal individual and the worst individual of the local search performed in the current round according to the clustering fitness of the matched individuals; among them, the solution space positions of the optimal individual and the worst individual are used in the next round of search process to adjust the movement path of the search agent individual by applying the set path algorithm.
[0048] The individuals formed in Step 130 can be understood as the inherent individuals of the population. A search agent individual refers to an individual used to search for the inherent individuals in the population in the solution space in the spider wasp algorithm, which is an additional individual set in the population. Each search agent individual contains a set number of search video frames. The search video frames can be the video frames obtained by splitting the surveillance video, or the video frames randomly created within the solution space range. The set of feature vectors of all the search video frames included in each search agent individual constitutes the solution space position of this search agent individual. In this embodiment, the process of the search agent individual performing individual search is realized by adjusting the feature vector values of the video frames of the search agent individual. As the feature vector values of the video frames of the search agent individual are adjusted, the solution space position of the search agent individual will change. That is, in this embodiment, the solution space position of the search agent individual is variable, and the solution space position of the inherent individuals in the population is fixed. The search agent individual can be analogized to the wasp in the spider wasp algorithm, and the globally optimal individual can be analogized to the spider in the spider wasp algorithm. The position update of the search agent individual simulates the search behavior of the wasp looking for the spider in the spider wasp algorithm.
[0049] The set path algorithm refers to the strategy used to guide the movement of the search agent individual in the solution space. This algorithm simulates the behavior of spider wasps in nature and adjusts the position of the search agent individual through specific update rules. The individuals with matching solution space positions refer to the individuals that are closest to the position of the current search agent individual in the solution space. The distance of the positions can be measured by the vector similarity between the current search agent individual and the individuals around the current search agent individual. The individual that is closest to the position of the current search agent individual can be the individual with the highest vector similarity to the current search agent individual. The current search agent individual can be any one of the multiple search agent individuals.
[0050] The local search stop condition refers to the condition under which the search agent individual stops moving during the local search phase, which may include that the total number of steps of the search agent individual reaches the set number of steps, and / or the total step length reaches the set step length.
[0051] The optimal individual is the individual with the highest clustering fitness obtained after the current round of search. The worst individual is the individual with the lowest clustering fitness obtained after the current round of search. The clustering fitness is the clustering effect value when the feature vectors of each video frame of the individual are used as the clustering centers to cluster the surveillance video. The clustering effect value can be calculated using the compactness and / or dispersion of the clustering. The compactness can be represented by the average distance between the clustering center and its members (i.e., video frames) below, that is, the mean within-class distance. The smaller the mean within-class distance, the closer the clustering center is to its members below, the tighter the clustering, and the better the clustering effect. The dispersion can be represented by the average distance between different clustering centers, that is, the between-class distance. The larger the between-class distance, the better the discrimination between different clusters, and the better the clustering effect.
[0052] The solution space positions of the optimal individual and the worst individual are used in the next round of search process to apply to the set path algorithm to adjust the moving path of the search agent individual, that is, the bidirectional search strategy. The bidirectional search strategy is a hybrid search strategy used to guide the search agent individual to move towards the solution space position of the optimal individual and away from the solution space position of the worst individual in the next round of search process. After each round of search, according to the solution space positions of the optimal individual and the worst individual, the algorithm dynamically adjusts the moving path of the search agent individual. This dynamic adjustment mechanism enables the search agent individual to search more efficiently in the solution space, avoid falling into local optimal solutions, and at the same time accelerate the convergence speed.
[0053] Step 150, return to start the individual search for the next round. Until the global search stop condition is reached, the finally determined global optimal individual is used as the initial clustering center of the surveillance video.
[0054] The global search stop condition refers to the condition under which the algorithm stops searching during the global search phase. These conditions are used to determine when to end the global search process. The final global optimal individual is the individual with the highest clustering fitness among the search results of multiple rounds.
[0055] The initial clustering center includes a set of clustering centers, and each clustering center in this set of clustering centers has the corresponding feature vector of the video frame. After obtaining the initial clustering center, the feature vectors of the video frames corresponding to the initial clustering center can be used to cluster all the video frames of the surveillance video to obtain the target clustering center. According to the target clustering center, all the video frames are clustered into clusters, and the feature vectors of each cluster are matched with the template vectors of abnormal behaviors to identify abnormal behaviors in the surveillance video.
[0056] Clustering is to group similar video frames into one category according to the feature vectors of video frames, and the data points in each category are close to each other in the feature space. In this embodiment, the K-means clustering algorithm can be used to cluster all video frames of the surveillance video. K-means clustering is a commonly used clustering algorithm, and its goal is to divide data points into K clusters to minimize the within-cluster distance. In conventional K-means clustering, the initial cluster centers are randomly selected. In this embodiment, a set of cluster centers found by the spider wasp algorithm improved by the bidirectional search strategy is used as the initial cluster centers of K-means clustering, so as to improve the quality of the initial cluster centers. The target cluster centers are the final cluster centers obtained after K-means iteration converges, representing the stable division of the data distribution.
[0057] The process of specifically performing K-means clustering on video frames according to the initial cluster centers is as follows:
[0058] (1) Assign video frames to the nearest cluster center. That is, calculate the distance (such as Euclidean distance) between each video frame and all current cluster centers, and assign each video frame to the nearest cluster center to obtain clusters composed of video frames.
[0059] (2) Update the cluster centers. For each cluster, calculate the average value of the feature vectors of all video frames assigned to this cluster, and use this average value as the new cluster center.
[0060] Repeat steps (1) and (2) until the cluster centers no longer change or reach the preset number of iterations.
[0061] After obtaining the target cluster centers, all video frames can be clustered based on the target cluster centers. The distance between each video frame and each target cluster center can be calculated according to the feature vector of each video frame and the feature vector of the video frame corresponding to each target cluster center, and each video frame is assigned to the target cluster center closest to itself, so as to obtain multiple clusters.
[0062] Specifically, in this embodiment, the surveillance video can be the surveillance video of a bank branch, and the abnormal behavior can be the behavior that does not conform to the bank's operation specifications, security management system or normal business process, and these behaviors may pose potential threats to the bank's security, operation order or customer rights and interests. Specifically, the abnormal behavior can include abnormal behavior in the operation room and abnormal behavior in the repository. The abnormal behavior in the operation room can include: single-person operation, opening multiple safes simultaneously, entering and leaving the operation room with a bag-shaped object (such as a shopping bag, handbag, backpack, etc.). The abnormal behavior in the repository can include: single-person entering and leaving the repository, entering and leaving the repository with a bag-shaped object (such as a shopping bag, handbag, backpack, etc.), smuggling items, too short inventory checking time, suspected formal inventory checking, etc. The template vector of the abnormal behavior can be a pre-defined feature vector database of the abnormal behavior, and this database can be created manually or automatically generated according to expert rules. There can be multiple template vectors of the abnormal behavior, and each template vector can represent a type of abnormal behavior, and each type of abnormal behavior's template vector can be preset with a type label corresponding to the abnormal behavior. Type labels such as: single-person entering and leaving the repository, entering and leaving the repository with a bag-shaped object.
[0063] The feature vectors of the video frames included in each cluster can be averaged to obtain the feature vector corresponding to the cluster. It is also possible to use the feature vectors of the video frames corresponding to the cluster centers of each cluster as the feature vectors corresponding to the clusters. The feature vectors of each cluster are matched with the pre-defined template vectors of abnormal behaviors, and the matching process can be achieved by calculating similarity or distance. The abnormal behavior in the surveillance video can be identified according to the type label of the template vector of the abnormal behavior matched by the feature vector of each cluster. For example, if the feature vector of a certain cluster matches the template vector with the type label of single-person entering and leaving the repository, it indicates that there is an abnormal behavior of single-person entering and leaving the repository in the surveillance video. The identified abnormal behavior can be reported to the management personnel for further measures to be taken for processing and intervention.
[0064] In this embodiment, the surveillance video is framed, the feature information of each video frame is extracted and vectorized. Through feature vectorization, the video frame is converted into quantifiable data, which is convenient for subsequent clustering analysis. This method can process massive video data more efficiently and reduce the computational complexity. The spider wasp algorithm is used to construct a population. Each individual in the population consists of a set of video frames. This set of video frames forms the solution space position of the corresponding individual. The solution space positions of all individuals in the population constitute the solution space of the population. The spider wasp algorithm has strong global search ability. Using the spider wasp algorithm to search for the global optimal individual in the population and taking the final global optimal individual as the initial clustering center of the surveillance video can effectively avoid the sensitivity problem of the traditional K-means algorithm to the selection of the initial clustering center and improve the clustering accuracy. Multiple search agent individuals are set in the population. Each search agent individual starts from the current solution space position, moves by adjusting the numerical value of the feature vector, and calculates the clustering fitness of the individual that matches the solution space position, providing interpretability for the execution of the algorithm. In each round of search, each search agent individual shares the search results. Each search agent individual dynamically adjusts its own movement path according to the solution space positions of the optimal individual and the worst individual, that is, combines the two-way search strategy to dynamically adjust the search path, which can better adapt to the changes in the surveillance video data. The cooperation of multiple search agent individuals improves the accuracy and speed of the search. When the global search stop condition is met, the finally determined global optimal individual is used as the initial clustering center of the surveillance video. Through the determination of the global optimal individual, a high-quality initial clustering center is provided for the subsequent clustering analysis. Using the initial clustering center to cluster all video frames can finally identify the risk points in the surveillance video more quickly and accurately, improving the safety management efficiency.
[0065] In the embodiment of the present invention, the spider wasp algorithm improved by the two-way search strategy is used to cluster video frames. First, the principle of the spider wasp algorithm is introduced below. The spider wasp algorithm can be mainly divided into a search stage, a following and escaping stage, and a nest building stage, where:
[0066] (1) In the search stage, the behavior of a wasp looking for a spider is simulated, and the behavior of the wasp searching for the area where the spider fell after losing the spider's trail is simulated.
[0067] For wasp i, it is initially set according to the following formula:
[0068]
[0069] represents the position of wasp i at the current time t, represents the lower limit of the optimization boundary, represents the upper limit of the optimization boundary, and r is a random number between 0 and 1.
[0070] The simulation equation for the behavior of the wasp looking for a spider is as follows:
[0071]
[0072] Among them, represents the position of wasp i at the next moment t + 1. a and b represent two random individuals from the population, which are used to determine the search direction of wasp i. represents the position of individual a in the population at time t, represents the position of individual b in the population at time t. μ1 represents the current position update factor of wasp i, and μ1 is determined by the following formula;
[0073] μ1 = |rn| * r1; Formula 2
[0074] r1 is a random number between 0 and 1, and rn is a random number generated by a normal distribution.
[0075] After the wasp loses the trail of the spider, the simulation equation for the behavior of searching for the area where the spider fell is as follows:
[0076]
[0077] μ2 = B * cos(2πl); Formula 4
[0078]
[0079] Among them, c is the individual selected in the area where the spider fell, represents the position of individual c in the population at time t. μ2 is the moving step size control factor, represents the upper limit of the optimization boundary, represents the lower limit of the optimization boundary, represents a multi-dimensional randomly initialized vector between 0 and 1. B is the calculation coefficient, and l is a randomly generated integer between -2 and 1.
[0080] Finally, the next position of wasp i is randomly selected between Formula 1 and Formula 3, and the selection formula is as follows:
[0081]
[0082] Among them, r3 and r4 are random numbers in [0, 1].
[0083] (2) Follow - and - flee phase: Define and simulate two behaviors of wasps. When the speed of the wasp is faster than that of the spider by \(C>0.5\), simulate the wasp following the spider; when the speed of the spider is faster than that of the wasp by \(C < 0.5\), the wasp makes a strategic retreat, simulating the wasp fleeing from the spider. Here, \(C\) is the distance - control factor that determines the speed of the wasp. It starts from speed 2 and linearly decreases to zero. If \(C>0.5\), it means the wasp is actively approaching the spider, simulating the behavior of the wasp catching the spider; if \(C < 0.5\), it indicates that the moving speed of the wasp slows down, which simulates that if the spider moves quickly, the wasp needs to adjust its strategy to adapt to this rapid change.
[0084] The simulation equation for the following behavior is as follows:
[0085]
[0086] Where, represents a vector of values randomly generated in the interval \([0, 1]\), \(r6\) is a random number in the interval \([0, 1]\), where \(T\) and \(T\) max represent the current iteration number and the maximum iteration number respectively.
[0087] The simulation equation for the fleeing behavior is as follows:
[0088]
[0089] Where, is a normal - distribution vector between \(k\) and \(-k\). Use formula 10 to generate \(k\) to increase the distance between wasp \(i\) and the spider.
[0090]
[0091] Use the following formula to make a random balance between the following and fleeing behaviors:
[0092]
[0093] Adopt the following formula to balance the search behavior and the following behavior:
[0094]
[0095] Where, \(p\) is a random number in \([0, 1]\).
[0096] (3) Nest - building phase: Simulate the behavior of the wasp building a nest and pulling the spider into the nest.
[0097] It can build a nest at the position of the optimal solution. The simulation equation for this nest - building behavior is as follows:
[0098]
[0099] Where, Indicates the position information of the optimal solution searched by the wasp i.
[0100] It can nest at the position of randomly selected individuals. The simulation equation of this nesting behavior is as follows:
[0101]
[0102] Among them, r3 is a random number between 0 and 1, and γ is a number randomly generated by Lévy flight. Indicates a binary vector.
[0103] Binary vector It is calculated according to the following formula:
[0104]
[0105] Among them, and are random vectors between 0 and 1.
[0106] Randomly exchange Formula 13 and Formula 14 using the following formula:
[0107]
[0108] Finally, the trade-off between search and following and nesting behaviors is achieved by the following formula.
[0109]
[0110] Next, in combination with the above three main stages of the spider wasp algorithm, the process of finding the global optimal individual in the embodiments of the present invention will be described. As Figure 2 shown, the method of this embodiment may include:
[0111] Step 210, determine whether the current round is the first round of search. If it is the first round of search, execute Step 221; if it is not the first round of search, execute Step 223.
[0112] Step 221, within the solution space range of the population, randomly create multiple sets of feature vector sets of search video frames. A set of search video frames constitutes a search agent individual, and a set of search video frames includes a set number of search video frames. The randomly created multiple sets of feature vector sets of search video frames are used as the current solution space positions of the corresponding search agent individuals.
[0113] Since the current round is the first round of search, search agent individuals need to be created. There are multiple created search agent individuals. After these search agent individuals complete local search (i.e., meet the local search stop condition) respectively, they will share the local search results with each other to optimize the search process and improve the search efficiency.
[0114] Step 222: Each search agent individual takes the individual with the higher clustering fitness among the two individuals randomly selected by itself as the reference individual. Each search agent individual determines the vector adjustment direction of each search agent individual according to the solution space position of its respective reference individual, and each search agent individual adjusts the numerical value of its feature vector according to its own vector adjustment direction and a preset first adjustment step length to move towards the solution space position of its respective reference individual.
[0115] Since the current round is the first round of search and there is no search result for reference. Therefore, each search agent individual can randomly select two individuals, identify the individual with the higher clustering fitness from the two randomly selected individuals, and take the individual with the higher clustering fitness as the reference individual of the corresponding search agent individual. Each search agent individual determines its own vector adjustment direction according to the solution space position of the reference individual, and the vector adjustment direction is, for example, to increase or decrease the vector value. Specifically, the adjustment process will cause the feature vector of each search agent individual to approach the feature vector of the reference individual. The first adjustment step length is the step length for each search agent individual to move each time when moving towards the solution space position of the reference individual, that is, the adjustment amplitude of the feature vector. This step length is preset and used to control the moving speed of the search agent individual.
[0116] By approaching the reference individual with the higher clustering fitness, the search agent individual can explore the solution space more efficiently and reduce ineffective search.
[0117] Step 223: Take the multiple search agent individuals in the previous round as the multiple search agent individuals in the current round, and take the solution space positions of the multiple search agent individuals at the end of the previous round of search as the current solution space positions of the corresponding search agent individuals.
[0118] When the current round is not the first round of search, the multiple search agent individuals in the previous round will be retained, and the solution space positions of each search agent individual at the end of the previous round of search will be retained. By inheriting the search agent individuals and their solution space positions in the previous round, the algorithm can maintain the continuity of the search and avoid losing the existing search information due to re-initialization. The inheritance mechanism enables the search agent individuals to continue to optimize based on the previous round, reduces repeated search, and improves the search efficiency.
[0119] Step 224: Each search agent individual determines the vector adjustment direction of each search agent individual according to the solution space positions of the optimal individual and the worst individual in the local search of the previous round.
[0120] The optimal individual of the previous local search is the individual with the highest clustering fitness determined by comprehensively considering the local search results of each search agent individual after the previous local search of each search agent individual. Its solution space position is used to guide the search agent individuals to approach this position. The worst individual of the previous local search is the individual with the lowest clustering fitness determined by comprehensively considering the local search results of each search agent individual after the previous local search of each search agent individual. Its solution space position is used to guide the search agent individuals to move away from this position.
[0121] After the previous local search is completed, the algorithm determines the solution space positions of the optimal and worst individuals. Each search agent individual calculates its own vector adjustment direction based on the solution space positions of the optimal and worst individuals. Specifically, the feature vector of the search agent individual will approach the feature vector of the optimal individual and move away from the feature vector of the worst individual. By approaching the optimal individual, the search agent individual can explore the solution space more efficiently and reduce ineffective searches. By moving away from the position of the worst individual, the search agent individual can avoid falling into local optimal solutions and thus improve its global search ability.
[0122] Step 225, identify the number of search agent individuals within the preset range of the solution space position of the optimal individual of the previous local search.
[0123] The preset range is a specific range set with the solution space position of the optimal individual as the center in the solution space. Within this range, the number of search agent individuals will be identified and counted.
[0124] Step 226, determine whether the number of search agent individuals within the preset range exceeds the preset number. If it exceeds, execute step 227; if it does not exceed, execute step 228.
[0125] If the number of search agent individuals within the preset range exceeds the preset number, it indicates that the distribution of the search agent individuals is relatively concentrated. If the number of search agent individuals within the preset range does not exceed the preset number, it indicates that the distribution of the search agent individuals is relatively dispersed.
[0126] Step 227, the search agent individuals within the preset range move towards the solution space position of the optimal individual of the previous local search and move away from the solution space position of the worst individual of the previous local search; the search agent individuals outside the preset range move away from the solution space position of the optimal individual of the previous local search and move away from the solution space position of the worst individual of the previous local search.
[0127] In the case where the distribution of search agent individuals is relatively concentrated, in order to avoid the search agent individuals clustering together and affecting the search effect, the search agent individuals within the preset range can move towards the solution space position of the optimal individual while moving away from the solution space position of the worst individual, so as to explore the solution space more efficiently and reduce ineffective searches; the search agent individuals outside the preset range can move away from the solution space position of the optimal individual and move away from the solution space position of the worst individual, so as to explore new solution space regions and increase the diversity of searches.
[0128] Step 228, each search agent individual adjusts the numerical value of its feature vector in the direction of its own vector and the preset first adjustment step size, so as to move towards the solution space position of the optimal individual in the previous round of local search and move away from the solution space position of the worst individual in the previous round of local search.
[0129] By approaching the optimal individual, the search agent individual can explore the solution space more efficiently and reduce ineffective searches. By moving away from the position of the worst individual, the search agent individual can avoid falling into a local optimal solution, thereby improving the global search ability.
[0130] Step 230, during the movement process, calculate the clustering fitness of the individuals whose solution space positions match.
[0131] Specifically, the feature vectors of each video frame of the individuals whose solution space positions match can be used as the clustering centers respectively, and all the video frames of the surveillance video are clustered to obtain a set number of classes; calculate the mean within-class distance of the set number of classes, and / or calculate the between-class distance of the set number of classes to obtain the clustering fitness of the corresponding individuals. The higher the clustering fitness, the better the clustering effect of the surveillance video when the feature vectors of each video frame of the individual are used as the clustering centers respectively.
[0132] Step 240, determine whether the clustering fitness of multiple consecutive individuals that match decreases one by one. If so, execute Step 250. If not, continue to move for local search and return to execute Step 230.
[0133] If the clustering fitness of multiple consecutive individuals that match decreases one by one, it means that the current search direction is not good and the search direction needs to be re-determined, that is, the vector adjustment direction is re-determined.
[0134] Step 250, each search agent individual re-determines the vector adjustment direction of each search agent individual according to the solution space position of the individual with the highest clustering fitness among the individuals that match. Each search agent individual adjusts the numerical value of its feature vector in the direction of the vector it re-determines and the preset second adjustment step size, so as to move towards the solution space position of the individual with the highest clustering fitness among the individuals that each search agent individual matches.
[0135] The re-determined vector adjustment direction is to move towards the solution space position of the individual with the highest clustering fitness among the individuals matched by the corresponding search agent individual, that is, to perform a backward search. The second adjustment step size is smaller than the first adjustment step size, indicating that during the backward search, a more detailed search is carried out to avoid missing the optimal individual.
[0136] If it is found that the current search direction leads to a continuous decrease in the clustering fitness, the search agent individual will adjust the direction, perform a backward search, and search more carefully. By dynamically adjusting the search direction and step size, a balance can be achieved between exploration and exploitation, improving the search effect.
[0137] Step 260: Determine whether the local search stop condition is satisfied. If it is satisfied, execute step 270; if not, continue to move for local search and return to execute step 230.
[0138] The local search stop conditions include that the total number of moving steps reaches the set number of steps, and / or the total moving step size reaches the set step size.
[0139] Step 270: Determine the optimal individual and the worst individual for the local search in the current round according to the clustering fitness of the matched individuals.
[0140] The individual with the highest clustering fitness among the individuals matched by each search agent individual in the current round can be obtained to get the first individual set, and the individual with the highest clustering fitness in the first individual set is determined as the optimal individual for the local search in the current round. The individual with the lowest clustering fitness among the individuals matched by each search agent individual in the current round can be obtained to get the second individual set, and the individual with the lowest clustering fitness in the second individual set is determined as the worst individual for the local search in the current round. Thus, one round of global search is completed.
[0141] Step 280: Determine whether the global search stop condition is satisfied. If it is satisfied, execute step 290; if not, return to step 223 to continue the next round of search.
[0142] The global search stop conditions include that the cumulative number of search rounds reaches the set number of rounds, and / or the change range of the clustering fitness of the optimal individuals found in multiple consecutive rounds remains within the set range.
[0143] Step 290: Use the finally determined global optimal individual as the initial clustering center of the surveillance video.
[0144] The global optimal individuals determined in each round can be obtained to get the third individual set, and the individual with the highest clustering fitness in the third individual set is determined as the finally determined global optimal individual.
[0145] After obtaining the finally determined global optimal individual, the feature vector of the video frame corresponding to the finally determined global optimal individual can be used as the initial clustering center of the surveillance video. The subsequent processing can refer to the description of the previous embodiments and will not be elaborated here.
[0146] Figure 3 FIG. is an example diagram of the surveillance video processing method provided by the embodiments of the present invention. After clustering video frames using a set of clustering centers found by the spider wasp algorithm improved by the bidirectional search strategy as the initial clustering centers of K-means clustering, target clustering centers can be obtained. After clustering the video frames using the target clustering centers, the abnormal behaviors in the surveillance video can be identified according to the type labels of the template vectors matched by the feature vectors of each cluster. The risk points in the surveillance video can be identified according to the positions of the abnormal behaviors in the video, and corresponding optimization measures, such as alarm prompts, video clips, etc., can be taken for the risk points to improve the security and practicality of the surveillance system.
[0147] In this embodiment, a population is constructed using the spider wasp algorithm. Each individual in the population consists of a set of video frames, and this set of video frames forms the solution space position of the corresponding individual. The solution space positions of all individuals in the population constitute the solution space of the population. The spider wasp algorithm has strong global search ability. Searching for the global optimal individual of the population using the spider wasp algorithm and using the global optimal individual as the initial clustering center of the surveillance video can effectively avoid the sensitivity problem of the traditional K-means algorithm to the selection of the initial clustering center and improve the clustering accuracy; multiple search agent individuals are set in the population. Each search agent individual starts from the current solution space position, moves by adjusting the numerical values of the feature vectors, and calculates the clustering fitness of the individual matching the solution space position, providing interpretability for the execution of the algorithm; in each round of search, each search agent individual shares the search results, and each search agent individual dynamically adjusts its own movement path according to the solution space positions of the optimal individual and the worst individual, that is, dynamically adjusts the search path by combining the bidirectional search strategy, which can better adapt to the changes in the surveillance video data. The cooperation of multiple search agent individuals improves the accuracy and speed of the search; when the global search stop condition is met, the finally determined global optimal individual is used as the initial clustering center of the surveillance video. By determining the global optimal individual, a high-quality initial clustering center is provided for the subsequent clustering analysis. Clustering all video frames using the initial clustering center can finally identify the risk points in the surveillance video more quickly and accurately, improving the security management efficiency.
[0148] Figure 4 FIG. is a structural schematic diagram of the surveillance video processing device provided by the embodiments of the present invention. This device is suitable for executing the surveillance video processing method provided by the embodiments of the present invention, such as Figure 4 shown, this device may specifically include:
[0149] The frame division module 401 is used to perform frame division on the surveillance video to obtain multiple video frames;
[0150] The vectorization module 402 is used to obtain the feature information of each video frame and vectorize the feature information of each video frame to obtain the feature vector of each video frame;
[0151] The generation module 403 is used to form a population of the spider wasp algorithm according to the feature vectors of each video frame; wherein, the population includes multiple individuals, each individual includes a set number of video frames, and the set of feature vectors of all video frames of each individual is used as the solution space position of the individual, and the solution space positions of all individuals in the population constitute the solution space of the population;
[0152] The search module 404 is used to start the individual search of the current round by using the spider wasp algorithm; wherein, the individual search process of the current round includes:
[0153] Set multiple search agent individuals; each search agent individual includes a set number of search video frames, and the set of feature vectors of all search video frames of each search agent individual is used as the solution space position of the search agent individual;
[0154] Each search agent individual starts from its current solution space position and moves the solution space position according to the set path algorithm by adjusting the feature vector value, and calculates the clustering fitness of the individuals matching the solution space position during the movement; the clustering fitness is the clustering effect value of clustering the surveillance video when the feature vectors of each video frame of the individual are used as the clustering centers respectively;
[0155] When the local search stop condition is met, determine the optimal individual and the worst individual of the local search performed in the current round according to the clustering fitness of the matched individuals; wherein, the solution space positions of the optimal individual and the worst individual are used to adjust the movement path of the search agent individual in the next round of search process by applying the set path algorithm;
[0156] Return to start the individual search of the next round, and when the global search stop condition is reached, use the finally determined global optimal individual as the initial clustering center of the surveillance video.
[0157] In one embodiment, the search module 404 sets multiple search agent individuals, including:
[0158] If the current round is the first round, within the solution space range of the population, randomly create multiple sets of feature vector sets of search video frames, a set of search video frames constitutes a search agent individual, a set of search video frames includes a set number of search video frames, and use the randomly created multiple sets of feature vector sets of search video frames as the current solution space position of the corresponding search agent individual;
[0159] If the current round is not the first round, then use the multiple search agent individuals in the previous round as the multiple search agent individuals in the current round, and use the solution space positions of the multiple search agent individuals at the end of the previous round of search as the current solution space positions of the corresponding search agent individuals.
[0160] In one embodiment, the movement of the solution space position is performed according to the set path algorithm by adjusting the numerical values of the feature vectors, including:
[0161] If the current round is the first round, then each search agent individual uses the individual with a higher clustering fitness among the two individuals randomly selected by itself as the reference individual. Each search agent individual determines the vector adjustment direction of each search agent individual according to the solution space position of its respective reference individual, and each search agent individual adjusts the numerical value of its own feature vector according to its own vector adjustment direction and a preset first adjustment step length to move towards the solution space position of its respective reference individual.
[0162] If the current round is not the first round, then each search agent individual determines the vector adjustment direction of each search agent individual according to the solution space positions of the optimal individual and the worst individual in the previous round of local search. Each search agent individual adjusts the numerical value of its own feature vector according to its own vector adjustment direction and a preset first adjustment step length to move towards the solution space position of the optimal individual in the previous round of local search and move away from the solution space position of the worst individual in the previous round of local search.
[0163] In one embodiment, before moving towards the solution space position of the optimal individual in the previous round of local search and moving away from the solution space position of the worst individual in the previous round of local search, it further includes:
[0164] Identify the number of search agent individuals within the preset range of the solution space position of the optimal individual in the previous round of local search;
[0165] If the number of search agent individuals within the preset range does not exceed the preset number, then trigger the execution of moving towards the solution space position of the optimal individual in the previous round of local search and moving away from the solution space position of the worst individual in the previous round of local search.
[0166] In one embodiment, after identifying the number of search agent individuals within the preset range of the solution space position of the optimal individual in the previous round of local search, it further includes:
[0167] If the number of search agent individuals within the preset range exceeds the preset number, then the search agent individuals within the preset range move towards the solution space position of the optimal individual in the previous round of local search and move away from the solution space position of the worst individual in the previous round of local search; the search agent individuals outside the preset range move away from the solution space position of the optimal individual in the previous round of local search and move away from the solution space position of the worst individual in the previous round of local search.
[0168] In one embodiment, after moving the position in the solution space according to the set path algorithm by adjusting the numerical value of the feature vector and calculating the clustering fitness of the individuals matching the position in the solution space during the movement, the following steps are further included:
[0169] If the clustering fitness of each of the continuously multiple individuals matched by each search agent individual decreases one by one, each search agent individual re-determines the vector adjustment direction of each search agent individual according to the solution space position of the individual with the highest clustering fitness among the matched individuals, and each search agent individual adjusts the numerical value of its own feature vector according to the re-determined vector adjustment direction of itself and the preset second adjustment step length, so as to move towards the solution space position of the individual with the highest clustering fitness among the individuals matched by each search agent individual;
[0170] Wherein, the second adjustment step length is less than the first adjustment step length.
[0171] In one embodiment, calculating the clustering fitness of the individuals matching the position in the solution space includes:
[0172] Taking the feature vectors of each video frame of the individuals matching the position in the solution space as the clustering centers respectively, clustering all the video frames of the surveillance video to obtain a set number of classes;
[0173] Calculating the mean intra-class distance of the set number of classes, and / or calculating the inter-class distance of the set number of classes to obtain the clustering fitness of the corresponding individuals.
[0174] In one embodiment, determining the optimal individual and the worst individual for the local search in the current round according to the clustering fitness of the matched individuals includes:
[0175] Obtaining the individuals with the highest clustering fitness matched by each search agent individual in the current round to obtain a first individual set, and determining the individual with the highest clustering fitness in the first individual set as the optimal individual for the local search in the current round;
[0176] Obtaining the individuals with the lowest clustering fitness matched by each search agent individual in the current round to obtain a second individual set, and determining the individual with the lowest clustering fitness in the second individual set as the worst individual for the local search in the current round.
[0177] In one embodiment, the local search stop conditions include that the total number of movement steps reaches the set number of steps, and / or the total movement step length reaches the set step length.
[0178] In one embodiment, the global search stop conditions include that the cumulative number of search rounds reaches the set number of rounds, and / or the change range of the clustering fitness of the optimal individuals searched in multiple consecutive rounds remains within the set range.
[0179] In one embodiment, the following steps are further included:
[0180] A clustering module, which is used to cluster all video frames of a monitored video by using initial cluster centers to obtain target cluster centers;
[0181] An identification module, which is used to cluster all video frames according to the target cluster centers, and match the feature vectors of each cluster with the template vectors of abnormal behaviors to identify abnormal behaviors in the monitored video.
[0182] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional module is used as an example for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described functional modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0183] The device according to the embodiment of the present invention divides the monitored video into frames, extracts the feature information of each video frame and vectorizes it. Through feature vectorization, the video frames are converted into quantifiable data, which is convenient for subsequent clustering analysis. This method can process massive video data more efficiently and reduce the computational complexity; a population is constructed by using the spider wasp algorithm. Each individual in the population consists of a group of video frames. This group of video frames forms the solution space position of the corresponding individual. The solution space positions of all individuals in the population constitute the solution space of the population. The spider wasp algorithm has strong global search ability. Using the spider wasp algorithm to search for the globally optimal individual in the population and taking the final globally optimal individual as the initial cluster center of the monitored video can effectively avoid the sensitivity problem of the traditional K-means algorithm to the selection of the initial cluster center and improve the clustering accuracy; multiple search agent individuals are set in the population. Each search agent individual starts from the current solution space position, moves by adjusting the numerical values of the feature vectors, and calculates the clustering fitness of the individuals matching the solution space position, providing interpretability for the execution of the algorithm; in each round of search, each search agent individual shares the search results, and each search agent individual dynamically adjusts its own movement path according to the solution space positions of the optimal individual and the worst individual, that is, dynamically adjusts the search path by combining the bidirectional search strategy, which can better adapt to the changes in the monitored video data. The multiple search agent individuals cooperate to improve the accuracy and speed of the search; when the global search stop condition is met, the finally determined globally optimal individual is used as the initial cluster center of the monitored video. Through the determination of the globally optimal individual, a high-quality initial cluster center is provided for subsequent clustering analysis. Using the initial cluster center to cluster all video frames can finally identify the risk points in the monitored video more quickly and accurately, and improve the safety management efficiency.
[0184] Next, refer to Figure 5, which shows a schematic structural diagram of a computer system 500 of an electronic device suitable for implementing the embodiments of the present invention. Figure 5 The illustrated electronic device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0185] As Figure 5 shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 502 or the programs loaded from the storage section 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the computer system 500 are also stored. The CPU 501, ROM 502, and RAM 503 are connected to each other via a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0186] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as required. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as required, so that the computer program read from it can be installed into the storage section 508 as required.
[0187] Particularly, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above functions defined in the system of the present invention are executed.
[0188] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0189] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0190] The modules and / or units involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules and / or units can also be provided in a processor. For example, it can be described as: a processor includes a frame division module, a vectorization module, a generation module, and a search module. Among them, the names of these modules do not constitute a limitation to the module itself in some cases.
[0191] As another aspect, the present invention further provides a computer-readable medium, which can be included in the device described in the above embodiments; or it can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes:
[0192] Performing frame division processing on the monitored video to obtain a plurality of video frames; acquiring the feature information of each video frame, and vectorizing the feature information of each video frame to obtain the feature vector of each video frame; forming a population of the spider wasp algorithm according to the feature vectors of each video frame; wherein, the population includes a plurality of individuals, each individual includes a set number of video frames, and the set of feature vectors of all video frames of each individual is used as the solution space position of the individual, and the solution space positions of all individuals in the population constitute the solution space of the population; adopting the spider wasp algorithm to start the individual search of the current round; wherein, the individual search process of the current round includes: setting a plurality of search agent individuals; each search agent individual includes a set number of search video frames, and the set of feature vectors of all search video frames of each search agent individual is used as the solution space position of the search agent individual; each search agent individual starts from its own current solution space position, moves the solution space position according to the set path algorithm by adjusting the feature vector value, and calculates the clustering fitness of the individuals matching the solution space position during the movement; the clustering fitness is the clustering effect value when clustering the monitored video with the feature vectors of each video frame of the individual as the clustering centers respectively; when the local search stop condition is met, determining the optimal individual and the worst individual of the local search performed in the current round according to the clustering fitness of the matched individuals; wherein, the solution space positions of the optimal individual and the worst individual are used to adjust the movement path of the search agent individual in the set path algorithm in the next round of search process; returning to start the individual search of the next round, and when the global search stop condition is reached, taking the finally determined global optimal individual as the initial clustering center of the monitored video.
[0193] According to the technical solution of the embodiment of the present invention, the surveillance video is framed, the feature information of each video frame is extracted and vectorized. Through feature vectorization, the video frame is converted into quantifiable data, which is convenient for subsequent clustering analysis. This method can process massive video data more efficiently and reduce the computational complexity. The spider wasp algorithm is used to construct a population. Each individual in the population consists of a set of video frames, and this set of video frames forms the solution space position of the corresponding individual. The solution space positions of all individuals in the population constitute the solution space of the population. The spider wasp algorithm has strong global search ability. Using the spider wasp algorithm to search for the global optimal individual in the population and taking the final global optimal individual as the initial clustering center of the surveillance video can effectively avoid the sensitivity problem of the traditional K-means algorithm to the selection of the initial clustering center and improve the clustering accuracy. Multiple search agent individuals are set in the population. Each search agent individual starts from the current solution space position, moves by adjusting the numerical value of the feature vector, and calculates the clustering fitness of the individual that matches the solution space position, providing interpretability for the execution of the algorithm. In each round of search, each search agent individual shares the search results, and each search agent individual dynamically adjusts its own movement path according to the solution space positions of the optimal individual and the worst individual, that is, dynamically adjusts the search path by combining the bidirectional search strategy, which can better adapt to the changes in the surveillance video data. The multiple search agent individuals cooperate to improve the accuracy and speed of the search. When the global search stop condition is met, the finally determined global optimal individual is used as the initial clustering center of the surveillance video. By determining the global optimal individual, a high-quality initial clustering center is provided for the subsequent clustering analysis. Using the initial clustering center to cluster all video frames can finally identify the risk points in the surveillance video more quickly and accurately, and improve the safety management efficiency.
[0194] The embodiment of the present invention also provides a computer program product, including a computer program, which when executed by a processor, implements the surveillance video processing method provided in any embodiment of the present application.
[0195] In the process of implementing the computer program product, the computer program code for executing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages. The program code can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0196] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0197] It should be noted that in the technical solution of the present disclosure, in aspects such as the collection, gathering, updating, analysis, processing, use, transmission, storage, etc. of the user's personal information, all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken for the user's personal information to prevent illegal access to the user's personal information data, and to safeguard the security of the user's personal information, network security and national security.
[0198] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for processing surveillance videos, characterized in that, Including: Performing frame splitting on the monitored video to obtain multiple video frames; Obtaining the feature information of each video frame, and vectorizing the feature information of each video frame to obtain the feature vector of each video frame; Forming a population of the spider wasp algorithm according to the feature vectors of each video frame; wherein, the population includes multiple individuals, each individual includes a set number of video frames, and the set of feature vectors of all video frames of each individual is used as the solution space position of the individual, and the solution space positions of all individuals in the population constitute the solution space of the population; Adopting the spider wasp algorithm to start the individual search of the current round; wherein, the individual search process of the current round includes: Setting multiple search agent individuals; each search agent individual includes a set number of search video frames, and the set of feature vectors of all search video frames of each search agent individual is used as the solution space position of the search agent individual; Each search agent individual starts from its own current solution space position, moves the solution space position according to the set path algorithm by adjusting the feature vector value, and calculates the clustering fitness of the individuals matching the solution space position during the movement; the clustering fitness is the clustering effect value of clustering the monitored video when the feature vectors of each video frame of the individual are used as the clustering centers respectively; When the local search stop condition is satisfied, determining the optimal individual and the worst individual of the local search performed in the current round according to the clustering fitness of the matched individuals; wherein, the solution space positions of the optimal individual and the worst individual are used to apply to the set path algorithm to adjust the movement path of the search agent individual in the next round of search; Returning to start the individual search of the next round, until the global search stop condition is reached, and taking the finally determined global optimal individual as the initial clustering center of the monitored video.
2. The method according to claim 1, characterized in that, Setting multiple search agent individuals, including: If the current round is the first round, within the solution space range of the population, randomly creating multiple sets of feature vector sets of search video frames, a set of search video frames constitutes a search agent individual, a set of search video frames includes a set number of search video frames, and taking the randomly created multiple sets of feature vector sets of search video frames as the current solution space positions of the corresponding search agent individuals; If the current round is not the first round, taking the multiple search agent individuals of the previous round as the multiple search agent individuals of the current round, and taking the solution space positions of the multiple search agent individuals at the end of the previous round of search as the current solution space positions of the corresponding search agent individuals.
3. The method according to claim 1, wherein Moving the solution space position according to the set path algorithm by adjusting the feature vector value, including: If the current round is the first round, each search agent individual takes the individual with the higher clustering fitness among the two individuals randomly selected by itself as the reference individual, each search agent individual determines the vector adjustment direction of each search agent individual according to the solution space position of its respective reference individual, and each search agent individual adjusts the feature vector value of itself according to its own vector adjustment direction and a preset first adjustment step length to move towards the solution space position of its respective reference individual; If the current round is not the first round, each search agent individual determines the vector adjustment direction of each search agent individual according to the solution space positions of the optimal individual and the worst individual in the previous round of local search. Each search agent individual adjusts the numerical value of its own feature vector according to its own vector adjustment direction and a preset first adjustment step length, so as to move towards the solution space position of the optimal individual in the previous round of local search and move away from the solution space position of the worst individual in the previous round of local search.
4. The method according to claim 3, characterized in that, Before moving towards the solution space position of the optimal individual in the previous round of local search and moving away from the solution space position of the worst individual in the previous round of local search, it further includes: Identifying the number of search agent individuals within a preset range of the solution space position of the optimal individual in the previous round of local search; If the number of search agent individuals within the preset range does not exceed the preset number, trigger the execution of moving towards the solution space position of the optimal individual in the previous round of local search and moving away from the solution space position of the worst individual in the previous round of local search.
5. The method according to claim 4, characterized in that After identifying the number of search agent individuals within a preset range of the solution space position of the optimal individual in the previous round of local search, it further includes: If the number of search agent individuals within the preset range exceeds the preset number, the search agent individuals within the preset range move towards the solution space position of the optimal individual in the previous round of local search and move away from the solution space position of the worst individual in the previous round of local search; the search agent individuals outside the preset range move away from the solution space position of the optimal individual in the previous round of local search and move away from the solution space position of the worst individual in the previous round of local search.
6. The method according to claim 3, wherein After moving the solution space position according to the set path algorithm by adjusting the numerical value of the feature vector and calculating the clustering fitness of the individuals matching the solution space position during the movement, it further includes: Each search agent individual judges that if the clustering fitness of a continuous plurality of individuals it matches decreases one by one, then each search agent individual re-determines the vector adjustment direction of each search agent individual according to the solution space position of the individual with the highest clustering fitness among the individuals it matches. Each search agent individual adjusts the numerical value of its own feature vector according to its own re-determined vector adjustment direction and a preset second adjustment step length, so as to move towards the solution space position of the individual with the highest clustering fitness among the individuals each search agent individual matches; Wherein, the second adjustment step length is less than the first adjustment step length.
7. The method according to claim 1, wherein Calculating the clustering fitness of the individuals matching the solution space position includes: Taking the feature vectors of each video frame of the individuals matching the solution space position as clustering centers respectively, clustering all the video frames of the surveillance video to obtain a set number of classes; Calculating the mean intra-class distance of the set number of classes and / or calculating the inter-class distance of the set number of classes to obtain the clustering fitness of the corresponding individuals.
8. The method according to claim 1, wherein Determining the optimal individual and the worst individual in the local search performed in the current round according to the clustering fitness of the individuals that match, includes: Obtaining the individuals with the highest clustering fitness that each search agent individual matches in the current round to obtain a first individual set, and determining the individual with the highest clustering fitness in the first individual set as the optimal individual in the local search performed in the current round; Obtain the individual with the lowest clustering fitness among the individuals matched by each search agent in the current round to obtain a second set of individuals, and determine the individual with the lowest clustering fitness in the second set of individuals as the worst individual for the local search in the current round.
9. The method according to claim 1, wherein The local search stop conditions include that the total number of moving steps reaches the set number of steps, and / or the total moving step length reaches the set step length.
10. The method according to claim 1, characterized in that The global search stop conditions include that the cumulative number of search rounds reaches the set number of rounds, and / or the change range of the clustering fitness of the optimal individuals found in multiple consecutive rounds remains within the set range.
11. The method according to claim 1, wherein It also includes: Use the initial clustering centers to cluster all video frames of the surveillance video to obtain target clustering centers; Cluster all video frames according to the target clustering centers, and match the feature vectors of each cluster with the template vectors of abnormal behaviors to identify abnormal behaviors in the surveillance video.
12. A monitoring video processing device, characterized in that, It includes: A frame splitting module for splitting the surveillance video into frames to obtain multiple video frames; A vectorization module for obtaining the feature information of each video frame and vectorizing the feature information of each video frame to obtain the feature vector of each video frame; A generation module for forming a population of the spider wasp algorithm according to the feature vectors of each video frame; wherein, the population includes multiple individuals, each individual includes a set number of video frames, the set of feature vectors of all video frames of each individual serves as the solution space position of the individual, and the solution space positions of all individuals in the population constitute the solution space of the population; A search module for starting the individual search in the current round using the spider wasp algorithm; wherein, the individual search process in the current round includes: Set multiple search agent individuals; each search agent individual includes a set number of search video frames, and the set of feature vectors of all search video frames of each search agent individual serves as the solution space position of the search agent individual; Each search agent individual starts from its current solution space position and moves the solution space position according to the set path algorithm by adjusting the numerical values of the feature vectors, and calculates the clustering fitness of the individuals matching the solution space position during the movement; the clustering fitness is the clustering effect value when the feature vectors of each video frame of the individual are used as the clustering centers to cluster the surveillance video; When the local search stop conditions are met, determine the optimal individual and the worst individual for the local search in the current round according to the clustering fitness of the matched individuals; wherein, the solution space positions of the optimal individual and the worst individual are used to adjust the movement paths of the search agent individuals in the next round of search process using the set path algorithm; Return to start the individual search in the next round, and until the global search stop conditions are reached, use the finally determined global optimal individual as the initial clustering center of the surveillance video.
13. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the surveillance video processing method according to any one of claims 1 to 11.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the surveillance video processing method according to any one of claims 1 to 11.
15. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by the processor, it implements the surveillance video processing method according to any one of claims 1 to 11.