Intelligent Analysis Method and System for Video Streams Based on Neural Networks

By applying intelligent analysis methods based on neural networks in video stream analysis, the abnormal images are automatically identified, and the problems of large resource consumption and inaccurate judgment in the prior art are solved, thereby achieving more efficient and accurate video stream analysis.

CN119723426BActive Publication Date: 2025-06-20SHENZHEN GEEK INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510211513.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-20
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The prior art relies on manual frame-by-frame analysis in video stream analysis, resulting in large resource consumption and the judgment of abnormal images relies on personal experience, and there are inaccurate problems.

Method used

Using an intelligent analysis method based on neural network, the video decomposition unit, image recognition unit, image comparison unit and result feedback unit are used to automatically analyze the video stream, identify and confirm abnormal images.

Benefits of technology

It improves the accuracy and intelligence of video stream analysis and reduces the resource consumption required during the analysis process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723426B_ABST
    Figure CN119723426B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology. An intelligent analysis method and system for video streams based on a neural network includes: confirming receipt of a video decomposition instruction from a video decomposition unit, parsing the video decomposition instruction to obtain a first decomposition frame number, using the first decomposition frame number and an initial video stream to obtain a decomposed image time sequence, obtaining a set of decomposed video time points based on the decomposed image time sequence, obtaining a set of target video streams based on the set of decomposed video time points and the initial video stream, obtaining a second decomposition frame number based on the target video stream, identifying a set of abnormal images based on the second decomposition frame number and the target video stream, and using a result feedback unit to send the set of abnormal images to the initiator of the analysis instruction, thereby realizing intelligent analysis of the initial video stream. The present invention can improve the accuracy and intelligence level of video stream analysis and reduce the resources required for video stream analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to an intelligent analysis method and system for video streams based on a neural network. Background Art

[0002] With the popularization of Internet of Things technology, more and more monitoring is applied to daily life. Through monitoring, real-time shooting of different regions can be achieved. If an emergency occurs in the region, the captured video will be used as important video evidence. Correspondingly, how to intelligently and resource-savingly analyze the video stream is of great significance for confirming the accuracy of the initial video stream.

[0003] Currently, for the analysis of video streams, most are manually analyzed frame by frame to screen out possible abnormal images from the video stream.

[0004] Although the above method can achieve the analysis of video streams, when performing frame-by-frame analysis on the video stream, it requires a large amount of human resources, and the judgment of abnormal images in the video mostly depends on personal experience, which may lead to inaccurate screening of abnormal images. Therefore, how to intelligently, accurately and energy-savingly analyze the video stream has become an urgent problem to be solved. Summary of the Invention

[0005] The present invention provides an intelligent analysis method for video streams based on a neural network and a computer-readable storage medium, and its main purpose is to improve the accuracy and intelligence of analyzing video streams and reduce the resources required for analyzing video streams.

[0006] To achieve the above object, an intelligent analysis method for video streams based on a neural network provided by the present invention includes:

[0007] Receiving an analysis instruction, and confirming an intelligent analysis environment based on the analysis instruction. The intelligent analysis environment includes an initial video stream and a video analysis system, and the video analysis system includes: a video decomposition unit, an image recognition unit, an image comparison unit, and a result feedback unit;

[0008] Confirming to receive a video decomposition instruction from the video decomposition unit, parsing the video decomposition instruction to obtain a first decomposition frame number, and using the first decomposition frame number and the initial video stream to obtain a decomposition image time series, where the decomposition image time series includes multiple decomposition image nodes, and the decomposition image nodes include decomposition images and image times;

[0009] Obtain a set of decomposition video time points based on the decomposed image time series, where the set of decomposition video time points includes M decomposition video time points, and M is an integer greater than or equal to 0. Obtain a set of target video streams based on the set of decomposition video time points and the initial video stream, where the set of target video streams includes one or more target video streams. Perform the following operations on each target video stream in the set of target video streams:

[0010] Obtain a second decomposition frame number based on the target video stream. Based on the second decomposition frame number and the target video stream, identify a set of abnormal images, where the set of abnormal images includes N abnormal images, and N is an integer greater than or equal to 0. Use the result feedback unit to send the set of abnormal images to the initiator of the analysis instruction to achieve intelligent analysis of the initial video stream.

[0011] Optionally, the obtaining of the set of decomposition video time points based on the decomposed image time series includes:

[0012] Obtain an analysis grayscale image time series based on the decomposed image time series. Perform the following operations on each analysis grayscale image in the analysis grayscale image time series:

[0013] Obtain an analysis grayscale mean value based on the analysis grayscale image, where the analysis grayscale mean value is the mean of multiple grayscale values corresponding to the analysis grayscale image;

[0014] Associate the analysis grayscale mean value and the image time to obtain an initial analysis node. Aggregate the initial analysis nodes to obtain a set of initial analysis nodes;

[0015] Use a pre-constructed clustering model and the set of initial analysis nodes to obtain one or more sets of target analysis nodes;

[0016] Based on one or more sets of target analysis nodes, identify the set of decomposition video time points in the initial video stream.

[0017] Optionally, the obtaining of the second decomposition frame number based on the target video stream includes:

[0018] Obtain the initial frame number and video duration of the target video stream. Based on the initial frame number and video duration, obtain the number of images. Obtain an image extraction gradient set for extracting images. Based on the number of images and the image extraction gradient set, extract an image time series from the target video stream, where the image extraction gradient set includes multiple image extraction ratios, and the image time series includes multiple initial images;

[0019] Confirm receiving an image recognition instruction from the image recognition unit. Based on the image recognition instruction, identify a set of image recognition models, where the set of image recognition models includes multiple image recognition models. Sequentially extract the initial images from the image time series, and perform the following operations on the extracted initial images:

[0020] Randomly extract two image recognition models from the set of image recognition models to obtain a first recognition model and a second recognition model. Based on the first recognition model and the second recognition model, identify the extracted initial images to obtain a first recognition name set and a second recognition name set. Among them, the first recognition name set includes multiple first recognition nodes, and each first recognition node includes a first recognition name and a first recognition quantity. The second recognition name set includes multiple second recognition nodes, and each second recognition node includes a second recognition name and a second recognition quantity;

[0021] After confirming the target recognition name set by using the first recognition name set and the second recognition name set, where the target recognition name set includes multiple target recognition nodes, and each target recognition node includes a target recognition name and a target recognition quantity, perform the following operations on each target recognition name in the target recognition name set:

[0022] Identify the local region image corresponding to the target recognition name in the initial image, and use the target recognition name and the target recognition quantity to label the local region image to obtain a target image with an identification serial number. Summarize the target images to obtain a target image set, and use the target image set to obtain the second decomposition frame number.

[0023] Optionally, the confirmation of the target recognition name set by using the first recognition name set and the second recognition name set includes:

[0024] Perform the following operations on each first recognition node in the first recognition name set:

[0025] Judge whether there is a second recognition node in the second recognition name set that is the same as the first recognition node;

[0026] If there is no second recognition node in the second recognition name set that is the same as the first recognition node, randomly extract multiple initial image recognition models from the set of image recognition models. The number of multiple initial image recognition models is a preset extraction quantity. Use the multiple initial image recognition models and the initial images to obtain multiple initial recognition name sets. Among them, the initial image recognition models and the initial recognition name sets are in one-to-one correspondence, and each initial recognition name set includes multiple initial recognition nodes;

[0027] In a combined form, combine the multiple initial recognition name sets to obtain multiple combined recognition name sets. Among them, each combined recognition name set includes two initial recognition name sets. Identify one or more fusion recognition name sets from the multiple combined recognition name sets. Among the fusion recognition name sets, the two initial recognition name sets are the same;

[0028] Eliminate any one of the initial recognition name sets corresponding to each of the one or more fused recognition name sets from multiple initial recognition name sets to obtain an updated recognition name set. Use the updated recognition name set as the multiple initial recognition name sets, and return the step of combining the multiple initial recognition name sets in a combined form until one or more initial fused name sets are obtained, where the multiple initial recognition name sets corresponding to the initial fused name sets are all the same;

[0029] Statistically count the number of initial recognition name sets corresponding to each of the one or more initial fused name sets in the one or more initial fused name sets to obtain one or more initial fusion quantities, where the initial fusion quantities correspond one-to-one to the initial fused name sets;

[0030] Extract the largest initial fusion quantity from the one or more initial fusion quantities to obtain a target verification quantity, calculate the ratio of the target verification quantity to the extraction quantity to obtain a correct ratio. When the correct ratio is greater than or equal to a preset ratio threshold, use the fused recognition name set corresponding to the target verification quantity as the target recognition name set;

[0031] Otherwise, return the step of randomly extracting multiple initial image recognition models from the image recognition model set until a target recognition name set is obtained.

[0032] Optionally, the obtaining the second decomposition frame number by using the target image set includes:

[0033] Summarize the target image set to obtain multiple target image sets, and use the target recognition name, serial number, and target image to obtain the target image time sequence;

[0034] Sequentially extract initial analysis images from the target image time sequence, and perform the following operations on the extracted initial analysis images:

[0035] Based on the initial analysis image, confirm a target analysis image in the target image time sequence, where the target analysis image is adjacent to and lagging behind the initial analysis image in the target image time sequence;

[0036] Obtain an initial coordinate point set and a target coordinate point set based on the initial analysis image and the target analysis image;

[0037] Map the initial coordinate points in the initial coordinate point set and the target coordinate points in the target coordinate point set to a pre-constructed reference coordinate system respectively to obtain a mapped coordinate point set, where the mapped coordinate point set includes multiple mapped coordinate points;

[0038] Count the number of reference coordinate points corresponding to each mapped coordinate point in the set of mapped coordinate points to obtain a reference quantity set, where the reference coordinate points are initial coordinate points or target coordinate points, the reference quantity set includes multiple reference quantities, count the number of reference quantities in the reference quantity set that are 2 to obtain a target reference quantity, count the number of mapped coordinate points in the set of mapped coordinate points to obtain a comprehensive quantity, calculate the ratio of the target reference quantity to the comprehensive quantity to obtain a reference quantity ratio. If the reference quantity ratio is less than or equal to a preset reference ratio threshold, then use the target analysis image as the initial analysis image, and return the step of identifying the target analysis image in the target image time series based on the initial analysis image. If the reference quantity ratios of the reference quantity sets corresponding to the target image time series are all less than or equal to the reference ratio threshold, then label the target image as a background image; otherwise, label the target image as an initial detection image. After confirming that the initial detection image is a preset target detection image, use the target detection image to identify a set of adjacent images in the initial image, where the set of adjacent images includes one or more adjacent images, and the adjacent images are target detection images or background images, and obtain a second decomposition frame number based on the set of adjacent images and the detection image.

[0039] Optionally, the step of confirming that the initial detection image is a preset target detection image includes:

[0040] Obtain a detection image time series according to the target recognition name, serial number, and the initial detection image, where the detection image time series includes multiple initial detection images. Perform a grayscale operation on each initial detection image in the detection image time series to obtain a grayscale detection image time series. Based on a preset sliding window, sequentially extract a group of analysis images from the grayscale detection image time series, where the group of analysis images includes two grayscale detection images;

[0041] Perform the following operations on each grayscale detection image in the group of analysis images:

[0042] Calculate the image center coordinates based on the grayscale detection image, and the calculation formula is as follows:

[0043] ,

[0044] where respectively represent the abscissa and ordinate of the image center coordinates, represents that there are pixel points in the grayscale detection image, represents the th pixel point in the grayscale detection image, respectively represent the abscissa and ordinate of the th pixel point in the grayscale detection image;

[0045] Summarize the center coordinates of the images to obtain a set of image center coordinates, calculate the comprehensive image change degree based on the set of image center coordinates, and judge the comprehensive image change degree and a preset comprehensive change degree threshold;

[0046] If the comprehensive image change degree is greater than or equal to the comprehensive change degree threshold, confirm that the initial detection image is the target detection image.

[0047] Optionally, the calculating the comprehensive image change degree based on the set of image center coordinates includes:

[0048] Obtain the time series of image center coordinates based on the set of image center coordinates, and use the sliding window to sequentially extract analysis coordinate groups from the time series of image center coordinates. Each analysis coordinate group includes two image center coordinates. Obtain the Euclidean distance between the two image center coordinates in the analysis coordinate group to get the moving distance;

[0049] Extract the first image center coordinate and the last image center coordinate from the time series of image center coordinates to obtain the initial center coordinate and the target center coordinate, and obtain the evaluation distance based on the initial center coordinate and the target center coordinate;

[0050] Summarize the moving distances to obtain a set of moving distances, and use the set of moving distances to obtain the variance of the moving distances, where the variance of the moving distances is the variance of multiple moving distances in the set of moving distances;

[0051] Calculate the comprehensive image change degree based on the variance of the moving distances and the evaluation distance. The calculation formula is as follows:

[0052] ,

[0053] where, represents the comprehensive image change degree, are all preset coefficients, represents the evaluation distance, represents the variance of the moving distances.

[0054] Optionally, the obtaining the second decomposition frame number based on the adjacent image set and the detection image includes:

[0055] Obtain the adjacent grayscale image set and the detection grayscale image based on the adjacent image set and the detection image. Use the pre-constructed region growing algorithm and the pre-constructed image gray level gradient set to extract one or more local region images from the detection grayscale image, and identify one or more target region images in the one or more local region images. The image gray level gradient set includes multiple image gradient gray values, and the target region image is adjacent to at least one adjacent gray image in the adjacent gray image set;

[0056] Perform the following operations on each target region image in the one or more target region images:

[0057] Obtain the local region gray - level mean based on the target region image, where the local region gray - level mean is the mean of multiple gray - level values corresponding to the target region image, associate the target region image and the adjacent neighboring gray - level images of the target region image to obtain an associated image set;

[0058] Starting from the local region gray - level mean, use the region - growing algorithm to extract the associated region image from the associated image set, and identify the target discriminant image in the associated region image, where the target discriminant image is the region of the neighboring gray - level image adjacent to the target region image in the associated image set;

[0059] Obtain the target discriminant mean based on the target discriminant image, and calculate the discriminant frame number according to the target discriminant mean and the local region gray - level mean. The calculation formula is as follows:

[0060] ,

[0061] ,

[0062] Among them, represents the discriminant frame number, represents the initial frame number, represents the target discriminant mean, represents the local region gray - level mean, represents a preset coefficient, represents a preset coefficient, represents the rounding symbol;

[0063] Summarize the discriminant frame numbers to obtain a discriminant frame number set, and confirm the second decomposition frame number based on the discriminant frame number set, where the second decomposition frame number is the smallest discriminant frame number in the discriminant frame number set.

[0064] Optionally, the confirming the abnormal image set based on the second decomposition frame number and the target video stream includes:

[0065] Use the second decomposition frame number to extract the decomposed image time sequence from the target video stream. The decomposed image time sequence includes multiple initial decomposed images. Sequentially extract the initial decomposed images from the decomposed image time sequence, and perform the following operations on the extracted initial decomposed images:

[0066] Based on the initial decomposed image, confirm the target decomposed image in the decomposed image time sequence, where the target decomposed image is adjacent to the initial decomposed image and lags behind the extracted initial decomposed image;

[0067] Confirm the receipt of the image comparison instruction from the image comparison unit, parse the image comparison instruction to obtain the local name database, and use the image recognition model set and the local name database to respectively identify the target identification decomposition image set and the initial identification decomposition image set in the target decomposed image and the initial decomposed image, where the target identification decomposition image set includes multiple target identification decomposition images, and the initial identification decomposition image set includes multiple initial identification decomposition images;

[0068] Match the target identification decomposition images in the target identification decomposition image set and the initial identification decomposition images in the initial identification decomposition image set to obtain multiple identification decomposition nodes, where the identification decomposition nodes include the initial identification decomposition images and the target identification decomposition images;

[0069] Perform the following operations on each of the multiple identification decomposition nodes:

[0070] Obtain the local offset distance based on the identification decomposition node, summarize the local offset distances to obtain a set of local offset distances, extract the maximum local offset distance from the set of local offset distances to obtain the target evaluation distance, and compare the target evaluation distance with the preset evaluation distance threshold;

[0071] If the target evaluation distance is greater than or equal to the evaluation distance threshold, then both the extracted initial decomposed image and the target decomposed image are confirmed as abnormal images, and the target decomposed image is used as the extracted initial decomposed image, and return to the step of identifying the target decomposed image in the decomposed image time series based on the initial decomposed image;

[0072] Otherwise, obtain the local offset distance variance based on the set of local offset distances. After confirming that the local offset distance variance is greater than or equal to the preset local offset distance threshold, both the extracted initial decomposed image and the target decomposed image are confirmed as abnormal images, and the target decomposed image is used as the extracted initial decomposed image, and return to the step of identifying the target decomposed image in the decomposed image time series based on the initial decomposed image;

[0073] Summarize the abnormal images to obtain a set of abnormal images.

[0074] To achieve the above object, the present invention also provides an intelligent analysis system for video streams based on neural networks, including:

[0075] An analysis environment confirmation module, configured to receive an analysis instruction and confirm an intelligent analysis environment based on the analysis instruction, where the intelligent analysis environment includes an initial video stream and a video analysis system, and the video analysis system includes: a video decomposition unit, an image recognition unit, an image comparison unit, and a result feedback unit;

[0076] An initial frame number confirmation module, configured to confirm receiving a video decomposition instruction from a video decomposition unit, parse the video decomposition instruction to obtain a first decomposition frame number, and use the first decomposition frame number and an initial video stream to obtain a decomposition image time sequence, where the decomposition image time sequence includes a plurality of decomposition image nodes, and a decomposition image node includes a decomposition image and an image time;

[0077] An initial video division module, configured to obtain a set of decomposition video time points based on the decomposition image time sequence, where the set of decomposition video time points includes M decomposition video time points, and M is an integer greater than or equal to 0, and obtain a set of target video streams based on the set of decomposition video time points and the initial video stream, where the set of target video streams includes one or more target video streams, and perform the following operations on each target video stream in the set of target video streams:

[0078] An abnormal image recognition module, configured to obtain a second decomposition frame number based on a target video stream, confirm an abnormal image set based on the second decomposition frame number and the target video stream, where the abnormal image set includes N abnormal images, and N is an integer greater than or equal to 0, and use a result feedback unit to send the abnormal image set to an initiator of an analysis instruction, so as to implement intelligent analysis of the initial video stream.

[0079] To solve the above problems, the present invention further provides an electronic device, where the electronic device includes:

[0080] A memory, storing at least one instruction; and a processor, executing the instruction stored in the memory to implement the above-mentioned intelligent analysis method for video streams based on a neural network.

[0081] To solve the above problems, the present invention further provides a computer-readable storage medium, where at least one instruction is stored in the computer-readable storage medium, and the at least one instruction is executed by a processor in an electronic device to implement the above-mentioned intelligent analysis method for video streams based on a neural network.

[0082] To solve the problems described in the background art, the present invention receives and confirms a video decomposition instruction from a video decomposition unit, parses the video decomposition instruction to obtain a first decomposition frame number, and uses the first decomposition frame number and the initial video stream to obtain a decomposition image time sequence. The decomposition image time sequence includes multiple decomposition image nodes, and each decomposition image node includes a decomposition image and an image time. Based on the decomposition image time sequence, a set of decomposition video time points is obtained. The set of decomposition video time points includes M decomposition video time points, where M is an integer greater than or equal to 0. Based on the set of decomposition video time points and the initial video stream, a set of target video streams is obtained. The set of target video streams includes one or more target video streams. It can be seen that the present invention considers possible different situations in the initial video stream before analyzing the initial video stream. Therefore, the initial video stream is divided to obtain a set of target video streams, thereby improving the accuracy of analyzing the initial video stream and the degree of intelligence in analyzing the initial video stream. The present invention obtains a second decomposition frame number based on the target video stream, and confirms an abnormal image set based on the second decomposition frame number and the target video stream. It can be seen that the present invention combines the characteristics corresponding to each target video stream to obtain a second decomposition frame number for analyzing the target video stream. Obtaining different second decomposition frame numbers through different characteristics can save the resources required for analyzing the target video stream, thereby improving the degree of intelligence of the embodiments of the present invention. Therefore, the present invention can improve the accuracy and degree of intelligence in analyzing the video stream and reduce the resources required for analyzing the video stream. Description of the Drawings

[0083] Figure 1 Schematic flowchart of a method for intelligent analysis of a video stream based on a neural network provided by an embodiment of the present invention;

[0084] Figure 2 Functional module diagram of an intelligent analysis system for a video stream based on a neural network provided by an embodiment of the present invention;

[0085] Figure 3 Schematic structural diagram of an electronic device for implementing the method for intelligent analysis of a video stream based on a neural network provided by an embodiment of the present invention.

[0086] Description of the reference numerals:

[0087] 1. Electronic device; 10. Processor; 11. Memory; 12. Bus.

[0088] The realization, functional features, and advantages of the objectives of the present invention will be further described in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Embodiments

[0089] It should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0090] An embodiment of the present application provides an intelligent analysis method for video streams based on a neural network. The execution subject of the intelligent analysis method for video streams based on a neural network includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiment of the present application. In other words, the intelligent analysis method for video streams based on a neural network can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.

[0091] Refer to Figure 1 As shown, it is a flowchart of an intelligent analysis method for video streams based on a neural network provided by an embodiment of the present invention. In this embodiment, the intelligent analysis method for video streams based on a neural network includes:

[0092] S1. Receive an analysis instruction, and confirm an intelligent analysis environment based on the analysis instruction. Among them, the intelligent analysis environment includes an initial video stream and a video analysis system. Among them, the video analysis system includes: a video decomposition unit, an image recognition unit, an image comparison unit, and a result feedback unit.

[0093] It should be explained that the analysis instruction is an instruction for realizing the analysis of the video. The intelligent analysis environment is a necessary environment for realizing the intelligent analysis of the video, and the intelligent analysis environment includes an initial video stream and a video analysis system. Among them, the video analysis system includes: a video decomposition unit, an image recognition unit, an image comparison unit, and a result feedback unit. For the specific application of the unit, please refer to the subsequent embodiments. Generally, when analyzing the initial video stream, an object to be analyzed can be selected from the initial video stream to facilitate reducing the resources required for analyzing the initial video stream.

[0094] It can be understood that the initial video stream refers to the video to be intelligently analyzed. The embodiment of the present invention mainly aims to realize the recognition and analysis of the modified content in the video stream, so as to ensure the security and correctness of the video stream. And when analyzing the initial video stream, different frame numbers are set according to the specific content of the video, thereby saving the resources required for analyzing the initial video stream.

[0095] Exemplarily, in order to determine whether the video stream as evidence has been tampered with, therefore, the analysis instruction is issued by the holder of the video stream, and the intelligent analysis environment is confirmed according to the analysis instruction. Here, the video stream as evidence is the initial video stream.

[0096] S2. Confirm the receipt of the video decomposition instruction from the video decomposition unit, parse the video decomposition instruction to obtain the first decomposition frame number, and use the first decomposition frame number and the initial video stream to obtain the decomposition image time sequence, where the decomposition image time sequence includes multiple decomposition image nodes, and the decomposition image nodes include decomposition images and image times.

[0097] It should be explained that the first decomposition frame number refers to the preset number of frames used to divide the initial video stream. Optionally, the first decomposition frame number is obtained by manual setting. The decomposition image time sequence refers to the sequence corresponding to the decomposition images extracted from the initial video stream according to the first decomposition frame number. For example, if the initial video stream includes 1000 images in total and the preset first decomposition frame number is 10, then the 10th image, 20th image, 30th image,..., 990th image, and 1000th image are extracted from the initial video stream respectively according to the first decomposition frame number, and the extracted images are sorted in the order of the time corresponding to the shooting time of the images from the earliest to the latest to obtain the decomposition image time sequence, where the extracted images are the decomposition images, and the time corresponding to the shooting time of the images is the image time. Generally, the setting of the first decomposition frame number can be combined with the number of frames and duration of the initial video stream.

[0098] S3. Obtain a set of decomposition video time points based on the decomposition image time sequence, where the set of decomposition video time points includes M decomposition video time points, and M is an integer greater than or equal to 0. Obtain a set of target video streams based on the set of decomposition video time points and the initial video stream, where the set of target video streams includes one or more target video streams.

[0099] It should be understood that obtaining the set of decomposition video time points based on the decomposition image time sequence includes:

[0100] Obtain an analysis grayscale image time sequence based on the decomposition image time sequence, and perform the following operations on each analysis grayscale image in the analysis grayscale image time sequence:

[0101] Obtain an analysis grayscale mean value based on the analysis grayscale image, where the analysis grayscale mean value is the mean value of multiple grayscale values corresponding to the analysis grayscale image;

[0102] Associate the analysis grayscale mean value and the image time to obtain an initial analysis node, and summarize the initial analysis nodes to obtain an initial analysis node set;

[0103] Obtain one or more target analysis node sets by using a pre-constructed clustering model and the initial analysis node set;

[0104] Based on one or more target analysis node sets, confirm the set of decomposition video time points in the initial video stream.

[0105] Further, perform grayscale transformation on each decomposed image in the decomposed image time series to obtain the analyzed grayscale image time series. The grayscale transformation refers to the operation of converting a color image into a grayscale image, and the grayscale transformation is a prior art and will not be elaborated here. Optionally, the k-means clustering algorithm is used as the clustering model, and the same effect can be achieved by using other technologies, which will not be elaborated here. Through the clustering model, the classification of the initial video stream can be realized by combining the features of time and image. Here, the feature of the image refers to the mean value of the grayscale values corresponding to the image. For example, using the k-means clustering algorithm, cluster 10 initial analysis nodes to obtain three clusters, where each initial analysis node included in each cluster constitutes the target analysis node set.

[0106] It should be explained that each target analysis node set includes multiple different initial analysis nodes, and each initial analysis node corresponds to an image time. Therefore, based on each target analysis node set in one or more target analysis node sets, a time period can be confirmed, and the boundary point between adjacent time periods can be used as the decomposed video time point. For example, there are currently three target analysis node sets. Among them, the time period corresponding to the first target analysis node set is from 10 to 15.8 seconds, the time period corresponding to the second target analysis node set is from 15.9 to 20 seconds, and the time period corresponding to the third target analysis node set is from 0 to 9.9 seconds. Then, two decomposed video time points can be confirmed using these three target analysis node sets, where the two decomposed video time points are 9.9 seconds and 15.8 seconds respectively.

[0107] It can be understood that obtaining the target video stream set based on the decomposed video time point set and the initial video stream means dividing the initial video stream using each decomposed video time point in the decomposed video time point set, and the divided initial video stream is the target video stream.

[0108] S4. Obtain the second decomposition frame number based on the target video stream, and confirm the abnormal image set based on the second decomposition frame number and the target video stream. The abnormal image set includes N abnormal images, and N is an integer greater than or equal to 0. Use the result feedback unit to send the abnormal image set to the initiator of the analysis instruction to realize the intelligent analysis of the initial video stream.

[0109] It can be understood that obtaining the second decomposition frame number based on the target video stream includes:

[0110] Obtain the initial number of frames and the video duration of the target video stream, obtain the number of images based on the initial number of frames and the video duration, obtain an image extraction gradient set for extracting images, and extract an image time series from the target video stream based on the number of images and the image extraction gradient set. The image extraction gradient set includes multiple image extraction ratios, and the image time series includes multiple initial images;

[0111] Confirm the receipt of an image recognition instruction from the image recognition unit, and confirm an image recognition model set based on the image recognition instruction. The image recognition model set includes multiple image recognition models. Sequentially extract initial images from the image time series, and perform the following operations on the extracted initial images:

[0112] Randomly extract two image recognition models from the image recognition model set to obtain a first recognition model and a second recognition model. Recognize the extracted initial images based on the first recognition model and the second recognition model to obtain a first recognition name set and a second recognition name set. The first recognition name set includes multiple first recognition nodes, where each first recognition node includes a first recognition name and a first recognition quantity. The second recognition name set includes multiple second recognition nodes, where each second recognition node includes a second recognition name and a second recognition quantity;

[0113] After confirming a target recognition name set using the first recognition name set and the second recognition name set, where the target recognition name set includes multiple target recognition nodes, and each target recognition node includes a target recognition name and a target recognition quantity, perform the following operations on each target recognition name in the target recognition name set:

[0114] Identify a local region image corresponding to the target recognition name in the initial image, label the local region image using the target recognition name and the target recognition quantity to obtain a target image with an identification serial number, summarize the target images to obtain a target image set, and obtain a second decomposition number of frames using the target image set.

[0115] Further, the initial number of frames refers to the number of frames of the target video stream, the video duration refers to the duration of the target video stream, the number of images refers to the number of images included in the target video stream, and the number of images is the product of the initial number of frames and the video duration. The image extraction ratio refers to the ratio used to divide the number of images. For example: if the number of frames of the target video stream is 24 FPS and the video duration is 10 seconds, then the number of images is 240. The image extraction gradients collectively include multiple image extraction ratios, and the multiple image extraction ratios are respectively: 0.2, 0.4, 0.6, 0.8. Then, through the image extraction ratio, it is calculated that the 48th frame image, the 96th frame image, the 144th frame image, and the 192nd frame image need to be extracted from the initial video stream. The 48th frame image, the 96th frame image, the 144th frame image, and the 192nd frame image constitute the image time sequence, and each frame image in the image time sequence is the initial image.

[0116] It should be understood that the image recognition model refers to a model used to recognize information in an image. Optionally, a pre-trained neural network model is used as the image recognition model. The first recognition name refers to the name of the information recognized in the initial image using the first recognition model, the second recognition name refers to the name of the information recognized in the initial image using the second recognition model, the first recognition quantity refers to the quantity corresponding to the first recognition name, and the second recognition quantity refers to the quantity corresponding to the second recognition name. For example: an image includes 5 pedestrians and 3 vehicles. Then, in an ideal situation, using a neural network model, two recognition nodes can be recognized. Here, the recognition node is the first recognition node or the second recognition node. Among them, the recognition name and recognition quantity corresponding to the first recognition node are respectively: person and 5, and the recognition name and recognition quantity corresponding to the second recognition node are respectively: vehicle and 3.

[0117] It should be noted that the local area image refers to the area of the image corresponding to the target recognition name. Generally, a neural network model can be used to recognize the initial image and divide the local area image, and the same effect can be achieved by using other technologies, which will not be elaborated here. The purpose of using the target recognition name to label the local area image is to distinguish the content of the images corresponding to different local areas in the initial image. In order to achieve better distinction of different areas, when labeling different contents corresponding to the same recognition name, the form of marking serial numbers can be used. The way to obtain the serial numbers here can be to mark from left to right or sort from top to bottom, and the largest number in the serial numbers is the target recognition quantity. For example: if there are 3 people in an image, then the areas of each of the 3 people are recognized in the initial image, and the area is labeled as a person, and the areas corresponding to the initial image labeled as a person are respectively labeled as person-1, person-2, person-3, and the target recognition quantity corresponding to the target recognition node is 3. Generally, when labeling the same information, the same serial number should be used to facilitate the analysis of the same image.

[0118] Further, the confirmation of the target recognition name set by using the first recognition name set and the second recognition name set includes:

[0119] Perform the following operations on each first recognition node in the first recognition name set:

[0120] Judge whether there is a second recognition node in the second recognition name set that is the same as the first recognition node;

[0121] If there is no second recognition node in the second recognition name set that is the same as the first recognition node, randomly extract multiple initial image recognition models from the image recognition model set, where the number of multiple initial image recognition models is the preset extraction quantity. Use multiple initial image recognition models and the initial image to obtain multiple initial recognition name sets, where the initial image recognition models and the initial recognition name sets are in one-to-one correspondence, and the initial recognition name set includes multiple initial recognition nodes;

[0122] In a combined form, combine multiple initial recognition name sets to obtain multiple combined recognition name sets, where the combined recognition name set includes two initial recognition name sets. Confirm one or more fusion recognition name sets in the multiple combined recognition name sets, where the two initial recognition name sets in the fusion recognition name set are the same;

[0123] Eliminate any one of the initial recognition name sets corresponding to each fused recognition name set in one or more fused recognition name sets from multiple initial recognition name sets to obtain an updated recognition name set. Use the updated recognition name set as the multiple initial recognition name sets, and return the step of combining the multiple initial recognition name sets in a combined form until one or more initial fused name sets are obtained, where the multiple initial recognition name sets corresponding to the initial fused name set are all the same;

[0124] Count the number of initial recognition name sets corresponding to each initial fused name set in one or more initial fused name sets respectively to obtain one or more initial fusion quantities, where the initial fusion quantity corresponds to the initial fused name set one by one;

[0125] Extract the largest initial fusion quantity from one or more initial fusion quantities to obtain a target verification quantity, calculate the ratio of the target verification quantity to the extraction quantity to obtain a correct ratio. When the correct ratio is greater than or equal to a preset ratio threshold, use the fused recognition name set corresponding to the target verification quantity as the target recognition name set;

[0126] Otherwise, return the step of randomly extracting multiple initial image recognition models from the image recognition model set until a target recognition name set is obtained.

[0127] It can be understood that when there is a second recognition node in the second recognition name set that is different from the first recognition node, the first recognition name set is different from the second recognition name set, that is, the results of recognizing the initial image using different image recognition models are different. For example, if the preset extraction quantity is 5, then randomly extract 5 initial image recognition models from the image recognition model set, and use the 5 initial image recognition models to recognize the initial image respectively to obtain 5 initial recognition name sets. Through a combined form, the 5 initial image recognition models can be combined into 10 combined recognition name sets. If there are two initial recognition name sets corresponding to a combined recognition name set that are the same among the 10 combined recognition name sets, then record this combined recognition name set as a fused recognition name set. Generally, when specific application scenarios are different, different methods can be adopted to realize the construction of the neural network model, and further improve the accuracy of information recognition in the image using the neural network model.

[0128] It should be explained that using different neural network models to recognize the same object may result in different results. Therefore, in the embodiment of the present invention, in a cyclic form, the same initial recognition name sets are all merged together, and during the cycle, the same initial recognition name sets are continuously eliminated from the combination main body to save resources required during the cycle. Here, the step of continuously updating the combination main body refers to using the more recognition name set as the multiple initial recognition name sets.

[0129] It should be understood that different initial image recognition models may recognize different information for the same content in an image. Therefore, a combined method can be used to classify the results of image recognition by the initial image recognition models as much as possible. When the correct ratio is greater than or equal to a preset ratio threshold, it indicates that the result of recognizing the initial image by the initial image recognition model is relatively accurate. Otherwise, there may be a large error in the result of recognizing the initial image.

[0130] Furthermore, after the correct ratio is greater than or equal to the ratio threshold, a relatively accurate set of target recognition names is obtained. Furthermore, using the relatively accurate set of target recognition names can improve the accuracy of obtaining the second decomposition frame number, and thus save the energy consumption required for analyzing the initial video stream. For the specific implementation method of the beneficial effects, please refer to the subsequent embodiments.

[0131] It should be explained that obtaining the second decomposition frame number by using the target image set includes:

[0132] Summarize the target image set to obtain multiple target image sets, and use the target recognition name, serial number, and target image to obtain the target image time sequence;

[0133] Successively extract the initial analysis images from the target image time sequence, and perform the following operations on the extracted initial analysis images:

[0134] Based on the initial analysis image, confirm the target analysis image in the target image time sequence, where the target analysis image is adjacent to and lags behind the initial analysis image in the target image time sequence;

[0135] Obtain the initial coordinate point set and the target coordinate point set based on the initial analysis image and the target analysis image;

[0136] Map the initial coordinate points in the initial coordinate point set and the target coordinate points in the target coordinate point set to a pre-constructed reference coordinate system respectively to obtain a mapped coordinate point set, where the mapped coordinate point set includes multiple mapped coordinate points;

[0137] Count the number of reference coordinate points corresponding to each mapped coordinate point in the statistical mapped coordinate point set to obtain a reference quantity set, where the reference coordinate points are initial coordinate points or target coordinate points, and the reference quantity set includes multiple reference quantities. Count the number of reference quantities in the reference quantity set that are 2 to obtain a target reference quantity. Count the number of mapped coordinate points in the mapped coordinate point set to obtain a comprehensive quantity. Calculate the ratio of the target reference quantity to the comprehensive quantity to obtain a reference quantity ratio. If the reference quantity ratio is less than or equal to a preset reference ratio threshold, use the target analysis image as the initial analysis image and return the step of confirming the target analysis image in the target image time series based on the initial analysis image. If the reference quantity ratios of the reference quantity sets corresponding to the target image time series are all less than or equal to the reference ratio threshold, label the target image as a background image; otherwise, label the target image as an initial detection image. After confirming that the initial detection image is a preset target detection image, use the target detection image to identify an adjacent image set in the initial image, where the adjacent image set includes one or more adjacent images, and the adjacent images are target detection images or background images. Obtain a second decomposition frame number based on the adjacent image set and the detection image.

[0138] It can be understood that the target image time series refers to an image sequence corresponding to different initial images of the same object. For example, in each target image of the target image set, a person labeled 1 is recognized, and the target images in the target image set are sorted in ascending order of the time corresponding to the target images to obtain the target image time series.

[0139] It should be understood that the initial coordinate point set refers to the set of pixel coordinates corresponding to each pixel point in the initial analysis image. The target coordinate point set is obtained in the same way as the initial coordinate point set, which will not be elaborated here. Optionally, the image coordinate system is used as the reference coordinate system, and the same effect can be achieved by using other technologies, which will not be elaborated here. Generally speaking, if the initial coordinate points in the initial coordinate point set are the same as the target coordinate points in the target coordinate point set, it indicates that the object corresponding to the target image is stationary. Therefore, when the reference quantity corresponding to the target image time series is 2, it can be shown that the object corresponding to the target image in the initial video stream is stationary. Here, the number of the reference quantity set corresponding to the target image time series is related to the number of target images included in the target image time series. For example, if there are 5 target images in the target image time series, 4 reference quantity sets can be obtained by using the 5 target images, and the 4 reference quantity sets are obtained from the first target image and the second target image, the second target image and the third target image, the third target image and the fourth target image, and the fourth target image and the fifth target image respectively. When the reference quantity ratio is greater than the reference ratio threshold, it indicates that the object corresponding to the target image is moving. Therefore, it is necessary to determine whether the object corresponding to the target image is in a moving state. Generally speaking, when the object is in a moving state, it is meaningful to analyze the object.

[0140] Further, the confirmation that the initial detection image is a preset target detection image includes:

[0141] Obtain the detection image time series according to the target recognition name, serial number, and the initial detection image. Among them, the detection image time series includes multiple initial detection images. Perform grayscale operation on each initial detection image in the detection image time series to obtain the grayscale detection image time series. Based on the preset sliding window, sequentially extract the analysis image groups from the grayscale detection image time series. Among them, the analysis image group includes two grayscale detection images;

[0142] Perform the following operations on each grayscale detection image in the analysis image group:

[0143] Calculate the image center coordinates based on the grayscale detection image. The calculation formula is as follows:

[0144] ,

[0145] Among them, respectively represent the abscissa and ordinate of the image center coordinates, represents that there are pixel points in the grayscale detection image, represents the th pixel point in the grayscale detection image, respectively represent the The abscissa and ordinate corresponding to a pixel point

[0146] Summarize the image center coordinates to obtain an image center coordinate set, calculate the comprehensive image change degree based on the image center coordinate set, and judge the comprehensive image change degree and a preset comprehensive change degree threshold

[0147] If the comprehensive image change degree is greater than or equal to the comprehensive change degree threshold, confirm that the initial detection image is the target detection image

[0148] Furthermore, the sliding window refers to a window with a fixed size, and a sliding step of the sliding window is set. The grayscale operation and the grayscale transformation can achieve the same effect, which will not be elaborated here. The method for obtaining the detection image time series is the same as the method for obtaining the target image time series, which will not be elaborated here

[0149] It should be explained that calculating the comprehensive image change degree based on the image center coordinate set includes

[0150] Obtain the image center coordinate time series based on the image center coordinate set, use the sliding window to sequentially extract analysis coordinate groups from the image center coordinate time series. Among them, an analysis coordinate group includes two image center coordinates, and obtain the Euclidean distance between the two image center coordinates in the analysis coordinate group to get the moving distance

[0151] Extract the first image center coordinate and the last image center coordinate from the image center coordinate time series to obtain the initial center coordinate and the target center coordinate, and obtain the evaluation distance based on the initial center coordinate and the target center coordinate

[0152] Summarize the moving distances to obtain a moving distance set, and use the moving distance set to obtain the moving distance variance. Among them, the moving distance variance is the variance of multiple moving distances in the moving distance set

[0153] Calculate the comprehensive image change degree based on the moving distance variance and the evaluation distance. The calculation formula is as follows

[0154] ,

[0155] Among them represents the comprehensive image change degree are all preset coefficients represents the evaluation distance represents the moving distance variance

[0156] It is understandable that the method for obtaining the evaluation distance based on the initial center coordinates and the target center coordinates is the same as the method for obtaining the movement distance, which will not be elaborated here. Generally, there may be moving objects that belong to the background in the initial video stream. For example, the leaves of a tree blown by the wind. Here, the leaves are moving while the tree is stationary. Therefore, by integrating the image change degree, the dynamic objects that are regarded as the background can be distinguished from the actual moving objects to be analyzed. Furthermore, the energy consumption for analyzing the objects that actually need to be analyzed can be saved, and the timeliness of analyzing the objects that actually need to be analyzed can be improved.

[0157] Further, the obtaining of the second decomposition frame number based on the adjacent image set and the detection image includes:

[0158] Based on the adjacent image set and the detection image, an adjacent grayscale image set and a detection grayscale image are obtained. Using a pre-constructed region growing algorithm and a pre-constructed image grayscale gradient set, one or more local region images are extracted from the detection grayscale image, and one or more target region images are identified in the one or more local region images. The image grayscale gradient set includes multiple image gradient grayscale values, and the target region image is adjacent to at least one adjacent grayscale image in the adjacent grayscale image set;

[0159] The following operations are performed on each of the one or more target region images in the one or more target region images:

[0160] Based on the target region image, the local region grayscale mean value is obtained. The local region grayscale mean value is the mean value of the multiple grayscale values corresponding to the target region image. The target region image and the adjacent grayscale image adjacent to the target region image are associated to obtain an associated image set;

[0161] Starting from the local region grayscale mean value, using the region growing algorithm, an associated region image is extracted from the associated image set, and a target discrimination image is identified in the associated region image. The target discrimination image is the region of the adjacent grayscale image adjacent to the target region image in the associated image set;

[0162] Based on the target discrimination image, the target discrimination mean value is obtained. According to the target discrimination mean value and the local region grayscale mean value, the discrimination frame number is calculated, and the calculation formula is as follows:

[0163] ,

[0164] ,

[0165] Wherein, represents the discrimination frame number, represents the initial frame number, represents the target discrimination mean value, represents the local region grayscale mean value, represents a preset coefficient, represents a preset coefficient, represents the rounding symbol;

[0166] Sum up the discrimination frame numbers to obtain a discrimination frame number set, and confirm the second decomposition frame number based on the discrimination frame number set, where the second decomposition frame number is the smallest discrimination frame number in the discrimination frame number set.

[0167] It should be explained that grayscale operations are respectively performed on the adjacent images and the detection image in the adjacent image set to obtain an adjacent grayscale image set and a detection grayscale image. The image grayscale image set includes multiple image gradient grayscale values. The method for extracting one or more local region images from the detection grayscale image by using the region growing algorithm and the image grayscale gradient set is as follows: sequentially extract the image gradient grayscale values from the image grayscale gradient set, and perform the following operations on the extracted image gradient grayscale values:

[0168] Taking the image gradient grayscale value as the starting point, use the region growing algorithm to retrieve in the detection grayscale image to obtain an initial retrieval region, count the number of pixel points corresponding to the initial retrieval region to obtain a statistical quantity. If the statistical quantity is greater than or equal to a preset statistical threshold, record the initial retrieval region as a local region image, and sum up the local region images to obtain one or more local region images.

[0169] Furthermore, the technique of using a fixed grayscale value as the starting point of the region growing algorithm and retrieving in the image by using the region growing algorithm is a prior art and will not be elaborated here. Generally, there may be abnormal factors such as noise in the image. Therefore, by setting the form of the statistical threshold, the noise in the image can be removed to improve the accuracy of obtaining the target region image. Generally, currently, the fact that the region image is adjacent to the adjacent grayscale image means that the regions corresponding to the two images are adjacent.

[0170] It should be understood that the target discrimination image refers to an image region with grayscale values similar to those of the target region image. The higher the similarity between the target discrimination image and the target region image, the smaller the required second decomposition frame number. For example: when there are significant differences between the target discrimination image and the target region image, it is easy to distinguish the target discrimination image from the target region image in the initial video stream. The target discrimination mean refers to the mean of multiple grayscale values corresponding to the target discrimination image.

[0171] It should be understood that the confirmation of the abnormal image set based on the second decomposition frame number and the target video stream includes:

[0172] Extract a decomposed image sequence from the target video stream using the second number of decomposed frames. The decomposed image sequence includes multiple initial decomposed images. Sequentially extract the initial decomposed images from the decomposed image sequence, and perform the following operations on the extracted initial decomposed images:

[0173] Based on the initial decomposed image, identify a target decomposed image in the decomposed image sequence. The target decomposed image is adjacent to the initial decomposed image and lags behind the extracted initial decomposed image;

[0174] Confirm receiving an image comparison instruction from the image comparison unit, parse the image comparison instruction to obtain a local name database, and use the image recognition model set and the local name database to respectively identify a target labeled decomposed image set and an initial labeled decomposed image set in the target decomposed image and the initial decomposed image. The target labeled decomposed image set includes multiple target labeled decomposed images, and the initial labeled decomposed image set includes multiple initial labeled decomposed images;

[0175] Match the target labeled decomposed images in the target labeled decomposed image set and the initial labeled decomposed images in the initial labeled decomposed image set to obtain multiple labeled decomposition nodes. The labeled decomposition nodes include the initial labeled decomposed images and the target labeled decomposed images;

[0176] Perform the following operations on each of the multiple labeled decomposition nodes:

[0177] Obtain a local offset distance based on the labeled decomposition node, summarize the local offset distances to obtain a set of local offset distances, extract the maximum local offset distance from the set of local offset distances to obtain a target evaluation distance, and compare the target evaluation distance with a preset evaluation distance threshold;

[0178] If the target evaluation distance is greater than or equal to the evaluation distance threshold, both the extracted initial decomposed image and the target decomposed image are identified as abnormal images, and the target decomposed image is used as the extracted initial decomposed image, and return to the step of identifying the target decomposed image in the decomposed image sequence based on the initial decomposed image;

[0179] Otherwise, obtain a local offset distance variance based on the set of local offset distances. After confirming that the local offset distance variance is greater than or equal to a preset local offset distance threshold, both the extracted initial decomposed image and the target decomposed image are identified as abnormal images, and the target decomposed image is used as the extracted initial decomposed image, and return to the step of identifying the target decomposed image in the decomposed image sequence based on the initial decomposed image;

[0180] Summarize the abnormal images to obtain a set of abnormal images.

[0181] It should be noted that the method for obtaining the decomposed image time series is the same as that for obtaining the detected image time series, which will not be elaborated here. The local name database refers to a database that stores multiple local names. For example, the local names corresponding to a person may include: hands, arms, mouth, eyes, etc. The methods for respectively identifying the target identification decomposed image set and the initial identification decomposed image set in the target decomposed image and the initial decomposed image by using the image recognition model set and the local name database are the same as the method for obtaining the target image, which will not be elaborated here. Matching the target identification decomposed image and the initial identification decomposed image means using the local name and serial number to match the target identification decomposed image and the initial identification decomposed image. For example: the target identification decomposed image set includes hand-1, hand-2, and the initial identification decomposed image set includes hand-1, hand-2. Then, through the local name and serial number, the hand-1 and hand-1, hand-2 and hand-2 in the target identification decomposed image set and the initial identification decomposed image set can be matched to obtain two identification decomposition nodes. The method for obtaining the local offset distance is the same as the method for obtaining the evaluation distance, which will not be elaborated here. The local offset distance variance refers to the variance of multiple local offset distances corresponding to the local offset distance set. Generally, when the target evaluation distance is greater than or equal to the evaluation distance threshold, it indicates that there are traces of modification in the local area of the initial video stream. When the local offset distance variance is greater than or equal to the local offset distance threshold, it indicates that the initial video stream may have been clipped. Therefore, both the qualified initial decomposed image and the target decomposed image are confirmed as abnormal images.

[0182] It can be understood that in the same target video stream, there may be clipping or modification of the target video stream. When a certain part of the target video stream is modified, it will cause the target evaluation distance to be greater than or equal to the evaluation distance threshold. When the target video stream is clipped, spliced, etc., it will cause the local offset distance variance to be greater than or equal to the local offset distance threshold, that is, there will be a visual effect of fragmentation between two frames in the target video stream. Therefore, after sending the abnormal image set to the initiator of the analysis instruction, the identification of abnormal image frames in the initial video stream can be achieved.

[0183] To solve the problems described in the background art, the present invention receives and confirms a video decomposition instruction from a video decomposition unit, parses the video decomposition instruction to obtain a first decomposition frame number, and uses the first decomposition frame number and an initial video stream to obtain a decomposed image time sequence. The decomposed image time sequence includes a plurality of decomposed image nodes, and each decomposed image node includes a decomposed image and an image time. Based on the decomposed image time sequence, a set of decomposed video time points is obtained. The set of decomposed video time points includes M decomposed video time points, where M is an integer greater than or equal to 0. Based on the set of decomposed video time points and the initial video stream, a set of target video streams is obtained. The set of target video streams includes one or more target video streams. It can be seen that the present invention considers possible different situations in the initial video stream before analyzing the initial video stream. Therefore, the initial video stream is divided to obtain a set of target video streams, thereby improving the accuracy of analyzing the initial video stream and the degree of intelligence in analyzing the initial video stream. The present invention obtains a second decomposition frame number based on the target video stream, and confirms an abnormal image set based on the second decomposition frame number and the target video stream. It can be seen that the present invention combines the characteristics corresponding to each target video stream to obtain a second decomposition frame number for analyzing the target video stream. Obtaining different second decomposition frame numbers through different characteristics can save the resources required for analyzing the target video stream, thereby improving the degree of intelligence of the embodiments of the present invention. Therefore, the present invention can improve the accuracy and degree of intelligence in analyzing the video stream and reduce the resources required for analyzing the video stream.

[0184] As Figure 2 shown, it is a functional module diagram of an intelligent analysis system for video streams based on a neural network provided by an embodiment of the present invention.

[0185] The intelligent analysis system 100 for video streams based on a neural network according to the present invention can be installed in an electronic device. According to the functions achieved, the intelligent analysis system 100 for video streams based on a neural network can include an analysis environment confirmation module 101, an initial frame number confirmation module 102, an initial video division module 103, and an abnormal image recognition module 104. The modules in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0186] The analysis environment confirmation module 101 is configured to receive an analysis instruction and confirm an intelligent analysis environment based on the analysis instruction. The intelligent analysis environment includes an initial video stream and a video analysis system. The video analysis system includes: a video decomposition unit, an image recognition unit, an image comparison unit, and a result feedback unit;

[0187] The initial frame number confirmation module 102 is configured to confirm receiving a video decomposition instruction from a video decomposition unit, parse the video decomposition instruction to obtain a first decomposition frame number, and use the first decomposition frame number and an initial video stream to obtain a decomposition image timing sequence, where the decomposition image timing sequence includes a plurality of decomposition image nodes, and each decomposition image node includes a decomposition image and an image time;

[0188] The initial video division module 103 is configured to obtain a set of decomposition video time points based on the decomposition image timing sequence, where the set of decomposition video time points includes M decomposition video time points, and M is an integer greater than or equal to 0, and obtain a set of target video streams based on the set of decomposition video time points and the initial video stream, where the set of target video streams includes one or more target video streams, and the following operations are performed on each target video stream in the set of target video streams:

[0189] The abnormal image recognition module 104 is configured to obtain a second decomposition frame number based on a target video stream, confirm an abnormal image set based on the second decomposition frame number and the target video stream, where the abnormal image set includes N abnormal images, and N is an integer greater than or equal to 0, and send the abnormal image set to the initiator of the analysis instruction by using a result feedback unit to implement intelligent analysis of the initial video stream.

[0190] Specifically, each module in the intelligent analysis system 100 for video stream based on neural network in the embodiment of the present invention adopts the same technical means as those in the above-mentioned Figure 1 intelligent analysis method for video stream based on neural network, and can produce the same technical effects, which will not be elaborated here.

[0191] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing an intelligent analysis method for video stream based on neural network provided by an embodiment of the present invention.

[0192] The electronic device 1 may include a processor 10, a memory 11, and a bus 12, and may further include a computer program stored in the memory 11 and executable on the processor 10, such as an intelligent analysis method program for video stream based on neural network.

[0193] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 11 can be an internal storage unit of the electronic device 1 in some embodiments, such as the mobile hard disk of the electronic device 1. The memory 11 can also be an external storage device of the electronic device 1 in some other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 also includes the internal storage unit of the electronic device 1 and also includes an external storage device. The memory 11 can not only be used to store application software installed on the electronic device 1 and various types of data, such as the code of the intelligent analysis method program for video stream based on neural network, etc., but also be used to temporarily store data that has been output or will be output.

[0194] The processor 10 can be composed of integrated circuits in some embodiments. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged, including the combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as the intelligent analysis method program for video stream based on neural network, etc.), and calling the data stored in the memory 11, to execute various functions of the electronic device 1 and process data.

[0195] The bus 12 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus 12 can be divided into an address bus, a data bus, a control bus, etc. The bus 12 is set to realize the connection and communication between the memory 11 and at least one processor 10, etc.

[0196] Figure 3 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 3The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0197] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to the at least one processor 10 through a power management system, so as to implement functions such as charging management, discharging management, and power consumption management through the power management system. The power source may also include any components such as one or more DC or AC power sources, a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 1 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0198] Furthermore, the electronic device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.

[0199] Optionally, the electronic device 1 may further include a user interface. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.

[0200] The program of the intelligent analysis method for video streams implemented based on a neural network stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can implement:

[0201] Receive an analysis instruction, and confirm an intelligent analysis environment based on the analysis instruction. Among them, the intelligent analysis environment includes an initial video stream and a video analysis system. The video analysis system includes: a video decomposition unit, an image recognition unit, an image comparison unit, and a result feedback unit;

[0202] Confirm to receive a video decomposition instruction from the video decomposition unit, parse the video decomposition instruction to obtain a first decomposition frame number, and use the first decomposition frame number and the initial video stream to obtain a decomposition image time series. The decomposition image time series includes multiple decomposition image nodes, and the decomposition image nodes include decomposition images and image times;

[0203] Obtain a set of decomposition video time points based on the decomposition image time series. Among them, the set of decomposition video time points includes M decomposition video time points, and M is an integer greater than or equal to 0. Obtain a set of target video streams based on the set of decomposition video time points and the initial video stream. Among them, the set of target video streams includes one or more target video streams. Perform the following operations on each target video stream in the set of target video streams:

[0204] Obtain a second decomposition frame number based on the target video stream. Based on the second decomposition frame number and the target video stream, confirm an abnormal image set. Among them, the abnormal image set includes N abnormal images, and N is an integer greater than or equal to 0. Use the result feedback unit to send the abnormal image set to the initiator of the analysis instruction to realize the intelligent analysis of the initial video stream.

[0205] Specifically, for the specific implementation method of the above instructions by the processor 10, reference can be made to Figures 1 to 3 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.

[0206] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or system capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory).

[0207] The present invention also provides a computer-readable storage medium. The readable storage medium stores a computer program, and when the computer program is executed by the processor of the electronic device, it can realize:

[0208] Receive an analysis instruction, and confirm an intelligent analysis environment based on the analysis instruction. Among them, the intelligent analysis environment includes an initial video stream and a video analysis system. Among them, the video analysis system includes: a video decomposition unit, an image recognition unit, an image comparison unit, and a result feedback unit;

[0209] Confirm to receive a video decomposition instruction from the video decomposition unit, parse the video decomposition instruction to obtain a first decomposition frame number, and use the first decomposition frame number and the initial video stream to obtain a decomposition image time series. Among them, the decomposition image time series includes multiple decomposition image nodes, and the decomposition image nodes include decomposition images and image times;

[0210] Obtain a set of decomposition video time points based on the decomposed image time series. Among them, the set of decomposition video time points includes M decomposition video time points, and M is an integer greater than or equal to 0. Obtain a set of target video streams based on the set of decomposition video time points and the initial video stream. Among them, the set of target video streams includes one or more target video streams. Perform the following operations on each target video stream in the set of target video streams:

[0211] Obtain a second decomposition frame number based on the target video stream. Identify a set of abnormal images based on the second decomposition frame number and the target video stream. Among them, the set of abnormal images includes N abnormal images, and N is an integer greater than or equal to 0. Use the result feedback unit to send the set of abnormal images to the initiator of the analysis instruction to realize the intelligent analysis of the initial video stream.

[0212] In several embodiments provided by the present invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative, and there can be other partitioning methods in actual implementation.

[0213] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0214] In addition, the functional modules in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.

[0215] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0216] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for intelligent analysis of video streams based on neural networks, characterized in that: The method comprises: Receiving an analysis instruction, and confirming an intelligent analysis environment based on the analysis instruction, wherein the intelligent analysis environment includes an initial video stream and a video analysis system, wherein the video analysis system includes: a video decomposition unit, an image recognition unit, an image comparison unit, and a result feedback unit; Confirming receipt of a video decomposition instruction from a video decomposition unit, parsing the video decomposition instruction to obtain a first decomposition frame number, and obtaining a decomposition image time sequence using the first decomposition frame number and an initial video stream, wherein the decomposition image time sequence includes a plurality of decomposition image nodes, and the decomposition image nodes include decomposition images and image time; A decomposed video time point set is obtained based on the decomposed image time sequence, wherein the decomposed video time point set includes M decomposed video time points, and M is an integer greater than or equal to 0; a target video stream set is obtained based on the decomposed video time point set and the initial video stream, wherein the target video stream set includes one or more target video streams, and the following operations are performed on each target video stream in the target video stream set: Acquiring a second decomposed frame number based on a target video stream, and confirming an abnormal image set based on the second decomposed frame number and the target video stream, wherein the abnormal image set includes N abnormal images, and N is an integer greater than or equal to 0, and confirming the abnormal image set based on the second decomposed frame number and the target video stream includes: A decomposed image sequence is extracted from the target video stream using the second decomposed frame number, wherein the decomposed image sequence includes a plurality of initial decomposed images, initial decomposed images are sequentially extracted from the decomposed image sequence, and the following operations are performed on the extracted initial decomposed images: Based on the initial decomposition image, a target decomposition image is identified in the decomposition image time sequence, wherein the target decomposition image is adjacent to the initial decomposition image and lags behind the extracted initial decomposition image; Confirming receiving an image comparison instruction from an image comparison unit, parsing the image comparison instruction, obtaining a local name database, and using the image recognition model set and the local name database to respectively identify a target identification decomposition image set and an initial identification decomposition image set in the target decomposition image and the initial decomposition image, wherein the target identification decomposition image set includes a plurality of target identification decomposition images, and the initial identification decomposition image set includes a plurality of initial identification decomposition images; Matching the target identification decomposition image in the target identification decomposition image set and the initial identification decomposition image in the initial identification decomposition image set to obtain a plurality of identification decomposition nodes, wherein the identification decomposition nodes include the initial identification decomposition image and the target identification decomposition image; The following operations are performed on each of the multiple identity decomposition nodes: Obtaining local offset distances based on the identified decomposition nodes, summarizing the local offset distances to obtain a local offset distance set, extracting the maximum local offset distance from the local offset distance set to obtain a target evaluation distance, and comparing the target evaluation distance with a preset evaluation distance threshold; If the target evaluation distance is greater than or equal to the evaluation distance threshold, both the extracted initial decomposition image and the target decomposition image are confirmed as abnormal images, and the target decomposition image is used as the extracted initial decomposition image, and the process returns to the step of confirming the target decomposition image in the decomposition image time sequence based on the initial decomposition image; Otherwise, the local offset distance variance is obtained based on the local offset distance set, and after confirming that the local offset distance variance is greater than or equal to the preset local offset distance threshold, the extracted initial decomposition image and the target decomposition image are both confirmed as abnormal images, and the target decomposition image is used as the extracted initial decomposition image, and the process returns to the step of confirming the target decomposition image in the decomposition image time sequence based on the initial decomposition image; Summarizing the abnormal images to obtain an abnormal image set; The abnormal image set is sent to the initiator of the analysis instruction by using the result feedback unit to realize intelligent analysis of the initial video stream.

2. The method for realizing intelligent analysis of video stream based on neural network as claimed in claim 1, characterized in that: The step of obtaining a decomposed video time point set based on the decomposed image time sequence comprises: Based on the decomposed image sequence, an analysis grayscale image sequence is obtained, and the following operations are performed on each analysis grayscale image in the analysis grayscale image sequence: Acquire an analysis grayscale mean based on the analysis grayscale image, wherein the analysis grayscale mean is the mean of multiple grayscale values ​​corresponding to the analysis grayscale image; Associating the analyzed grayscale mean and image time to obtain initial analysis nodes, and summarizing the initial analysis nodes to obtain an initial analysis node set; Obtain one or more target analysis node sets using a pre-built clustering model and an initial analysis node set; Based on one or more target analysis node sets, a set of decomposed video time points is identified in the initial video stream.

3. The method for realizing intelligent analysis of video stream based on neural network as claimed in claim 2, characterized in that: The obtaining of the second decomposition frame number based on the target video stream includes: Obtaining an initial number of frames and a video duration of a target video stream, obtaining a number of images based on the initial number of frames and the video duration, obtaining an image extraction gradient set for extracting images, and extracting an image time sequence from the target video stream based on the number of images and the image extraction gradient set, wherein the image extraction gradient set includes a plurality of image extraction ratios, and the image time sequence includes a plurality of initial images; Confirming receipt of an image recognition instruction from an image recognition unit, confirming an image recognition model set based on the image recognition instruction, wherein the image recognition model set includes a plurality of image recognition models, sequentially extracting an initial image from the image sequence, and performing the following operations on the extracted initial image: Randomly extracting two image recognition models from the image recognition model set to obtain a first recognition model and a second recognition model, recognizing the extracted initial image based on the first recognition model and the second recognition model to obtain a first recognition name set and a second recognition name set, wherein the first recognition name set includes a plurality of first recognition nodes, wherein the first recognition nodes include a first recognition name and a first recognition quantity, and the second recognition name set includes a plurality of second recognition nodes, wherein the second recognition nodes include a second recognition name and a second recognition quantity; After the target identification name set is confirmed by using the first identification name set and the second identification name set, wherein the target identification name set includes a plurality of target identification nodes, and the target identification node includes a target identification name and a target identification quantity, the following operations are performed on the target identification names in the target identification name set: A local area image corresponding to a target recognition name is identified in the initial image, the local area image is labeled using the target recognition name and the target recognition number to obtain a target image with a serial number, the target images are aggregated to obtain a target image set, and the target image set is used to obtain a second decomposition frame number.

4. The method for realizing intelligent analysis of video stream based on neural network as claimed in claim 3, characterized in that: The step of using the first identification name set and the second identification name set to identify a target identification name set includes: The following operations are performed for each first identification node in the first identification name set: Determine whether there is a second identification node in the second identification name set that is the same as the first identification node; If there is no second recognition node in the second recognition name set that is the same as the first recognition node, then randomly extracting a plurality of initial image recognition models from the image recognition model set, wherein the number of the plurality of initial image recognition models is a preset number of extractions, and obtaining a plurality of initial recognition name sets using the plurality of initial image recognition models and the initial image, wherein the initial image recognition models correspond to the initial recognition name sets one by one, and the initial recognition name sets include a plurality of initial recognition nodes; Combining the multiple initial identification name sets in a combined form to obtain multiple combined identification name sets, wherein the combined identification name set includes two initial identification name sets, and identifying one or more fused identification name sets in the multiple combined identification name sets, wherein the two initial identification name sets in the fused identification name set are the same; Eliminate any one of the initial identification name sets corresponding to each fused identification name set in one or more fused identification name sets from the multiple initial identification name sets to obtain an updated identification name set, use the updated identification name set as multiple initial identification name sets, and return to the step of combining the multiple initial identification name sets in a combined form until one or more initial fused name sets are obtained, wherein the multiple initial identification name sets corresponding to the initial fused name set are all the same; Respectively counting the number of initial identification name sets corresponding to each of the one or more initial fusion name sets to obtain one or more initial fusion numbers, wherein the initial fusion numbers correspond to the initial fusion name sets one by one; Extracting the largest initial fusion number from one or more initial fusion numbers to obtain a target verification number, calculating the ratio of the target verification number to the extraction number to obtain a correct ratio, and when the correct ratio is greater than or equal to a preset ratio threshold, taking the fusion recognition name set corresponding to the target verification number as the target recognition name set; Otherwise, return to the step of randomly extracting a plurality of initial image recognition models from the image recognition model set until a target recognition name set is obtained.

5. The method for realizing intelligent analysis of video stream based on neural network as claimed in claim 4, characterized in that: The step of obtaining a second decomposition frame number by using the target image set includes: Summarizing the target image sets to obtain multiple target image sets, and obtaining the target image time sequence using the target identification name, serial number and target image; Sequentially extract initial analysis images from the target image sequence, and perform the following operations on the extracted initial analysis images: Based on the initial analysis image, identifying a target analysis image in the target image time sequence, wherein the target analysis image is adjacent to the initial analysis image and lags behind the initial analysis image in the target image time sequence; Acquire an initial coordinate point set and a target coordinate point set based on an initial analysis image and a target analysis image; Mapping the initial coordinate points in the initial coordinate point set and the target coordinate points in the target coordinate point set to a pre-constructed reference coordinate system to obtain a mapping coordinate point set, wherein the mapping coordinate point set includes a plurality of mapping coordinate points; The number of reference coordinate points corresponding to each mapping coordinate point in the mapping coordinate point set is counted to obtain a reference quantity set, wherein the reference coordinate point is an initial coordinate point or a target coordinate point, and the reference quantity set includes multiple reference quantities. The reference quantity in the reference quantity set is counted as 2 to obtain a target reference quantity. The number of mapping coordinate points in the mapping coordinate point set is counted to obtain a comprehensive quantity. The ratio of the target reference quantity to the comprehensive quantity is calculated to obtain a reference quantity ratio. If the reference quantity ratio is less than or equal to a preset reference ratio threshold, the target analysis image is taken as the initial analysis image, and the step of confirming the target analysis image in the target image time sequence based on the initial analysis image is returned. If the reference quantity ratios of the reference quantity set corresponding to the target image time sequence are all less than or equal to the reference ratio threshold, the target image is identified as a background image, otherwise, the target image is identified as an initial detection image. After confirming that the initial detection image is a preset target detection image, the target detection image is used to identify an adjacent image set in the initial image, wherein the adjacent image set includes one or more adjacent images, and the adjacent image is a target detection image or a background image. The second decomposition frame number is obtained based on the adjacent image set and the detection image.

6. The method for realizing intelligent analysis of video stream based on neural network as claimed in claim 5, characterized in that: The step of confirming that the initial detection image is a preset target detection image includes: Acquire a detection image sequence according to the target recognition name, serial number and initial detection image, wherein the detection image sequence includes multiple initial detection images, perform a grayscale operation on each initial detection image in the detection image sequence, and obtain a grayscale detection image sequence, and extract analysis image groups in sequence from the grayscale detection image sequence based on a preset sliding window, wherein the analysis image group includes two grayscale detection images; The following operations are performed for each grayscale detection image in the analysis image group: The image center coordinates are calculated based on the grayscale detection image. The calculation formula is as follows: , in, Respectively represent the horizontal and vertical coordinates of the image center coordinates, Indicates that there are a total of pixels, Indicates the grayscale detection image The gray value corresponding to each pixel point, Respectively represent the grayscale detection image The horizontal and vertical coordinates corresponding to the pixel points; Summarizing the image center coordinates to obtain an image center coordinate set, calculating a comprehensive image change degree based on the image center coordinate set, and determining whether the comprehensive image change degree is equal to a preset comprehensive change degree threshold; If the comprehensive image change degree is greater than or equal to the comprehensive change degree threshold, the initial detection image is confirmed to be a target detection image.

7. The method for realizing intelligent analysis of video stream based on neural network as claimed in claim 6, characterized in that: The calculating of the comprehensive image change degree based on the image center coordinate set includes: Based on the image center coordinate set, an image center coordinate time series is obtained, and the sliding window is used to sequentially extract analysis coordinate groups from the image center coordinate time series, wherein the analysis coordinate group includes two image center coordinates, and the Euclidean distance between the two image center coordinates in the analysis coordinate group is obtained to obtain the moving distance; Extract the first image center coordinate and the last image center coordinate in the image center coordinate sequence, obtain the initial center coordinate and the target center coordinate, and obtain the evaluation distance based on the initial center coordinate and the target center coordinate; Summarizing the moving distances to obtain a moving distance set, and using the moving distance set to obtain a moving distance variance, wherein the moving distance variance is a variance of multiple moving distances in the moving distance set; The comprehensive image change degree is calculated based on the moving distance variance and the evaluation distance. The calculation formula is as follows: , in, represents the comprehensive image change degree, are all preset coefficients. represents the evaluation distance, represents the variance of moving distance.

8. The method for realizing intelligent analysis of video stream based on neural network as claimed in claim 7, characterized in that: The step of obtaining a second decomposed frame number based on the adjacent image set and the detection image includes: Based on the adjacent image set and the detection image, a set of adjacent grayscale images and a detection grayscale image are obtained, one or more local area images are extracted from the detection grayscale image using a pre-constructed region growing algorithm and a pre-constructed image grayscale gradient set, and one or more target area images are identified in the one or more local area images, wherein the image grayscale gradient set includes a plurality of image gradient grayscale values, and the target area image is adjacent to at least one adjacent grayscale image in the adjacent grayscale image set; The following operations are performed on each of the one or more target area images: Based on the target area image, a local area grayscale mean is obtained, wherein the local area grayscale mean is the mean of multiple grayscale values ​​corresponding to the target area image, and the target area image and adjacent grayscale images adjacent to the target area image are associated to obtain an associated image set; Taking the local region grayscale mean as a starting point, extracting a related region image from the related image set using the region growing algorithm, and identifying a target discrimination image in the related region image, wherein the target discrimination image is a region of an adjacent grayscale image in the related image set that is adjacent to the target region image; The target discrimination mean is obtained based on the target discrimination image, and the discrimination frame number is calculated according to the target discrimination mean and the local area grayscale mean. The calculation formula is as follows: , , in, Indicates the number of frames to be judged. Indicates the initial frame number, represents the target discrimination mean, represents the grayscale mean of the local area, represents the preset coefficient, represents the preset coefficient, Indicates the rounding symbol; The determination frame numbers are summarized to obtain a determination frame number set, and a second decomposition frame number is determined based on the determination frame number set, wherein the second decomposition frame number is the smallest determination frame number in the determination frame number set.

9. An intelligent analysis system for video streams based on neural networks, characterized in that: The system comprises: An analysis environment confirmation module is used to receive an analysis instruction and confirm an intelligent analysis environment based on the analysis instruction, wherein the intelligent analysis environment includes an initial video stream and a video analysis system, wherein the video analysis system includes: a video decomposition unit, an image recognition unit, an image comparison unit and a result feedback unit; An initial frame number confirmation module is used to confirm receiving a video decomposition instruction from a video decomposition unit, parse the video decomposition instruction, obtain a first decomposition frame number, and obtain a decomposition image sequence using the first decomposition frame number and an initial video stream, wherein the decomposition image sequence includes a plurality of decomposition image nodes, and the decomposition image node includes a decomposition image and an image time; The initial video partitioning module is used to obtain a decomposed video time point set based on the decomposed image time sequence, wherein the decomposed video time point set includes M decomposed video time points, and M is an integer greater than or equal to 0, and obtain a target video stream set based on the decomposed video time point set and the initial video stream, wherein the target video stream set includes one or more target video streams, and perform the following operations on each target video stream in the target video stream set: An abnormal image recognition module is used to obtain a second decomposition frame number based on a target video stream, and to identify an abnormal image set based on the second decomposition frame number and the target video stream, wherein the abnormal image set includes N abnormal images, and N is an integer greater than or equal to 0. The identification of the abnormal image set based on the second decomposition frame number and the target video stream includes: A decomposed image sequence is extracted from the target video stream using the second decomposed frame number, wherein the decomposed image sequence includes a plurality of initial decomposed images, initial decomposed images are sequentially extracted from the decomposed image sequence, and the following operations are performed on the extracted initial decomposed images: Based on the initial decomposition image, a target decomposition image is identified in the decomposition image time sequence, wherein the target decomposition image is adjacent to the initial decomposition image and lags behind the extracted initial decomposition image; Confirming receiving an image comparison instruction from an image comparison unit, parsing the image comparison instruction, obtaining a local name database, and using the image recognition model set and the local name database to respectively identify a target identification decomposition image set and an initial identification decomposition image set in the target decomposition image and the initial decomposition image, wherein the target identification decomposition image set includes a plurality of target identification decomposition images, and the initial identification decomposition image set includes a plurality of initial identification decomposition images; Matching the target identification decomposition image in the target identification decomposition image set and the initial identification decomposition image in the initial identification decomposition image set to obtain a plurality of identification decomposition nodes, wherein the identification decomposition nodes include the initial identification decomposition image and the target identification decomposition image; The following operations are performed on each of the multiple identity decomposition nodes: Obtaining local offset distances based on the identified decomposition nodes, summarizing the local offset distances to obtain a local offset distance set, extracting the maximum local offset distance from the local offset distance set to obtain a target evaluation distance, and comparing the target evaluation distance with a preset evaluation distance threshold; If the target evaluation distance is greater than or equal to the evaluation distance threshold, both the extracted initial decomposition image and the target decomposition image are confirmed as abnormal images, and the target decomposition image is used as the extracted initial decomposition image, and the process returns to the step of confirming the target decomposition image in the decomposition image time sequence based on the initial decomposition image; Otherwise, the local offset distance variance is obtained based on the local offset distance set, and after confirming that the local offset distance variance is greater than or equal to the preset local offset distance threshold, the extracted initial decomposition image and the target decomposition image are both confirmed as abnormal images, and the target decomposition image is used as the extracted initial decomposition image, and the process returns to the step of confirming the target decomposition image in the decomposition image time sequence based on the initial decomposition image; Summarizing the abnormal images to obtain an abnormal image set; The abnormal image set is sent to the initiator of the analysis instruction by using the result feedback unit to realize intelligent analysis of the initial video stream.

Citation Information

Patent Citations

  • Video-based health degree monitoring method and system

    CN115209134A

  • Video frame stream processing method and system

    CN118540515A