A deep learning-based video analysis extraction method and system

By employing deep learning-based video analysis methods, combined with object detection and tracking algorithms using convolutional neural networks and Siamese networks, the accuracy and reliability issues of information extraction in video analysis are addressed, enabling efficient, accurate, and intelligent processing of video data.

CN117830909BActive Publication Date: 2026-02-06HUA RONG INFORMATION IND
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202410112855.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-26
Publication Date
2026-02-06
Estimated Expiration
2044-01-26

AI Technical Summary

Technical Problem

Existing technologies in video image processing, when extracting target location information from video analysis, are not very accurate or reliable, and cannot effectively extract key information from video data sources.

Method used

This paper employs a deep learning-based video analysis method, combining convolutional neural networks and Siamese networks for target detection and tracking algorithms. It continuously identifies the features of detected targets in video images, calculates the feature information and position information of the targets, and uses a deep learning-trained model for position comparison and information extraction.

Benefits of technology

It improves the accuracy and reliability of video information extraction, and enables efficient, accurate and intelligent processing of information such as text, charts, map points, trajectory information and number of people in videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117830909B_ABST
    Figure CN117830909B_ABST
Patent Text Reader

Abstract

The application provides a deep learning-based video analysis extraction method and system, which comprises the following steps: continuously identifying the detection target features in the preprocessed video images to be processed, and extracting the feature information of the detection targets between the continuous frame video images to be processed; labeling the targets in the continuous frame video images to be processed according to the extracted feature information of the detection targets, and obtaining the quantity statistical information of the detection targets; calculating the first actual position information of the targets in the continuous frame video images to be processed according to the first type calculation information of the targets; comparing the first actual position information with the second actual position information obtained according to the second type calculation information of the targets; if the difference between the first actual position information and the second actual position information is less than a preset difference threshold, the first actual position information and the quantity statistical information of the targets are output, and the accuracy and reliability of video information extraction are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of video image processing, and in particular to a video analysis extraction method and system based on deep learning. BACKGROUND

[0002] At present, with the popularization of digital information and the wide application of video monitoring systems, the large-scale growth of video data sources makes it challenging to effectively extract key information from videos.

[0003] Traditional video image extraction methods generally separately process aspects such as text extraction, geographic location information acquisition, and personnel quantity statistics in video images, and when extracting target location information from video images, only one calculation or extraction method is considered, resulting in low accuracy and reliability when analyzing and extracting target location information from videos.

[0004] To address this problem, the present application provides a video analysis extraction method and system based on deep learning to solve the above problems. SUMMARY

[0005] The present application proposes a video analysis extraction method and system based on deep learning to effectively solve the problem of low accuracy and reliability of video information extraction caused by existing technologies, and effectively improve the accuracy and reliability of video information extraction.

[0006] The present application provides a video analysis extraction method based on deep learning, comprising:

[0007] Obtaining a video image to be processed, and preprocessing the video image to be processed;

[0008] Sequentially passing through a target detection algorithm of a convolutional neural network and a target tracking algorithm based on a Siamese network, continuously identifying the detection target features in the preprocessed video image to be processed, and extracting the feature information of the detection target between the continuous frames of the video image to be processed;

[0009] According to the extracted feature information of the detection target, labeling the target in the continuous frames of the video image to be processed, and obtaining the quantity statistical information of the detection target between the continuous frames of the video image to be processed;

[0010] The first actual position information of the target in the to-be-processed continuous frame video image and the number statistical information are output if the difference between the first actual position information and the second actual position information is less than a preset difference threshold, the first type calculation information is a moving track of a plurality of preset points in the target in the to-be-processed continuous frame video image and a camera parameter, and the second type calculation information is other information except the first type calculation information.

[0011] Optionally, the method further comprises:

[0012] The text chart in the preprocessed to-be-processed video image is recognized, the text information and the chart information in the to-be-processed video image are extracted, and the text information and / or the chart information in the to-be-processed video image are taken as the second type calculation information.

[0013] The map point position information corresponding to the image where the target is located in the preprocessed to-be-processed video image is recognized, and the map point position information corresponding to the image where the target is located is taken as the second type calculation information.

[0014] Further, when the second type calculation information is the text information and / or the chart information in the to-be-processed video image, the second actual position information is actual position information corresponding to the target included in the text information and / or the chart information in the to-be-processed video image; and when the second type calculation information is the map point position information corresponding to the image where the target is located, the second actual position information is the map point position information corresponding to the image where the target is located.

[0015] Optionally, the second type calculation information is historical position information of the target in the to-be-processed continuous frame video image and historical pixel coordinate information of the target in the to-be-processed continuous frame video image.

[0016] Further, when the second type calculation information is the historical position information of the target in the to-be-processed continuous frame video image and the historical pixel coordinate information of the target in the to-be-processed continuous frame video image, the second position information is obtained in the following manner:

[0017] The historical position information of the target in the to-be-processed continuous frame video image and the historical pixel coordinate information of the target in the to-be-processed continuous frame video image are obtained.

[0018] The deep learning training is performed according to historical position information of the target in the continuous frames of video images to be processed and historical pixel coordinate information of the target in the continuous frames of video images to be processed, to obtain a position corresponding relationship model, which is used to represent a corresponding relationship between the position information of the target in the continuous frames of video images to be processed and the pixel coordinate information of the target in the continuous frames of video images to be processed.

[0019] The second actual position information is obtained according to actual pixel coordinate information of the target in the continuous frames of video images to be processed and the position corresponding relationship model.

[0020] Optionally, the calculation of the first actual position information of the target in the continuous frames of video images to be processed according to the first type of calculation information of the target specifically includes:

[0021] A first number of preset points in the target in the continuous frames of video images to be processed are selected;

[0022] The coordinates of a three-dimensional space point where each preset point is located in a camera coordinate system are obtained according to the displacement of each preset point in the continuous adjacent frames of video images to be processed and internal parameters of the camera.

[0023] The three-dimensional space position coordinates where the preset point is located are obtained according to the coordinates of the three-dimensional space point where the preset point is located in the camera coordinate system and external parameters of the camera.

[0024] The first actual position information of the target in the continuous frames of video images to be processed is obtained according to the three-dimensional space position coordinates where a plurality of preset points in the target are located.

[0025] Further, the coordinates of the three-dimensional space point where each preset point is located in the camera coordinate system are obtained according to the displacement of each preset point in the continuous adjacent frames of video images to be processed and the internal parameters of the camera specifically include:

[0026]

[0027]

[0028]

[0029] wherein, Z Ci is a vertical direction coordinate of the three-dimensional space point where the preset point i is located in the camera coordinate system, X Ci is a horizontal direction coordinate of the three-dimensional space point where the preset point i is located in the camera coordinate system, Y Ci is a vertical direction coordinate of the three-dimensional space point where the preset point i is located in the camera coordinate system, Δu i is a horizontal direction displacement of the preset point i in the continuous adjacent frames of video images to be processed, and Δv if is an internal parameter of the camera.

[0030] Further, u i , v i The following conditions need to be met:

[0031] I x ·u i +I y ·v i =-I t

[0032] Wherein, I x is the gradient of the preset point i in the horizontal direction of the continuous adjacent frame video image to be processed, I y is the gradient of the preset point i in the vertical direction of the continuous adjacent frame video image to be processed, u i is the displacement of the preset point i in the horizontal direction of the continuous multiple frame video image to be processed, v i is the displacement of the preset point i in the vertical direction of the continuous multiple frame video image to be processed, I t is the gradient of the preset point i in the continuous adjacent frame video image to be processed with time.

[0033] Optionally, according to the coordinates of the three-dimensional space point where the preset point is located in the camera coordinate system and the camera external parameters, the three-dimensional space position coordinates of the preset point are obtained, and the three-dimensional space position coordinates of the preset point are specifically:

[0034] Wherein, X i is the horizontal direction coordinate of the three-dimensional space position of the preset point i, Y i is the vertical direction coordinate of the three-dimensional space position of the preset point i, Z i is the vertical direction coordinate of the three-dimensional space position of the preset point i, and [R|t] represents a camera external parameter matrix, wherein the camera external parameter matrix includes a rotation matrix R and a translation vector t.

[0035] And, the value conditions of the rotation matrix R and the translation vector t are that, for the three-dimensional space point in the camera coordinate system, after the transformation through the camera external parameter matrix and the internal parameter matrix, the pixel coordinates projected onto the image plane are minimum error with the known pixel coordinates.

[0036] The second aspect of the present application provides a video analysis extraction system based on deep learning, comprising:

[0037] An acquisition and preprocessing module is configured to acquire a video image to be processed and pre-process the video image to be processed.

[0038] The recognition and extraction module sequentially passes through a target detection algorithm of a convolutional neural network and a target tracking algorithm based on a Siamese network, continuously recognizes detection target features in the preprocessed video image to be processed, and extracts feature information of the detection target between the video images of the continuous frames to be processed.

[0039] The labeling module labels the target in the video images of the continuous frames to be processed according to the extracted feature information of the detection target, and obtains quantity statistical information of the detection target between the video images of the continuous frames to be processed.

[0040] The calculation and output module calculates first actual position information of the target in the video images of the continuous frames to be processed according to first type calculation information of the target, compares the first actual position information with second actual position information obtained according to second type calculation information of the target, and if a difference range between the first actual position information and the second actual position information is less than a preset difference threshold, outputs the first actual position information of the target and the quantity statistical information, the first type calculation information is a moving track of a plurality of preset points in the target in the video images of the continuous frames to be processed and a camera parameter, and the second type calculation information is other information except the first type calculation information.

[0041] The technical scheme adopted by the present application includes the following technical effects:

[0042] 1、The recognition and extraction module sequentially passes through a target detection algorithm of a convolutional neural network and a target tracking algorithm based on a Siamese network, continuously recognizes detection target features in the preprocessed video image to be processed, and extracts feature information of the detection target between the video images of the continuous frames to be processed.

[0043] 2. In the technical solution of this invention, text and charts in the preprocessed video image to be processed are identified, and text and chart information in the video image to be processed are extracted. The text and / or chart information in the video image to be processed is used as the second type of calculation information. When the second type of calculation information is text and / or chart information in the video image to be processed, the second actual position information is the actual position information corresponding to the target, which not only enables the extraction of text and chart information in the video image, but also allows the extracted text and chart information to be used as the second actual position information of the target, thus improving the utilization rate of video extracted information.

[0044] 3. In the technical solution of the present invention, the map point information corresponding to the target image in the preprocessed video image to be processed is identified, and the map point information corresponding to the target image is used as the second type of calculation information. When the second type of calculation information is the map point information corresponding to the target image, the second actual position information is the map point information corresponding to the target image. This not only enables the extraction of map point information in the video image, but also allows the map point information corresponding to the target image to be used as the second actual position information of the target, thereby improving the utilization rate of video extracted information.

[0045] 4. The second type of calculated information in the technical solution of the present invention can also be the historical position information and historical pixel coordinate information of the target in the continuous frame video image to be processed. When the second type of calculated information is the historical position information and historical pixel coordinate information of the target in the continuous frame video image to be processed, the acquisition method of the second position information is as follows: acquire the historical position information and historical pixel coordinate information of the target in the continuous frame video image to be processed; perform deep learning training based on the historical position information and historical pixel coordinate information of the target in the continuous frame video image to be processed to obtain a position correspondence model; obtain the second actual position information based on the actual pixel coordinate information and position correspondence model of the target in the continuous frame video image to be processed. The first actual position information of the target can be determined based on the pixel coordinate information of the target in the continuous frame video image to be processed, thereby improving the accuracy and reliability of video information extraction.

[0046] 5. The technical solution of this invention is based on deep learning. By accurately extracting and calculating information such as text, chart data, map points, trajectory information, number of people, and actual location in the video, it achieves efficient, accurate, and intelligent video data processing.

[0047] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the application, as claimed. BRIEF DESCRIPTION OF DRAWINGS

[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced below. Obviously, for those skilled in the field, other drawings can also be obtained based on these drawings without any creative effort.

[0049] Figure 1 A flowchart of the method of example one in the present application scheme is shown.

[0050] Figure 2 A flowchart of step S4 in the method of example one in the present application scheme is shown.

[0051] Figure 3 Another flowchart of the method of example one in the present application scheme is shown.

[0052] Figure 4 A flowchart of the acquisition method of the second position information when the second type of calculation information is the historical position information of the target in the continuous frame video image to be processed and the historical pixel coordinate information of the target in the continuous frame video image to be processed in the method of example one in the present application scheme is shown.

[0053] Figure 5 A structural diagram of the system of example two in the present application scheme is shown. DETAILED DESCRIPTION

[0054] In order to clearly illustrate the technical features of the present application, the present application will be described in detail below through specific embodiments, and in combination with the drawings. The following disclosure provides many different embodiments or examples for implementing the different structures of the present application. In order to simplify the disclosure of the present application, the components and settings of specific examples are described below. In addition, the present application can repeatedly refer to numbers and / or letters in different examples. Such repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed. It should be noted that the components illustrated in the drawings are not necessarily drawn to scale. The present application omits the description of well-known components and processing techniques and processes to avoid unnecessary limitation of the present application.

[0055] Example one

[0056] As shown in the drawings, the present application provides a deep learning-based video analysis extraction method, which comprises: Figure 1

[0057] S1, obtaining a video image to be processed, and preprocessing the video image to be processed;​

[0058] S2, sequentially through the target detection algorithm of the convolutional neural network and the target tracking algorithm based on the Siamese network, continuously identify the detection target features in the pre-processed video images to be processed, and extract the feature information of the detection targets between the continuous frame video images to be processed;

[0059] S3, according to the extracted feature information of the detection targets, labeling the targets in the continuous frame video images to be processed, and obtaining the quantity statistical information of the detection targets between the continuous frame video images to be processed;

[0060] S4, according to the first type calculation information of the target, calculating the first actual position information of the target in the continuous frame video images to be processed; the first type calculation information is the movement trajectory of the plurality of preset points in the target in the continuous frame video images to be processed, and the camera parameter, and the second type calculation information is other information except the first type calculation information;

[0061] S5, comparing the first actual position information with the second actual position information obtained according to the second type calculation information of the target, judging whether the difference between the first actual position information and the second actual position information is less than a preset difference threshold value, if the judgment result is yes, executing step S6, if the judgment result is no, executing steps S4-S5;

[0062] S6, outputting the first actual position information of the target and the quantity statistical information.

[0063] In step S1, the video images to be processed are obtained from the video to be processed or the externally input picture file. The image preprocessing can adopt advanced image processing algorithms, including but not limited to image denoising, edge detection, image enhancement, to ensure efficient processing and accurate information extraction of the input image.

[0064] In step S2, the convolutional neural network is a pre-trained feature extraction network, responsible for extracting the features of all targets (which can be personnel or specific articles) for subsequent personnel and specific article statistics. The target detection algorithm can be an object detection algorithm or a personnel detection algorithm, such as Faster R-CNN, YOLO, etc., which accurately extracts personnel and specific articles in the video, and then continuously tracks the targets in the pre-processed video images to be processed determined by the target detection algorithm through the target tracking algorithm based on the Siamese network, to ensure continuous tracking of the targets between the continuous frames, continuously identify the detection target features in the pre-processed video images to be processed, and extract the feature information of the detection targets between the continuous frame video images to be processed.

[0065] In step S3, after obtaining all the feature information of the detection target, the target in the continuous frame video image to be processed can be labeled according to the feature information of the detection target, and the number of personnel and specific articles can be counted. In the process of counting the number of personnel and specific articles, a deep learning instance segmentation algorithm such as Mask R-CNN can be applied to label and count the target at the pixel level to obtain more detailed number counting results and obtain the number counting information of the detection target between the continuous frame video images to be processed.

[0066] In step S4, the first actual position information of the target in the continuous frame video image to be processed is calculated according to the first type calculation information in the target, as shown in the following formula: Figure 2

[0067] S41, selecting a first number of preset points in the target in the labeled continuous frame video image to be processed;

[0068] The number of preset points selected, i.e., the first number, is different according to the type of the target in the labeled continuous frame video image to be processed. Specifically, if the type of the target in the labeled continuous frame video image to be processed is a line-shaped regular object such as a rectangle, a square, a triangle, a parallelogram, a circle, etc., the value of the first number is relatively small. If the type of the target in the labeled continuous frame video image to be processed is a line-shaped irregular object such as a ring with inconsistent curvature or a person, the value of the first number is relatively large. Specifically, a corresponding relationship database between the type of the target and the number of preset points selected (the first number) can be established in advance, and once the type of the detection target is determined, the corresponding first number of preset points can be selected. In the selection process, if the type of the target is a line-shaped regular object, the selection rule can be uniform selection. If the type of the target is a line-shaped regular object, the selection rule can be uniform selection, or more preset points can be selected when the curvature change of the object shape is greater than a preset curvature change threshold, and fewer preset points can be selected when the curvature change of the object shape is not greater than the preset curvature change threshold. The selection rule can also be adjusted flexibly according to the actual situation, which is not limited in the present application.

[0069] S42, obtaining the coordinates of the three-dimensional space point where the preset point is located in the camera coordinate system according to the displacement of each preset point in the continuous adjacent frame video image to be processed and the internal parameters of the camera;

[0070] The coordinates of the three-dimensional space point (X i , Y i , Z i ) where the preset point i is located in the camera coordinate system (camera coordinate system) are specifically as follows:

[0071]

[0072]

[0073]

[0074] wherein, Z Ci is the vertical direction coordinate of the three-dimensional space point where the preset point i is located in the camera coordinate system, X Ci is the horizontal direction coordinate of the three-dimensional space point where the preset point i is located in the camera coordinate system, Y Ci is the vertical direction coordinate of the three-dimensional space point where the preset point i is located in the camera coordinate system, Δu i is the horizontal direction displacement of the preset point i in the continuous adjacent frame video images to be processed (the horizontal direction pixel change value of the preset point i in the continuous adjacent frame video images to be processed), Δv i is the vertical direction displacement of the preset point i in the continuous adjacent frame video images to be processed (the vertical direction pixel change value of the preset point i in the continuous adjacent frame video images to be processed), and f is the camera internal parameter (camera intrinsic matrix, that is, K, the camera internal parameter including the camera focal length).

[0075] wherein, u i , v i need to meet the following conditions:

[0076] I x ·u i +I y ·v i =-I t

[0077] wherein, I x is the gradient of the preset point i in the horizontal direction of the continuous adjacent frame video images to be processed, I y is the gradient of the preset point i in the vertical direction of the continuous adjacent frame video images to be processed, u i is the displacement of the preset point i in the horizontal direction of the continuous multiple frame video images to be processed, v i is the displacement of the preset point i in the vertical direction of the continuous multiple frame video images to be processed, I t is the gradient of the preset point i in the continuous adjacent frame video images to be processed changing with time.

[0078] By extracting the trajectory information in the video through the target tracking algorithm, not only the displacement of the preset point i in the horizontal direction and the vertical direction of the continuous multiple frame video images to be processed can be obtained, but also the trajectory analysis of the target can be performed, the motion mode, the key path or the abnormal trajectory can be identified, and the information value can be improved.

[0079] S43, obtaining the three-dimensional space position coordinate of the preset point according to the coordinate of the three-dimensional space point where the preset point is located in the camera coordinate system and the camera external parameter.

[0080] The specific three-dimensional spatial coordinates of the preset point are as follows:

[0081] Among them, X i Let Y be the horizontal coordinate of the three-dimensional spatial position of the preset point i. i Z represents the vertical coordinates of the three-dimensional spatial position of the preset point i. i Let [R|t] be the vertical coordinate of the three-dimensional spatial position of the preset point i, and let [R|t] represent the camera extrinsic parameter matrix, which includes the rotation matrix R and the translation vector t.

[0082] Furthermore, the rotation matrix R and translation vector t are determined such that the pixel coordinates projected onto the image plane after transformation by the camera's extrinsic and intrinsic parameters in the 3D space point in the camera coordinate system have the smallest error compared to the known pixel coordinates.

[0083] Specifically, the optimal rotation matrix R and translation vector t (which together constitute the camera extrinsic parameter matrix, or simply camera extrinsic matrix) ensure that for each point in three-dimensional space (i.e., X...)... i Y i Z i Projection P in the camera coordinate system i (P i , that is, X Ci Y Ci Z Ci After transformation by the extrinsic parameter matrix (rotation matrix R and translation vector t) and the intrinsic parameter matrix K, the pixel coordinates projected onto the image plane are (project(K[R|t]P)). i )) and known pixel coordinates (p i The error between ) is minimized, that is:

[0084]

[0085] Where project(·) represents the projection function, which projects 3D points in the camera coordinate system or camcorder coordinate system onto 2D pixel coordinates on the image plane; [R|t] represents the extrinsic parameter matrix, consisting of the rotation matrix R and the translation vector t; and K represents the intrinsic parameter matrix of the camera or camcorder. i and p i These represent the projected pixel coordinates of a point in three-dimensional space and a preset target point in the image, respectively.

[0086] S44: Based on the three-dimensional spatial coordinates of multiple preset points in the target, obtain the first actual position information of the target in the continuous frame video images to be processed.

[0087] The three-dimensional space position coordinates of the plurality of preset points in the target are combined and transformed to obtain first actual position information of the target in the continuous frame video image to be processed. For example, the mean value of the three-dimensional space position coordinates of all preset points in the target can be used as the three-dimensional space position coordinates of the target in the continuous frame video image to be processed. Then, the three-dimensional space position coordinates of the target are converted into latitude and longitude information by using a geographic information system (GIS), i.e., the first actual position information.

[0088] In step S5, the first actual position information is compared with second actual position information obtained according to the second type of calculation information of the target, and it is judged whether the difference between the first actual position information and the second actual position information is less than a preset difference threshold. The preset difference threshold can be flexibly adjusted according to the actual situation, and the present application does not limit this.

[0089] In step S6, the number of targets in the video, such as the number of people, can be detected by using a deep learning model. For example, a people flow calculation algorithm can be used to analyze the density and flow trend of people, thereby providing more information for the management of the place.

[0090] As shown in Figure 3 The present application also provides a video analysis and extraction method based on deep learning, which further comprises:

[0091] S7, recognizing the text chart in the preprocessed video image to be processed, extracting the text information and chart information in the video image to be processed, and taking the text information and / or chart information in the video image to be processed as the second type of calculation information; or,

[0092] S8, recognizing the map point information corresponding to the image where the target is located in the preprocessed video image to be processed, and taking the map point information corresponding to the image where the target is located as the second type of calculation information.

[0093] The deep learning model can be used to recognize the text information in the video. For example, a convolutional neural network (CNN) based text detection algorithm, a recurrent neural network (RNN) based text recognition algorithm, an algorithm for analyzing and extracting chart data (such as PaddleOCR, a super-lightweight OCR model library), etc. can be used to accurately recognize the text chart in the image and extract information. Natural language processing (NLP) technology is used for text analysis to extract keywords, topics or emotional information from the text information. Deep learning algorithms can be used to analyze the structure of the chart in the video and extract data. Data cleaning and statistical methods are used to ensure the accuracy of the extracted chart data.

[0094] Deep learning technology is used to detect the map point information in the video, and geographic information system (GIS) data is integrated to ensure the accuracy and integrity of the map point information.

[0095] When the second type of calculation information is text information and / or chart information in the to-be-processed video image, the second actual position information is actual position information corresponding to the text information and / or chart information in the to-be-processed video image in which the target is located; for example, whether the text information and / or chart information in the to-be-processed video image includes a cell name, a shopping mall name, a geographical position, and the like in which the target is located, if yes, actual position information (longitude and latitude information) corresponding to the cell name, the shopping mall name, and the geographical position coordinate in which the target is located in the text information and / or chart information is taken as the second actual position information.

[0096] When the second type of calculation information is map point information corresponding to an image in which the target is located, the second actual position information is the map point information corresponding to the image in which the target is located. Whether the to-be-processed video image includes map point information, if yes, the map point information (longitude and latitude information) is taken as the second actual position information, that is, when the second type of calculation information is the map point information corresponding to the image in which the target is located, the second actual position information is the second type of calculation information.

[0097] Further, if the text information and / or chart information in the to-be-processed video image does not include a cell name, a shopping mall name, and a geographical position in which the target is located, and the to-be-processed video image does not include map point information, the second type of calculation information can also be historical position information of the target in the to-be-processed continuous frame video image and historical pixel coordinate information of the target in the to-be-processed continuous frame video image. As shown in FIG. 6, when the second type of calculation information is the historical position information of the target in the to-be-processed continuous frame video image and the historical pixel coordinate information of the target in the to-be-processed continuous frame video image, the second position information is obtained in the following manner: Figure 4

[0098] The historical position information (longitude and latitude information) of the target in the to-be-processed continuous frame video image and the historical pixel coordinate information of the target in the to-be-processed continuous frame video image are obtained.

[0099] Deep learning training is performed according to the historical position information of the target in the to-be-processed continuous frame video image and the historical pixel coordinate information of the target in the to-be-processed continuous frame video image, to obtain a position corresponding relationship model, which is used to represent a corresponding relationship between the position information of the target in the to-be-processed continuous frame video image and the pixel coordinate information of the target in the to-be-processed continuous frame video image.

[0100] The second actual position information is obtained according to the actual pixel coordinate information of the target in the to-be-processed continuous frame video image and the position corresponding relationship model.

[0101] ​The position correspondence model is acquired through a deep learning training technique, so as to ensure accurate matching of coordinate information (pixel coordinates) and an actual scene (second actual position information).

[0102] It should be noted that the application field of the technical solution in the first embodiment of the present application can be:

[0103] Monitoring and security: used in intelligent monitoring systems to quickly identify and respond to abnormal events.

[0104] Urban planning and traffic management: by analyzing the number of people, trajectory information, etc., to optimize urban planning and traffic flow management;

[0105] Enterprise decision support: provides in-depth analysis of chart data to provide strong support for enterprise decision-making;

[0106] Emergency rescue: through map points and personnel quantity information, to improve the response speed and effect of emergency rescue.

[0107] The application object of the technical solution in the first embodiment of the present application can be:

[0108] Security company: used to design intelligent monitoring systems;

[0109] Urban planning department: to optimize urban planning and traffic flow management;

[0110] Enterprise management: to support decision-making and strategic planning;

[0111] Rescue agency: to improve the response capability of emergency rescue.

[0112] The technical solution in the first embodiment of the present application combines deep learning with secondary processing of comparative analysis between information, to improve the efficiency and accuracy of information extraction and processing; uses deep learning to realize intelligent classification of multiple information, to improve application applicability; through secondary processing, to conduct multi-dimensional in-depth analysis, to provide more comprehensive information insight; can be applicable to multiple fields, to meet the needs of different industries for video data.

[0113] The application sequentially passes through a target detection algorithm of a convolutional neural network and a target tracking algorithm based on a Siamese network, continuously identifies detection target features in preprocessed video images to be processed, extracts feature information of the detection target between continuous frame video images to be processed, labels the target in the continuous frame video images to be processed according to the extracted feature information of the detection target, obtains quantity statistical information of the detection target between the continuous frame video images to be processed, calculates first actual position information of the target in the continuous frame video images to be processed according to first type calculation information of the target, compares the first actual position information with second actual position information obtained according to second type calculation information of the target, and outputs the first actual position information of the target and the quantity statistical information if a difference range between the first actual position information and the second actual position information is less than a preset difference threshold, effectively solving the problem of low accuracy and reliability of video information extraction caused by the prior art, and effectively improving the accuracy and reliability of video information extraction.

[0114] In the technical scheme of the application, text and charts in the preprocessed video images to be processed are identified, text information and chart information in the video images to be processed are extracted, the text information and / or the chart information in the video images to be processed are taken as second type calculation information, and when the second type calculation information is the text information and / or the chart information in the video images to be processed, the second actual position information is actual position information corresponding to the text information and / or the chart information in the video images to be processed, so that not only the text and chart information in the video images can be extracted, but also the extracted text and chart information can be taken as the second actual position information of the target, thereby improving the utilization rate of video extraction information.

[0115] In the technical scheme of the application, map point position information corresponding to an image where a target is located in preprocessed video images to be processed is identified, the map point position information corresponding to the image where the target is located is taken as second type calculation information, and when the second type calculation information is the map point position information corresponding to the image where the target is located, the second actual position information is the map point position information corresponding to the image where the target is located, so that not only the map point position information in the video images can be extracted, but also the map point position information corresponding to the image where the target is located can be taken as the second actual position information of the target, thereby improving the utilization rate of video extraction information.

[0116] The second type of calculation information in the technical scheme of the present application can also be historical position information of the target in the to-be-processed continuous frame video image and historical pixel coordinate information of the target in the to-be-processed continuous frame video image. When the second type of calculation information is the historical position information of the target in the to-be-processed continuous frame video image and the historical pixel coordinate information of the target in the to-be-processed continuous frame video image, the manner of obtaining the second position information is specifically: obtaining the historical position information of the target in the to-be-processed continuous frame video image and the historical pixel coordinate information of the target in the to-be-processed continuous frame video image; performing deep learning training according to the historical position information of the target in the to-be-processed continuous frame video image and the historical pixel coordinate information of the target in the to-be-processed continuous frame video image to obtain a position correspondence relationship model; and obtaining the second actual position information according to the actual pixel coordinate information of the target in the to-be-processed continuous frame video image and the position correspondence relationship model. The first actual position information of the target can be determined according to the pixel coordinate information of the target in the to-be-processed continuous frame video image, and the accuracy and reliability of video information extraction are improved.

[0117] In the technical scheme of the present application, based on deep learning, accurate extraction and calculation of information such as text, chart data, map points, trajectory information, number of personnel, actual position and the like in a video are realized, and efficient, accurate and intelligent video data processing is realized.

[0118] Embodiment two

[0119] As shown in Figure 5 The technical scheme of the present application also provides a video analysis and extraction system based on deep learning, which comprises:

[0120] The acquisition and preprocessing module 101 acquires a to-be-processed video image and pre-processes the to-be-processed video image.

[0121] The recognition and extraction module 102 successively recognizes detection target features in the pre-processed to-be-processed video image through a target detection algorithm of a convolutional neural network and a target tracking algorithm based on a Siamese network, and extracts feature information of the detection target between to-be-processed continuous frame video images.

[0122] The labeling module 103 labels the target in the to-be-processed continuous frame video image according to the extracted feature information of the detection target, and acquires number statistical information of the detection target between the to-be-processed continuous frame video images.

[0123] The computing and outputting module 104 calculates first actual position information of the target in the to-be-processed continuous frame video image according to first type calculation information of the target; compares the first actual position information with second actual position information obtained according to second type calculation information of the target; and if a difference range between the first actual position information and the second actual position information is less than a preset difference threshold, outputs the first actual position information of the target and the quantity statistical information, the first type calculation information being movement trajectories of a plurality of preset points in the target in the to-be-processed continuous frame video image and camera parameters, and the second type calculation information being other information except the first type calculation information.

[0124] The application successively identifies the detection target features in the preprocessed to-be-processed video image through the target detection algorithm of the convolutional neural network and the target tracking algorithm based on the Siamese network, extracts the feature information of the detection target between the to-be-processed continuous frame video images, labels the target in the to-be-processed continuous frame video image according to the extracted feature information of the detection target, and obtains the quantity statistical information of the detection target between the to-be-processed continuous frame video images; calculates first actual position information of the target in the to-be-processed continuous frame video image according to first type calculation information of the target; compares the first actual position information with second actual position information obtained according to second type calculation information of the target; and if a difference range between the first actual position information and the second actual position information is less than a preset difference threshold, outputs the first actual position information of the target and the quantity statistical information, effectively solves the problem of low accuracy and reliability of video information extraction caused by the prior art, and effectively improves the accuracy and reliability of video information extraction.

[0125] In the technical scheme of the application, the text and chart in the preprocessed to-be-processed video image are identified, the text information and chart information in the to-be-processed video image are extracted, and the text information and / or chart information in the to-be-processed video image is taken as second type calculation information; when the second type calculation information is the text information and / or chart information in the to-be-processed video image, the second actual position information is actual position information corresponding to the target in the text information and / or chart information in the to-be-processed video image, which not only realizes extraction of the text and chart information in the video image, but also improves the utilization rate of video extraction information by taking the extracted text and chart information as the second actual position information of the target.

[0126] In the technical scheme of the present application, the map point position information corresponding to the image where the target is located in the pre-processed video image to be processed is identified, and the map point position information corresponding to the image where the target is located is taken as the second type of calculation information, when the second type of calculation information is the map point position information corresponding to the image where the target is located, the second actual position information is the map point position information corresponding to the image where the target is located, which not only realizes the extraction of the map point position information in the video image, but also takes the map point position information corresponding to the image where the target is located as the second actual position information of the target, thereby improving the utilization rate of video extraction information.

[0127] In the technical scheme of the present application, the second type of calculation information can also be the historical position information of the target in the continuous frame video image to be processed and the historical pixel coordinate information of the target in the continuous frame video image to be processed, when the second type of calculation information is the historical position information of the target in the continuous frame video image to be processed and the historical pixel coordinate information of the target in the continuous frame video image to be processed, the acquisition mode of the second position information is specifically: obtaining the historical position information of the target in the continuous frame video image to be processed and the historical pixel coordinate information of the target in the continuous frame video image to be processed; performing deep learning training according to the historical position information of the target in the continuous frame video image to be processed and the historical pixel coordinate information of the target in the continuous frame video image to be processed, to obtain a position correspondence model; obtaining the second actual position information according to the actual pixel coordinate information of the target in the continuous frame video image to be processed and the position correspondence model, which can determine the first actual position information of the target according to the pixel coordinate information of the target in the continuous frame video image to be processed, thereby improving the accuracy and reliability of video information extraction.

[0128] In the technical scheme of the present application, based on deep learning, accurate extraction and calculation of information such as text, chart data, map point position, trajectory information, number of personnel, actual position and the like in the video are realized, thereby realizing efficient, accurate and intelligent video data processing.

[0129] Although the specific embodiments of the present application have been described above with reference to the accompanying drawings, the present application is not limited to the above-described embodiments, and various modifications or changes can be made by those skilled in the art without creative labor, which are still within the scope of the present application.

Claims

1. A deep learning-based video analysis extraction method, characterized in that, The method comprises the following steps: acquiring a to-be-processed video image, and preprocessing the to-be-processed video image; sequentially passing through a target detection algorithm of a convolutional neural network and a target tracking algorithm based on a Siamese network, continuously identifying detection target features in the preprocessed to-be-processed video image, and extracting feature information of the detection target between to-be-processed continuous frame video images; annotating the target in the to-be-processed continuous frame video image according to the extracted feature information of the detection target, and acquiring quantity statistical information of the detection target between the to-be-processed continuous frame video images; calculating first actual position information of the target in the to-be-processed continuous frame video image according to first type calculation information of the target; comparing the first actual position information with second actual position information obtained according to second type calculation information of the target, and if a difference range between the first actual position information and the second actual position information is less than a preset difference threshold, outputting the first actual position information and the quantity statistical information of the target, the first type calculation information being movement trajectories of a plurality of preset points in the target in the to-be-processed continuous frame video image and camera parameters, and the second type calculation information being other information except the first type calculation information; if the difference range between the first actual position information and the second actual position information is not less than the preset difference threshold, recalculating the first actual position information of the target in the to-be-processed continuous frame video image according to the first type calculation information of the target until the difference range between the first actual position information and the second actual position information is less than the preset difference threshold. 2.The method of claim 1, further comprising The method comprises the following steps: identifying text and charts in the preprocessed to-be-processed video image, extracting text information and chart information in the to-be-processed video image, and taking the text information and / or the chart information in the to-be-processed video image as the second type calculation information; or identifying map point position information corresponding to an image where the target is located in the preprocessed to-be-processed video image, and taking the map point position information corresponding to the image where the target is located as the second type calculation information. 3.The method of claim 2, wherein, When the second type calculation information is the text information and / or the chart information in the to-be-processed video image, the second actual position information is actual position information included in the text information and / or the chart information in the to-be-processed video image; When the second type calculation information is the map point position information corresponding to the image where the target is located, the second actual position information is the map point position information corresponding to the image where the target is located.

4. The method of claim 1, wherein the method further comprises: The second type calculation information is historical position information of the target in the to-be-processed continuous frame video image and historical pixel coordinate information of the target in the to-be-processed continuous frame video image.

5. The method of claim 4, wherein the method further comprises: When the second type calculation information is the historical position information of the target in the to-be-processed continuous frame video image and the historical pixel coordinate information of the target in the to-be-processed continuous frame video image, the second position information is acquired in the following manner: acquiring the historical position information of the target in the to-be-processed continuous frame video image and the historical pixel coordinate information of the target in the to-be-processed continuous frame video image; The deep learning training is performed according to historical position information of the target in the continuous frames of video images to be processed and historical pixel coordinate information of the target in the continuous frames of video images to be processed, so as to obtain a position corresponding relationship model, which is used to represent a corresponding relationship between the position information of the target in the continuous frames of video images to be processed and the pixel coordinate information of the target in the continuous frames of video images to be processed; The second actual position information is obtained according to the actual pixel coordinate information of the target in the continuous frames of video images to be processed and the position corresponding relationship model.

6. The method of claim 1, wherein the method further comprises: The first actual position information of the target in the continuous frames of video images to be processed is calculated according to the first type calculation information of the target, and the calculation specifically includes: A first number of preset points in the target in the continuous frames of video images to be processed are selected; A coordinate of a three-dimensional space point where each preset point is located in a camera coordinate system is obtained according to a displacement of the preset point in the continuous adjacent frames of video images to be processed and an internal parameter of the camera; A three-dimensional space position coordinate where the preset point is located is obtained according to the coordinate of the three-dimensional space point where the preset point is located in the camera coordinate system and an external parameter of the camera. The first actual position information of the target in the continuous frames of video images to be processed is obtained according to the three-dimensional space position coordinates where the plurality of preset points in the target are located.

7. The method of claim 6, wherein the method further comprises: The coordinate of the three-dimensional space point where each preset point is located in the camera coordinate system specifically includes: , , , wherein, is the vertical direction coordinate of the three-dimensional space point where the preset point i is located in the camera coordinate system, is the horizontal direction coordinate of the three-dimensional space point where the preset point i is located in the camera coordinate system, is the vertical direction coordinate of the three-dimensional space point where the preset point i is located in the camera coordinate system, is the horizontal direction displacement of the preset point i in the continuous adjacent frame video images to be processed, is the vertical direction displacement of the preset point i in the continuous adjacent frame video images to be processed, and f is the camera internal parameter. 8.The method of claim 7, wherein the method further comprises, , The following conditions need to be met: wherein, is a gradient of the preset point i in a horizontal direction of the continuous adjacent frame video images to be processed, is a gradient of the preset point i in a vertical direction of the continuous adjacent frame video images to be processed, is a displacement of the preset point i in a horizontal direction of the continuous multiple frame video images to be processed, is a displacement of the preset point i in a vertical direction of the continuous multiple frame video images to be processed, is a gradient of the preset point i changing over time of the continuous adjacent frame video images to be processed.

9. The video analysis and extraction method based on deep learning according to claim 6, characterized in that, The three-dimensional space position coordinate where the preset point is located specifically includes: wherein, is a horizontal direction coordinate of a three-dimensional space position of the preset point i, is a vertical direction coordinate of a three-dimensional space position of the preset point i, is a vertical direction coordinate of a three-dimensional space position of the preset point i, and [R|t] represents a camera exterior parameter matrix, the camera exterior parameter matrix including a rotation matrix R and a translation vector t. In addition, the values of the rotation matrix R and the translation vector t satisfy a condition that, for a three-dimensional space point in the camera coordinate system, after being transformed by the camera external parameter matrix and the internal parameter matrix, the pixel coordinate of the three-dimensional space point projected onto an image plane is closest to the known pixel coordinate. 10.A deep learning based video analysis extraction system, characterized in that, The method comprises: An acquisition and preprocessing module is configured to acquire video images to be processed and pre-process the video images to be processed; A recognition and extraction module is configured to sequentially identify detection target features in the pre-processed video images to be processed by using a target detection algorithm of a convolutional neural network and a target tracking algorithm based on a Siamese network, and extract feature information of the detection target between the continuous frames of video images to be processed; A labeling module is configured to label the target in the continuous frames of video images to be processed according to the extracted feature information of the detection target, and acquire quantity statistical information of the detection target between the continuous frames of video images to be processed; A calculation and output module is configured to calculate first actual position information of the target in the continuous frames of video images to be processed according to first type calculation information of the target in the continuous frames of video images to be processed. comparing the first actual position information with second actual position information obtained according to second type calculation information of the target, if the difference between the first actual position information and the second actual position information is less than a preset difference threshold, outputting the first actual position information of the target and the quantity statistical information, the first type calculation information being movement tracks of a plurality of preset points in the target in the continuous frame video image to be processed and camera parameters, and the second type calculation information being other information except the first type calculation information; if the difference between the first actual position information and the second actual position information is not less than the preset difference threshold, recalculating the first actual position information of the target in the continuous frame video image to be processed according to the first type calculation information of the target, until the difference between the first actual position information and the second actual position information is less than the preset difference threshold.

Citation Information

Patent Citations

  • Traffic information counting method based on unmanned aerial vehicle aerial photographing video and system

    CN108320510A

  • Augmented reality positioning method and device based on environment visual feature point recognition technology

    CN110533719A

  • Target detection method and device, target tracking method and device, visual sensor and medium

    CN114170499A

  • Data processing method and device, electronic equipment and storage medium

    CN114627519A

  • Unmanned aerial vehicle video stitching method and device with position information, equipment and medium

    CN117201708A