A method and device for acquiring video data of a monitoring terminal

By realizing the video data acquisition method in the monitoring terminal, and using the online monitoring and identification model for real-time identification and upgrading, the problem of reduced recognition accuracy in the monitoring terminal is solved, and the accuracy and adaptability of intelligent identification of video data is improved.

CN119851188BActive Publication Date: 2025-07-01ANHUI TRAFFIC CONTROL INFORMATION IND CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510330092.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-01
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

The monitoring and identification model loaded in the monitoring terminal has a problem of decreasing recognition accuracy, especially when background changes frequently. If the model is not updated in time, it will lead to a decrease in recognition accuracy.

Method used

A method for obtaining video data from a monitoring terminal is proposed. By obtaining the initial video data, it determines whether there is an identification target. If it does not exist, the online monitoring and identification model is used to identify the target picture, and determine whether the model upgrade conditions are met based on the recognition results. If it is met, the model is upgraded, and the upgraded model is obtained for intelligent identification.

Benefits of technology

By updating the monitoring and identification model in real time, the accuracy of intelligent identification of video data can be effectively improved, adapt to background changes, and the recognition ability of monitoring terminals can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851188B_ABST
    Figure CN119851188B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for obtaining video data of a monitoring terminal, which relates to the technical field of video data processing. The method includes: when it is determined that there is no recognition target in the initial video data, obtaining a target picture in the initial video data, recognizing the target picture based on an online monitoring recognition model, and obtaining the recognition result of the online monitoring recognition model; if the recognition result meets the preset model upgrade condition, performing an upgrade operation on the online monitoring recognition model according to the target picture to obtain an upgraded online monitoring recognition model. The monitoring terminal in the present invention closely cooperates with the online monitoring recognition model, can realize real-time intelligent recognition of monitoring data, and the online monitoring recognition model can upgrade and update itself according to the recognition situation of the current video, so that the online monitoring recognition model can well recognize the latest monitoring background picture, thereby improving the accuracy of intelligent recognition of video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video data processing, and in particular, to a method and device for acquiring video data of a monitoring terminal. Background Art

[0002] An intelligent monitoring terminal obtained by combining a monitoring terminal with a monitoring recognition model can achieve intelligent monitoring. For example, in road monitoring, intelligent recognition of target objects (vehicles or pedestrians), etc., and even intelligent reporting of monitoring events. Different from existing intelligent monitoring systems, the intelligent monitoring terminal itself has intelligence, and monitoring data can be intelligently processed inside the monitoring terminal, and then the data sent carries the results of intelligent recognition, which improves the efficiency of intelligent monitoring. Therefore, the intelligent monitoring terminal can improve the efficiency of intelligent monitoring in road monitoring or Internet of Things monitoring.

[0003] However, loading a monitoring recognition model into a monitoring terminal will also increase the cost of the monitoring terminal. Therefore, the volume (memory occupation and calculation amount) of the monitoring recognition model loaded into the monitoring terminal is generally not large, because the position of the monitoring terminal is fixed or the usage scenario is fixed. For example, a monitoring system fixed on a road gantry, a home intelligent monitor fixed at a certain position in the home, etc. At this time, as long as the monitoring system can recognize the background of the monitoring screen, the recognition target can be quickly recognized from the background. In this method, the monitoring recognition model algorithm is very simple, its volume is also small, and correspondingly, the intelligent recognition efficiency is higher. On this basis, if more accurate or customized intelligent recognition results are to be obtained, the video data can be further processed.

[0004] In the above background, the monitoring terminal should have a sensitive perception ability for its surrounding environment. Therefore, it is required that the monitoring recognition model can accurately recognize the background of the monitoring environment. However, the background often changes over time. If the monitoring recognition model is not updated in time, the recognition accuracy will be reduced. Therefore, a new method for acquiring video data of a monitoring terminal needs to be proposed to improve the accuracy of intelligent recognition of video data. Summary of the Invention

[0005] The present invention provides a method and device for acquiring video data of a monitoring terminal, which are used to improve the accuracy of intelligent recognition of video data.

[0006] To solve the above technical problems, in the first aspect of the present invention, a method for acquiring video data of a monitoring terminal is disclosed, and the method includes:

[0007] Acquire the initial video data captured by the monitoring terminal, and based on the initial video data, determine whether there is a recognition target in the initial video data;

[0008] When it is determined that the recognition target does not exist in the initial video data, obtain the target picture in the initial video data, identify the target picture based on the online monitoring recognition model, and obtain the recognition result of the online monitoring recognition model;

[0009] Judge whether the recognition result meets the preset model upgrade condition. If the recognition result meets the preset model upgrade condition, perform an upgrade operation on the online monitoring recognition model according to the target picture to obtain an upgraded online monitoring recognition model;

[0010] Perform an intelligent recognition operation on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain target video data.

[0011] As an optional implementation manner, in the first aspect of the present invention, the identifying the target picture based on the online monitoring recognition model and obtaining the recognition result of the online monitoring recognition model includes:

[0012] Divide the target picture into multiple sub-regions according to a preset division strategy;

[0013] Insert a preset pattern to be recognized into each sub-region. For each sub-region, determine the pixel region within a preset range around the edge position of the pattern to be recognized as a transition region, and perform a Gaussian filtering operation on the transition region to obtain a target processed picture;

[0014] Identify the target processed picture based on the online monitoring recognition model, and obtain the pattern recognition result of the online monitoring recognition model, where the pattern recognition result includes the pattern recognized by the online monitoring recognition model from the target processed picture;

[0015] And, the judging whether the recognition result meets the preset model upgrade condition includes:

[0016] Compare the pattern recognition result with the pattern to be recognized, and calculate the recognition accuracy rate of the pattern recognition result;

[0017] If the accuracy rate of the pattern recognition result is lower than the preset accuracy rate threshold, it is determined that the recognition result meets the preset model upgrade condition.

[0018] As an optional implementation manner, in the first aspect of the present invention, the dividing the target picture into multiple sub-regions according to a preset division strategy includes:

[0019] According to the lens parameters corresponding to the monitoring terminal and the target picture, determine the distances from a preset number of reference points in the target picture to the monitoring terminal;

[0020] Divide all reference points into multiple subsets of reference points. Among them, within each subset of reference points, the distances from all the sub-reference points to the monitoring terminal are within the preset distance range corresponding to this subset of reference points;

[0021] Divide the target picture into multiple sub-regions according to all the subsets of reference points, where each subset of reference points corresponds to a sub-region.

[0022] As an optional implementation manner, in the first aspect of the present invention, the performing a Gaussian filtering operation on the transition region to obtain a target processed picture includes:

[0023] Determine Gaussian filtering algorithm parameters according to the preset distance range corresponding to each sub-region;

[0024] Perform a Gaussian filtering operation on the transition region according to the Gaussian filtering algorithm parameters corresponding to the transition region to obtain a target processed picture.

[0025] As an optional implementation manner, in the first aspect of the present invention, the performing an intelligent recognition operation on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain target video data includes:

[0026] Perform an intelligent recognition operation on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain an intelligent recognition result;

[0027] Based on the intelligent recognition result, extract target characters from a pre-loaded static font library, where the target characters are used to represent the intelligent recognition result through a preset expression manner;

[0028] Overlay the target characters onto the video stream corresponding to the initial video data through a rendering operation to obtain target video data.

[0029] As an optional implementation manner, in the first aspect of the present invention, the method further includes:

[0030] Pull a target video stream corresponding to the target video data from the monitoring terminal, and decode the target video stream according to a preset decoding scheme to obtain a decoded target video stream;

[0031] According to the decoded target video stream, obtain multiple derivative video streams corresponding to the target video stream according to a preset data multiplexing scheme, and perform re-encoding operations on each derivative video stream according to the encoding scheme corresponding to the preset decoding scheme to obtain multiple derivative video data.

[0032] As an alternative implementation, in the first aspect of the present invention, performing intelligent recognition operations on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain target video data includes:

[0033] Performing intelligent recognition operations on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain an intelligent recognition result;

[0034] Based on the intelligent recognition result, extracting the derivative target characters corresponding to each path of derivative video data from a pre-loaded static font library, where the derivative target characters correspond to the resolution and bit rate of the corresponding derivative video data, and the derivative target characters are used to represent the intelligent recognition result through a preset expression method;

[0035] For each path of the derivative video data, superimposing the derivative target characters onto the video stream corresponding to the derivative video data through a rendering operation to obtain target derivative video data.

[0036] As an alternative implementation, in the first aspect of the present invention, the method further includes:

[0037] Obtaining the original audio data recorded when the monitoring terminal captures the initial video data, extracting the audio features in the original audio data and generating an audio feature set;

[0038] Judging whether there are matching audio features corresponding to a pre-determined recognition target in the audio feature set;

[0039] And judging whether there is a recognition target in the initial video data according to the initial video data, including:

[0040] Based on the online monitoring recognition model to recognize the initial video data, if the online monitoring recognition model recognizes a recognition target from the initial video data, or judges that there is the matching audio feature in the audio feature set, it is determined that there is a recognition target in the initial video data.

[0041] The second aspect of the present invention discloses a video data acquisition device for a monitoring terminal, and the device includes:

[0042] A target recognition module, configured to obtain the initial video data captured by the monitoring terminal, and judge whether there is a recognition target in the initial video data according to the initial video data;

[0043] A model testing module, configured to, when it is determined that the recognition target does not exist in the initial video data, obtain a target picture in the initial video data, recognize the target picture based on an online monitoring recognition model, and obtain the recognition result of the online monitoring recognition model;

[0044] A model upgrade module, configured to determine whether the recognition result meets a preset model upgrade condition. If the recognition result meets the preset model upgrade condition, perform an upgrade operation on the online monitoring recognition model according to the target picture to obtain an upgraded online monitoring recognition model;

[0045] An intelligent recognition module, configured to perform an intelligent recognition operation on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain target video data.

[0046] As an optional implementation manner, in the second aspect of the present invention, the specific operation manner for the model testing module to recognize the target picture based on the online monitoring recognition model and obtain the recognition result of the online monitoring recognition model includes:

[0047] Divide the target picture into multiple sub-regions according to a preset division strategy;

[0048] Insert a preset pattern to be recognized into each sub-region. For each sub-region, determine the pixel region within a preset range around the edge position of the pattern to be recognized as a transition region, and perform a Gaussian filtering operation on the transition region to obtain a target processed picture;

[0049] Recognize the target processed picture based on the online monitoring recognition model to obtain the pattern recognition result of the online monitoring recognition model, where the pattern recognition result includes the pattern recognized by the online monitoring recognition model from the target processed picture;

[0050] And, the specific operation manner for the model upgrade module to determine whether the recognition result meets the preset model upgrade condition includes:

[0051] Compare the pattern recognition result with the pattern to be recognized, and calculate the recognition accuracy rate of the pattern recognition result;

[0052] If the accuracy rate of the pattern recognition result is lower than a preset accuracy rate threshold, determine that the recognition result meets the preset model upgrade condition.

[0053] As an optional implementation manner, in the second aspect of the present invention, the specific operation manner for the model testing module to divide the target picture into multiple sub-regions according to a preset division strategy includes:

[0054] Determine the distances from a plurality of preset reference points in the target picture to the monitoring terminal according to the lens parameters corresponding to the monitoring terminal and the target picture;

[0055] Divide all the reference points into a plurality of sub-reference point sets. Among them, within each sub-reference point set, the distances from all the sub-reference points to the monitoring terminal are within the preset distance range corresponding to the sub-reference point set;

[0056] Divide the target picture into a plurality of sub-regions according to all the sub-reference point sets, where each sub-reference point set corresponds to a sub-region.

[0057] As an optional implementation manner, in the second aspect of the present invention, the specific operation manner of the model testing module to perform Gaussian filtering operation on the transition region to obtain the target processed picture includes:

[0058] Determine the Gaussian filtering algorithm parameters according to the preset distance range corresponding to each sub-region;

[0059] Perform Gaussian filtering operation on the transition region according to the Gaussian filtering algorithm parameters corresponding to the transition region to obtain the target processed picture.

[0060] As an optional implementation manner, in the second aspect of the present invention, the specific operation manner of the intelligent recognition module to perform intelligent recognition operation on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain the target video data includes:

[0061] Perform intelligent recognition operation on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain an intelligent recognition result;

[0062] Extract target characters from a pre-loaded static font library based on the intelligent recognition result, where the target characters are used to represent the intelligent recognition result through a preset expression manner;

[0063] Overlay the target characters onto the video stream corresponding to the initial video data through a rendering operation to obtain the target video data.

[0064] As an optional implementation manner, in the second aspect of the present invention, the device further includes:

[0065] A video decoding module, configured to pull a target video stream corresponding to the target video data from the monitoring terminal, and decode the target video stream according to a preset decoding scheme to obtain a decoded target video stream;

[0066] A data derivation module, configured to obtain multiple derived video streams corresponding to the target video stream according to a preset data multiplexing scheme based on the decoded target video stream, and perform re-encoding operations on each derived video stream according to an encoding scheme corresponding to the preset decoding scheme to obtain multiple derived video data.

[0067] As an alternative implementation manner, in the second aspect of the present invention, the specific operation manner of the intelligent recognition module for performing intelligent recognition operations on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain target video data includes:

[0068] Performing intelligent recognition operations on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain an intelligent recognition result;

[0069] Based on the intelligent recognition result, extracting the derived target characters corresponding to each derived video data from a pre-loaded static font library, where the derived target characters correspond to the resolution and bit rate of the corresponding derived video data, and the derived target characters are used to represent the intelligent recognition result through a preset expression manner;

[0070] For each derived video data, superimposing the derived target characters onto the video stream corresponding to the derived video data through a rendering operation to obtain target derived video data.

[0071] As an alternative implementation manner, in the second aspect of the present invention, the device further includes:

[0072] An audio analysis module, configured to obtain the original audio data recorded when the monitoring terminal captures the initial video data, extract the audio features in the original audio data and generate an audio feature set; determine whether there are matching audio features corresponding to a pre-determined recognition target in the audio feature set;

[0073] And the specific operation manner of the target recognition module for determining whether there is a recognition target in the initial video data based on the initial video data includes:

[0074] Identifying the initial video data based on the online monitoring recognition model. If the online monitoring recognition model identifies a recognition target from the initial video data or determines that there are the matching audio features in the audio feature set, it is determined that there is a recognition target in the initial video data.

[0075] The third aspect of the present invention discloses another video data acquisition system for a monitoring terminal, the system includes:

[0076] A memory storing executable program code;

[0077] A processor coupled to the memory;

[0078] The processor calls the executable program code stored in the memory and executes the video data acquisition method of the monitoring terminal disclosed in the first aspect of the present invention.

[0079] Compared with the prior art, in the present invention, the monitoring terminal closely cooperates with the online monitoring recognition model to realize real-time intelligent recognition of monitoring data, and the online monitoring recognition model can upgrade and update itself according to the recognition situation of the current video, so that the online monitoring recognition model can well recognize the latest monitoring background picture, thereby improving the accuracy of intelligent recognition of video data. Description of the Drawings

[0080] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0081] Figure 1 It is a schematic flowchart of a video data acquisition method of a monitoring terminal disclosed in an embodiment of the present invention;

[0082] Figure 2 It is a schematic structural diagram of a video data acquisition device of a monitoring terminal disclosed in an embodiment of the present invention;

[0083] Figure 3 It is a schematic structural diagram of a video data acquisition system of a monitoring terminal disclosed in an embodiment of the present invention. Detailed Embodiments

[0084] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0085] In the description, claims, and above-mentioned drawings of the present invention, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or terminal that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or terminals.

[0086] Referring to "embodiments" herein means that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present invention. The phrase appears in various places in the description and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0087] The present invention discloses a method and device for obtaining video data of a monitoring terminal, which is used to improve the accuracy of intelligent recognition of video data.

[0088] Embodiment 1

[0089] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for obtaining video data of a monitoring terminal disclosed in an embodiment of the present invention. Among them, Figure 1 the described method for obtaining video data of the monitoring terminal can be integrated in a video data obtaining device of the monitoring terminal or the monitoring terminal, and the video data obtaining device of the monitoring terminal can be integrated in a cloud server or a local server. As Figure 1 shown, the method for obtaining video data of the monitoring terminal may include the following operations:

[0090] Step 101: Obtain the initial video data captured by the monitoring terminal, and based on the initial video data, determine whether there is an identification target in the initial video data.

[0091] In an embodiment of the present invention, the monitoring terminal cooperates with an online monitoring recognition model to implement intelligent processing of video data. Among them, the online monitoring recognition model can be integrated in the monitoring terminal or in other devices connected to the monitoring terminal. Based on the cooperation between the monitoring terminal and the online monitoring recognition model, it is possible to determine whether there is an identification target in the initial video data according to the initial video data.

[0092] Step 102: When it is determined that there is no identification target in the initial video data, obtain the target picture in the initial video data, and based on the online monitoring recognition model, identify the target picture to obtain the recognition result of the online monitoring recognition model.

[0093] In the embodiments of the present invention, when there is an identification target in the initial video data, the online monitoring identification model needs to perform identification operations in real time, and since there is an identification target in the picture and a complete background picture cannot be obtained, no upgrade operation is performed on the online monitoring identification model at this time. If there is no identification target in the initial video data, the picture of the video data at this time is the latest background picture, and a further model upgrade operation needs to be performed so that the latest model can recognize the latest background picture.

[0094] Specifically, based on the online monitoring identification model to identify the target picture, the identification result of the online monitoring identification model is obtained. Here, the identification result can be the identification result obtained by the online monitoring identification model when identifying the target picture according to the normal processing flow; or the online monitoring identification model can be set to the test mode, and the online identification model in the test mode can identify the target picture and extract the feature of each element in the target picture.

[0095] In the embodiments of the present invention, optionally, the online monitoring identification model can identify the background of the monitoring picture, and then quickly identify the identification target from the background. At this time, the algorithm of the monitoring identification model is simple and the intelligent identification efficiency is high.

[0096] Step 103: Determine whether the identification result meets the preset model upgrade condition. If the identification result meets the preset model upgrade condition, perform an upgrade operation on the online monitoring identification model according to the target picture to obtain an upgraded online monitoring identification model.

[0097] In the embodiments of the present invention, it is also necessary to further determine whether the identification result indicates that the current online monitoring identification model needs to be upgraded, that is, to determine whether the current online monitoring identification model recognizes the current target picture. Specifically, in the embodiments of the present invention, it is determined whether the identification result meets the preset model upgrade condition. If the identification result meets the preset model upgrade condition, it means that an upgrade operation needs to be performed on the current online monitoring identification model. Then, an upgrade operation is performed on the online monitoring identification model according to the target picture, so that the online monitoring identification model can recognize the current target picture as the background, thereby obtaining an upgraded online monitoring identification model.

[0098] Step 104: Perform an intelligent identification operation on the initial video data captured by the monitoring terminal through the upgraded online monitoring identification model to obtain target video data.

[0099] In the embodiment of the present invention, after the monitoring terminal captures the initial video data, the upgraded online monitoring recognition model performs intelligent recognition operations on the initial video data, such as identifying the recognition target in the video data, judging the state of the recognition target in the video data, giving an alarm according to the video data, and judging whether the video data meets the execution conditions of certain operations. Combining the intelligent recognition operations, the finally obtained data is the target video data.

[0100] It can be seen that in the embodiment of the present invention, the monitoring terminal and the online monitoring recognition model cooperate closely to realize the intelligent recognition of real-time monitoring data, and the online monitoring recognition model can upgrade and update itself according to the recognition situation of the current video, so that the online monitoring recognition model can well recognize the latest monitoring background picture, thereby improving the accuracy of the intelligent recognition of video data.

[0101] In an optional embodiment, based on the online monitoring recognition model to recognize the target picture and obtain the recognition result of the online monitoring recognition model, it may include:

[0102] The target picture is divided into multiple sub-regions according to a preset division strategy. Among them, the recognition ability of the target picture should be the recognition ability of its whole. Therefore, the target picture can be divided into multiple sub-regions, and then each sub-region is recognized separately, so as to comprehensively judge the recognition ability of the online monitoring recognition model.

[0103] In each sub-region, a preset pattern to be recognized is inserted. In this optional embodiment, the core principle of judging the recognition ability of the online monitoring recognition model is: in each sub-region, a preset pattern to be recognized is inserted. Optionally, the pattern to be recognized can be a specific graphic, shape, non-vector symbol, or a pattern of a recognition target such as a vehicle, a pedestrian, a specific object, etc. Among them, the pattern to be recognized is directly inserted into the sub-region. Because there are differences between the pattern to be recognized and the pixels of the sub-region, it is often very easy to recognize the pattern to be recognized from the sub-region. In order to make the insertion of the pattern to be recognized more natural and closer to the image obtained in the formal real scene, in this optional embodiment, the Gaussian filtering method is used to process the connection between the pattern to be recognized and the sub-region.

[0104] Specifically, for each sub-region, the pixel region within a preset range around the edge position of the pattern to be recognized is determined as the transition region, and the Gaussian filtering operation is performed on the transition region to obtain the target processed picture. Optionally, in order to improve the operation efficiency, the preset range is the N pixels on the left and the N pixels on the right of the edge pixels of the pattern to be recognized, where N is the preset number of pixels, which is more convenient for computer operation. In this embodiment, the Gaussian filtering method is used to process the connection between the pattern to be recognized and the sub-region, so that the pattern to be recognized and the sub-region are more naturally and realistically fused, and the reliability of the subsequent test results is improved.

[0105] Identify the target processed image based on the online monitoring recognition model, and obtain the pattern recognition result of the online monitoring recognition model. The pattern recognition result may include the pattern recognized by the online monitoring recognition model from the target processed image, that is, the area of the pattern to be recognized recognized by the online monitoring recognition model from the target processed image.

[0106] Moreover, the above judgment of whether the recognition result meets the preset model upgrade conditions may include:

[0107] Compare the pattern recognition result with the pattern to be recognized, and calculate the recognition accuracy rate of the pattern recognition result. For example, for each pattern to be recognized, use the method of pattern similarity analysis to judge the similarity between the pattern recognition result and the pattern to be recognized. When the similarity is lower than a certain threshold, it is determined that the recognition is incorrect. When all the patterns to be recognized, the proportion of the number of correctly recognized patterns is the recognition accuracy rate.

[0108] If the accuracy rate of the pattern recognition result is lower than the preset accuracy rate threshold, it is determined that the recognition result meets the preset model upgrade conditions.

[0109] In this optional embodiment, the target image is divided into multiple sub-regions according to the preset division strategy, and the Gaussian filtering method is used to process the connection between the pattern to be recognized and the sub-regions, making the pattern to be recognized and the sub-regions fuse more naturally and realistically, and improving the reliability of the judgment process for determining whether the recognition result meets the preset model upgrade conditions.

[0110] In another optional embodiment, for the monitoring terminal, without installing a special lens, the acquired video and image data generally satisfy the condition that objects are larger and clearer when closer and smaller and blurrier when farther away. Therefore, if the same recognition requirements are applied to each sub-region, it does not conform to the actual situation. For this reason, in this optional embodiment, the above-mentioned division of the target image into multiple sub-regions according to the preset division strategy may include:

[0111] Determine the distances from a preset number of reference points in the target image to the monitoring terminal according to the lens parameters corresponding to the monitoring terminal and the target image; the lens parameters include a series of parameters such as the distortion coefficient and focal length. According to the lens parameters and the target image, the image coordinates of each point on the target image can be converted into world coordinates, and the true position of each point can be calculated.

[0112] All reference points are divided into multiple subsets of reference points. Among them, within each subset of reference points, the distances from all the sub-reference points to the monitoring terminal are within the preset distance range corresponding to this subset of reference points. Through the above operations, the sub-reference points that are close in distance and clustered together are grouped into the same subset of reference points. Then, according to all the subsets of reference points, the target picture is divided into multiple sub-regions, where each subset of reference points corresponds to one sub-region. In this way, on the entire target picture, some positions closer to the lens and some positions farther from the lens will appear in different sub-regions.

[0113] It can be seen that when dividing different sub-regions in this alternative embodiment, the different sub-regions are divided according to the distance from the lens of the monitoring terminal, so that the true positions of the images within the same sub-region are close, and thus the sub-regions are divided more reasonably.

[0114] In this alternative embodiment, further optionally, the above-mentioned Gaussian filtering operation on the transition region to obtain the target processed picture may include:

[0115] Determine the Gaussian filtering algorithm parameters according to the preset distance range corresponding to each sub-region; specifically, the Gaussian filtering algorithm parameters generally include the filter kernel size (ksize) and the Gaussian filtering standard deviation (σ). Among them, the filter kernel size refers to the width and height of the convolution kernel, usually an odd number (such as 3×3, 5×5, etc.). It determines the range of neighboring pixels during the filtering process. The larger the filter kernel, the more obvious the smoothing effect, but the computational complexity will also increase. The standard deviation (σ) is the key parameter of the Gaussian function, which determines the diffusion degree of the Gaussian distribution. The larger σ is, the flatter the weight distribution of the Gaussian kernel, and the stronger the smoothing effect; the smaller σ is, the more the weights are concentrated in the center, and the weaker the smoothing effect. Generally, for the sub-regions with a farther preset distance range, the set filter kernel and standard deviation are smaller, and for the sub-regions with a closer preset distance range, the set filter kernel and standard deviation are larger. Thus, for the closer sub-regions, the pattern to be recognized combines more naturally with the sub-region, and the higher the requirement for the recognition ability of the online monitoring recognition model; for the farther sub-regions, the requirement for the naturalness of the combination of the pattern to be recognized and the sub-region is lower, and the requirement for the recognition ability of the online monitoring recognition model is also lower.

[0116] Perform a Gaussian filtering operation on the transition region according to the Gaussian filtering algorithm parameters corresponding to the transition region to obtain the target processed picture. Among them, for the target processed picture obtained by the above method, different sub-regions correspond to different Gaussian filtering algorithms according to their distances, so that the obtained target processed picture is closer to the real situation and improves the reliability of the subsequent judgment process.

[0117] In another optional embodiment, the results recognized by the online monitoring recognition model are generally described in computer language, but in some scenarios, users need to be able to see the description of the results, such as inserting corresponding characters or agreed symbols in the video screen, so that users can see the recognition results in real time. In the prior art, the traditional OSD character overlay technology process is relatively complicated. It is necessary to first generate the characters into a picture format, then decode the picture, and finally overlay the decoded picture with the decoded data of the target video. The whole process is extremely time-consuming and seriously affects the operating efficiency of the system.

[0118] Therefore, in this optional embodiment, the above-mentioned intelligent recognition operation is performed on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain the target video data, which may include:

[0119] Through the upgraded online monitoring recognition model, the initial video data captured by the monitoring terminal is intelligently recognized to obtain intelligent recognition results;

[0120] Based on the intelligent recognition result, a target character is extracted from a pre-loaded static character library, wherein the target character is used to represent the intelligent recognition result in a preset expression mode;

[0121] The target character is superimposed on the video stream corresponding to the initial video data through a rendering operation to obtain the target video data.

[0122] In this optional embodiment, a technical solution combining static loading and dynamic mapping is adopted. For example, a complete set of static character libraries can be generated in advance and loaded into the memory in advance. When superimposing characters, the character generation step can be directly skipped, and only the rendering operation is required to complete the character superposition, which greatly improves the efficiency of character superposition.

[0123] In another optional embodiment, the data sent by the monitoring terminal is often sent to different host computers, and different host computers may require different resolutions and bit rates. In the traditional solution, the above tasks are handed over to the downstream network for execution. When the number of monitoring terminals is large, it will bring great pressure to the downstream network bandwidth. Therefore, in this optional embodiment, the method may also include:

[0124] Pull a target video stream corresponding to a target video data from the monitoring terminal, decode the target video stream according to a preset decoding scheme, and obtain a decoded target video stream;

[0125] According to the decoded target video stream, multiple derivative video streams corresponding to the target video stream are obtained according to a preset data multiplexing scheme, and a re-encoding operation is performed on each derivative video stream according to a coding scheme corresponding to the preset decoding scheme to obtain multiple derivative video data.

[0126] The above scheme is illustrated as follows:

[0127] In the video compression and transcoding technology system, the RTSP protocol can be used. With its excellent compatibility, the protocol can achieve seamless connection with cameras of various brands and manufacturers. In the actual operation process, the software only needs to pull one high-definition video stream from the monitoring terminal and decode the H264 data in it once. Subsequently, with the help of data multiplexing technology and re-encoding technology, it can efficiently output multiple standard H264 video streams with different resolutions and bit rates. This solution not only fully meets the complex and diverse business needs, but also significantly reduces the transmission pressure of the downstream network bandwidth, providing a strong guarantee for the stable operation of the system.

[0128] In this optional embodiment, further optionally, in the case of a multiplexed scenario, the above-mentioned upgraded online monitoring recognition model is used to perform intelligent recognition operations on the initial video data captured by the monitoring terminal to obtain the target video data, which may also include:

[0129] Through the upgraded online monitoring recognition model, the initial video data captured by the monitoring terminal is intelligently recognized to obtain intelligent recognition results;

[0130] Based on the intelligent recognition result, the derived target characters corresponding to each channel of derived video data are extracted from the pre-loaded static character library, wherein the derived target characters correspond to the resolution and bit rate of the corresponding derived video data, and the derived target characters are used to represent the intelligent recognition result in a preset expression mode;

[0131] For each channel of derived video data, the derived target characters are superimposed on the video stream corresponding to the derived video data through a rendering operation to obtain the target derived video data.

[0132] In this optional embodiment, for video streams with different resolutions and bit rates, character information adapted thereto can be flexibly superimposed, thereby achieving an organic combination of static loading and dynamic mapping and improving the flexibility of character superposition.

[0133] In another optional embodiment, the identification target is often accompanied by its unique sound, such as human footsteps, vehicle engine sound, etc. In many cases, the identification target has not appeared in the video screen, but the sound it emits can be detected. At this time, it often means that the identification target is nearby and may appear in the video screen at any time. Therefore, in order to improve the identification reliability of the identification target, the method may also include:

[0134] Obtaining the original audio data recorded when the monitoring terminal captures the initial video data, extracting the audio features in the original audio data and generating an audio feature set;

[0135] Determine whether there are matching audio features corresponding to a pre-determined recognition target in the audio feature set; among them, the matching audio features corresponding to the pre-determined recognition target can be specific footsteps, door opening sounds, and any sounds made by humans, or vehicle engine sounds, vehicle horn sounds, etc.

[0136] And, according to the initial video data, determining whether there is a recognition target in the initial video data may include:

[0137] Based on the online monitoring recognition model, recognize the initial video data. If the online monitoring recognition model recognizes a recognition target from the initial video data, or determines that there are matching audio features in the audio feature set, it is determined that there is a recognition target in the initial video data.

[0138] In this optional embodiment, the determination of whether there is a recognition target is comprehensively judged by combining the video picture and the sound feature information, thereby improving the reliability of the determination of whether there is a recognition target.

[0139] Embodiment Two

[0140] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a video data acquisition device of a monitoring terminal disclosed in an embodiment of the present invention. As Figure 2 shown, the video data acquisition device of this monitoring terminal may include:

[0141] A target recognition module 201, configured to obtain the initial video data captured by the monitoring terminal, and determine whether there is a recognition target in the initial video data according to the initial video data;

[0142] A model testing module 202, configured to, when it is determined that there is no recognition target in the initial video data, obtain the target picture in the initial video data, recognize the target picture based on the online monitoring recognition model, and obtain the recognition result of the online monitoring recognition model;

[0143] A model upgrade module 203, configured to determine whether the recognition result meets a preset model upgrade condition. If the recognition result meets the preset model upgrade condition, perform an upgrade operation on the online monitoring recognition model according to the target picture to obtain an upgraded online monitoring recognition model;

[0144] An intelligent recognition module 204, configured to perform an intelligent recognition operation on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain target video data.

[0145] In an alternative embodiment, the specific operation method for the model testing module 202 to identify a target picture based on the online monitoring recognition model and obtain the recognition result of the online monitoring recognition model may include:

[0146] Divide the target picture into multiple sub-regions according to a preset division strategy;

[0147] Insert a preset pattern to be recognized into each sub-region. For each sub-region, determine the pixel region within a preset range around the edge position of the pattern to be recognized as the transition region, and perform a Gaussian filtering operation on the transition region to obtain a target processed picture;

[0148] Recognize the target processed picture based on the online monitoring recognition model to obtain the pattern recognition result of the online monitoring recognition model. Among them, the pattern recognition result may include the pattern recognized by the online monitoring recognition model from the target processed picture;

[0149] In addition, the specific operation method for the model upgrade module 203 to determine whether the recognition result meets the preset model upgrade conditions may include:

[0150] Compare the pattern recognition result with the pattern to be recognized, and calculate the recognition accuracy rate of the pattern recognition result;

[0151] If the accuracy rate of the pattern recognition result is lower than the preset accuracy rate threshold, it is determined that the recognition result meets the preset model upgrade conditions.

[0152] In another alternative embodiment, the specific operation method for the model testing module 202 to divide the target picture into multiple sub-regions according to a preset division strategy may include:

[0153] According to the lens parameters corresponding to the monitoring terminal and the target picture, determine the distances from multiple preset reference points in the target picture to the monitoring terminal;

[0154] Divide all the reference points into multiple sub-reference point sets. Among them, within each sub-reference point set, the distances from all sub-reference points to the monitoring terminal are within the preset distance range corresponding to the sub-reference point set;

[0155] According to all the sub-reference point sets, divide the target picture into multiple sub-regions, where each sub-reference point set corresponds to a sub-region.

[0156] In yet another alternative embodiment, the specific operation method for the model testing module 202 to perform a Gaussian filtering operation on the transition region to obtain a target processed picture may include:

[0157] Determine the Gaussian filtering algorithm parameters according to the preset distance range corresponding to each sub-region;

[0158] Perform Gaussian filtering operation on the transition region according to the Gaussian filtering algorithm parameters corresponding to the transition region to obtain the target processed picture.

[0159] In yet another optional embodiment, the specific operation mode of the intelligent recognition module 204 to perform intelligent recognition operation on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain the target video data may include:

[0160] Perform intelligent recognition operation on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain the intelligent recognition result;

[0161] Based on the intelligent recognition result, extract the target characters from the static font library loaded in advance, where the target characters are used to represent the intelligent recognition result through a preset expression method;

[0162] Overlay the target characters onto the video stream corresponding to the initial video data through a rendering operation to obtain the target video data.

[0163] In yet another optional embodiment, the device may further include:

[0164] A video decoding module, configured to pull a target video stream corresponding to a target video data from the monitoring terminal, and decode the target video stream according to a preset decoding scheme to obtain the decoded target video stream;

[0165] A data derivation module, configured to obtain multiple derived video streams corresponding to the target video stream according to a preset data multiplexing scheme based on the decoded target video stream, and perform re-encoding operations on each derived video stream according to an encoding scheme corresponding to the preset decoding scheme to obtain multiple derived video data.

[0166] In yet another optional embodiment, the specific operation mode of the intelligent recognition module 204 to perform intelligent recognition operation on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain the target video data may include:

[0167] Perform intelligent recognition operation on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain the intelligent recognition result;

[0168] Based on the intelligent recognition result, extract the derived target characters corresponding to each derived video data from the static font library loaded in advance, where the derived target characters correspond to the resolution and bit rate of the corresponding derived video data, and the derived target characters are used to represent the intelligent recognition result through a preset expression method;

[0169] For each derived video data, the derived target character is superimposed on the video stream corresponding to the derived video data through a rendering operation to obtain the target derived video data.

[0170] In yet another alternative embodiment, the device may further include:

[0171] An audio analysis module, configured to obtain the original audio data recorded when the monitoring terminal captures the initial video data, extract the audio features in the original audio data and generate an audio feature set; determine whether there is a matching audio feature corresponding to a pre-determined recognition target in the audio feature set;

[0172] Moreover, the specific operation method for the target recognition module 201 to determine whether there is a recognition target in the initial video data based on the initial video data may include:

[0173] Recognize the initial video data based on an online monitoring recognition model. If the online monitoring recognition model recognizes a recognition target from the initial video data, or determines that there is a matching audio feature in the audio feature set, it is determined that there is a recognition target in the initial video data.

[0174] Embodiment III

[0175] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a video data acquisition system of a monitoring terminal disclosed in an embodiment of the present invention. As Figure 3 shown, the video data acquisition system of the monitoring terminal may include:

[0176] A memory 301 storing executable program code;

[0177] A processor 302 coupled to the memory 301;

[0178] The processor 302 calls the executable program code stored in the memory 301 and executes the steps in the video data acquisition method of the monitoring terminal described in Embodiment I of the present invention.

[0179] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0180] Through the specific descriptions of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, which includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.

[0181] Finally, it should be noted that: The video data acquisition method and device of a monitoring terminal disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than limiting them; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for acquiring video data of a monitoring terminal, characterized in that: The method comprises: Acquire initial video data captured by the monitoring terminal, and determine whether there is an identification target in the initial video data according to the initial video data; When it is determined that the identification target does not exist in the initial video data, obtaining a target image in the initial video data, and dividing the target image into a plurality of sub-areas according to a preset division strategy; Inserting a preset pattern to be identified in each sub-region, for each sub-region, determining a pixel region within a preset range around an edge position of the pattern to be identified as a transition region, performing a Gaussian filter operation on the transition region, and obtaining a target processing image; Identify the target processing image based on the online monitoring recognition model, and obtain a pattern recognition result of the online monitoring recognition model, wherein the pattern recognition result includes a pattern recognized by the online monitoring recognition model from the target processing image; Comparing the pattern recognition result with the pattern to be recognized, and calculating the recognition accuracy of the pattern recognition result; If the accuracy of the pattern recognition result is lower than a preset accuracy threshold, performing an upgrade operation on the online monitoring recognition model according to the target image to obtain an upgraded online monitoring recognition model; The upgraded online monitoring recognition model is used to perform intelligent recognition operations on the initial video data captured by the monitoring terminal to obtain target video data.

2. The video data acquisition method of the monitoring terminal according to claim 1, characterized in that: The step of dividing the target image into a plurality of sub-areas according to a preset division strategy includes: Determining the distances from a plurality of reference points preset in the target image to the monitoring terminal according to the lens parameters corresponding to the monitoring terminal and the target image; Divide all reference points into a plurality of sub-reference point sets, wherein, in each of the sub-reference point sets, the distances from all the sub-reference points to the monitoring terminal are within a preset distance range corresponding to the sub-reference point set; According to all the sub-reference point sets, the target image is divided into a plurality of sub-regions, wherein each of the sub-reference point sets corresponds to a sub-region.

3. The video data acquisition method of the monitoring terminal according to claim 2, characterized in that: The performing of a Gaussian filtering operation on the transition region to obtain a target processed image includes: Determining Gaussian filter algorithm parameters according to the preset distance range corresponding to each sub-area; A Gaussian filtering operation is performed on the transition area according to Gaussian filtering algorithm parameters corresponding to the transition area to obtain a target processed image.

4. The video data acquisition method of the monitoring terminal according to claim 1, characterized in that: The upgraded online monitoring recognition model is used to perform intelligent recognition operations on the initial video data captured by the monitoring terminal to obtain target video data, including: Using the upgraded online monitoring recognition model, an intelligent recognition operation is performed on the initial video data captured by the monitoring terminal to obtain an intelligent recognition result; Based on the intelligent recognition result, extracting a target character from a static character library loaded in advance, wherein the target character is used to represent the intelligent recognition result in a preset expression; The target character is superimposed on the video stream corresponding to the initial video data through a rendering operation to obtain target video data.

5. The video data acquisition method of the monitoring terminal according to claim 1 or 4, characterized in that: The method further comprises: Pulling a target video stream corresponding to the target video data from the monitoring terminal, decoding the target video stream according to a preset decoding scheme, and obtaining a decoded target video stream; According to the decoded target video stream, multiple derivative video streams corresponding to the target video stream are obtained according to a preset data multiplexing scheme, and according to a coding scheme corresponding to the preset decoding scheme, a re-encoding operation is performed on each derivative video stream to obtain multiple derivative video data.

6. The video data acquisition method of the monitoring terminal according to claim 5, characterized in that: The upgraded online monitoring recognition model is used to perform intelligent recognition operations on the initial video data captured by the monitoring terminal to obtain target video data, including: Using the upgraded online monitoring recognition model, an intelligent recognition operation is performed on the initial video data captured by the monitoring terminal to obtain an intelligent recognition result; Based on the intelligent recognition result, extracting a derived target character corresponding to each channel of derived video data from a pre-loaded static character library, wherein the derived target character corresponds to the resolution and bit rate of the corresponding derived video data, and the derived target character is used to represent the intelligent recognition result in a preset expression; For each channel of the derived video data, the derived target characters are superimposed on the video stream corresponding to the derived video data through a rendering operation to obtain target derived video data.

7. The video data acquisition method of the monitoring terminal according to claim 1, characterized in that: The method further comprises: Acquire original audio data recorded when the monitoring terminal captures initial video data, extract audio features from the original audio data and generate an audio feature set; Determining whether there is a matching audio feature corresponding to a predetermined recognition target in the audio feature set; And, judging whether there is an identification target in the initial video data according to the initial video data, comprising: The initial video data is identified based on an online monitoring recognition model. If the online monitoring recognition model identifies a recognition target from the initial video data, or determines that the matching audio feature exists in the audio feature set, it is determined that the recognition target exists in the initial video data.

8. A video data acquisition device for a monitoring terminal, characterized in that: The device includes: The target recognition module is used to obtain the initial video data captured by the monitoring terminal, and determine whether there is an identification target in the initial video data according to the initial video data; A model testing module is used to obtain a target image in the initial video data when it is determined that the identification target does not exist in the initial video data, and divide the target image into multiple sub-areas according to a preset division strategy; insert a preset pattern to be identified in each sub-area, and for each of the sub-areas, determine that the pixel area within a preset range around the edge position of the pattern to be identified is a transition area, perform a Gaussian filtering operation on the transition area, and obtain a target processing image; identify the target processing image based on an online monitoring recognition model, and obtain a pattern recognition result of the online monitoring recognition model, wherein the pattern recognition result includes the pattern recognized by the online monitoring recognition model from the target processing image; A model upgrade module, used to compare the pattern recognition result with the pattern to be recognized, and calculate the recognition accuracy of the pattern recognition result; if the accuracy of the pattern recognition result is lower than a preset accuracy threshold, an upgrade operation is performed on the online monitoring recognition model according to the target image to obtain an upgraded online monitoring recognition model; The intelligent recognition module is used to perform intelligent recognition operations on the initial video data captured by the monitoring terminal through the upgraded online monitoring recognition model to obtain target video data.

9. A video data acquisition system for a monitoring terminal, characterized in that: The system includes: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the video data acquisition method of the monitoring terminal according to any one of claims 1-7.

Citation Information

Patent Citations

  • Smart river chief system based on cloud-edge collaborative silent upgrading and upgrading method

    CN114078233A