Video miniature method, system and storage medium
By sharing video microscope processing tasks between cloud servers and camera devices, and using microscope configuration information and algorithms to identify video frames, the problems of high cost and lack of universality in the prior art are solved, and efficient and universal video microscope effects are achieved.
Patent Information
- Application Number
- CN202510214050.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-26
AI Technical Summary
In the prior art, video microscopy is expensive and lacks versatility, making it difficult for users to quickly understand events occurring in the day.
The cloud server sends microscope configuration information to the specified camera device. The camera uses a specified algorithm to identify the video frame and intercept the target clips during the microscope activation period. The cloud server synthesizes these clips to generate microscope videos.
It reduces the computing burden and cost of cloud servers, improves the efficiency and quality of video processing, and realizes the universality of video microscopes and improves user experience.
Smart Images

Figure CN119697405B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a video miniature method, system and storage medium. Background Art
[0002] In recent years, with the continuous development of video surveillance technology, smart cameras have become an indispensable part of users' homes. Video storage has also gradually changed from memory card recording mode to cloud storage mode, but the number of cloud storage videos in a day is large, and users cannot quickly understand what happened in a day. They can only look through all the cloud playback videos, which is not only time-consuming but also difficult to find the video clips that need attention.
[0003] Currently, there are two main ways to synthesize playback video clips into a short video clip. The first way is to use the cloud server to analyze the surveillance video in the cloud storage, extract the clips that the user is interested in, and synthesize them. The user views the synthesized video through the client. The second way is to use a specific player to analyze and process the surveillance video through structured video, and display a miniature video of key information when the user plays it.
[0004] In the above existing technologies, the first method uses a cloud server to perform intelligent analysis on videos and select video clips that users are interested in. This method requires high computing resources and bandwidth of the cloud server and has high commercial costs. The second method uses a specific player, which lacks versatility and cannot provide users with the function of sharing thumbnail videos. In addition, the thumbnail processing is only performed during playback, which requires a long waiting time, so the user experience is poor. Summary of the invention
[0005] The embodiments of the present application provide a video miniature method, system and storage medium to solve the technical problems of high cost and lack of versatility in the prior art for video miniature.
[0006] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:
[0007] In a first aspect, an embodiment of the present application provides a video epitome method, which is applied to a cloud server, and the method includes:
[0008] According to the device identifier, the epitome configuration information is sent to the camera device specified by the user, wherein the epitome configuration information includes the epitome effective time period and the algorithm identifier;
[0009] Receive a target segment sent by the camera device; the target segment is a video segment corresponding to each recognition result intercepted from the continuous video frames by the camera device during the time period when the reduction is effective, by identifying the continuous video frames using the reduction algorithm corresponding to the algorithm identifier, and according to the recognition result;
[0010] A target miniature video is generated by synthesizing a plurality of the target segments.
[0011] In combination with the first aspect, in a possible design manner, before synthesizing the plurality of target fragments, the method further includes:
[0012] Receiving the recognition result sent by the camera device;
[0013] When the recognition result does not meet the preset accuracy, analyzing the target segment corresponding to the recognition result by using the epitome algorithm to obtain an analysis result;
[0014] If the analysis result meets the preset accuracy, the target segment corresponding to the recognition result is retained; if the analysis result does not meet the preset accuracy, the target segment corresponding to the recognition result is discarded.
[0015] In combination with the first aspect, in a possible design manner, the recognition result includes a first confidence level of the target object recognized by the camera device;
[0016] When the recognition result does not meet the preset accuracy, the target segment corresponding to the recognition result is analyzed by the miniature algorithm to obtain the analysis result, including:
[0017] If the first confidence level is less than a confidence level threshold preset by the cloud server, identifying the target segment corresponding to the recognition result by using the epitome algorithm to obtain the analysis result, wherein the analysis result includes a second confidence level of the target object recognized by the cloud server;
[0018] If the analysis result satisfies the preset accuracy, retaining the target segment corresponding to the recognition result; if the analysis result does not satisfy the preset accuracy, discarding the target segment corresponding to the recognition result, includes:
[0019] If the second confidence is greater than the confidence threshold, the target segment corresponding to the recognition result is retained; if the second confidence is less than or equal to the confidence threshold, the target segment corresponding to the recognition result is discarded.
[0020] In combination with the first aspect, in a possible design manner, the miniature configuration information further includes a duration of the miniature video; and the method further includes:
[0021] Calculating the duration of an initial miniature video after synthesizing the plurality of target segments;
[0022] If the duration of the initial miniature video is greater than the duration of the miniature video, then based on a preset elimination strategy, the target segment among the multiple target segments is eliminated;
[0023] The multiple target segments remaining after elimination are synthesized to generate a target miniature video, wherein the duration of the target miniature video is equal to or less than the duration of the miniature video.
[0024] In combination with the first aspect, in a possible design, the step of eliminating a target segment from the plurality of target segments based on a preset elimination strategy includes:
[0025] Based on the miniature effective time period, determining a miniature time midpoint;
[0026] Classifying the multiple target segments according to the midpoints of the miniature time to obtain two sets of target segment sets;
[0027] For a target segment set having a larger number of target segments in the two groups of target segment sets: if the number of target segments in the target segment set is greater than 1, the target segment set is classified according to the time midpoint of the target segment set until one target segment remains in the target segment set; and the remaining target segments in the target segment set are eliminated.
[0028] In combination with the first aspect, in a possible design manner, the method further includes:
[0029] Determine the device identification of the camera device according to the user's selection operation on the camera device through the client, and determine the miniature algorithm according to the focus content set by the user;
[0030] The miniature configuration information is generated according to the miniature effective time period, miniature background music and miniature video duration set by the user through the client, combined with the algorithm identifier corresponding to the miniature algorithm.
[0031] In combination with the first aspect, in a possible design manner, the method further includes:
[0032] Receiving the recognition result reported by the camera device;
[0033] Generate an analysis report based on the number of occurrences of various events, the time period of high-frequency events, and previous period comparison data within the effective time period of the epitome analyzed from the recognition results;
[0034] The step of synthesizing the plurality of target segments to generate a target miniature video includes:
[0035] Synthesizing a plurality of the target clips, and packaging the synthesis result with the miniature background music to generate a target miniature video;
[0036] A message is pushed to the client to instruct the user to view the target thumbnail video and analysis report.
[0037] In a second aspect, an embodiment of the present application provides a video epitome method, which is applied to a camera device, and the method includes:
[0038] Receiving the epitome configuration information sent by the cloud server according to the device identifier, wherein the epitome configuration information includes the epitome effective time period and the algorithm identifier;
[0039] During the epitome effective time period, the epitome algorithm corresponding to the algorithm identifier is used to process the continuous video frames to obtain the recognition results output by the epitome algorithm from the continuous video frames, and the target segment corresponding to each recognition result is intercepted from the continuous video frames;
[0040] The multiple target segments are sent to the cloud server so that the cloud server synthesizes the multiple target segments to generate a target miniature video.
[0041] In combination with the second aspect, in a possible design, the continuous video frames are acquired in real time by the camera device; within the effective time period of the epitome, the continuous video frames are processed by the epitome algorithm corresponding to the algorithm identifier to obtain the recognition result obtained from the continuous video frames output by the epitome algorithm, including:
[0042] At the miniature start time point of the miniature effective time period, the miniature algorithm corresponding to the algorithm identifier is used to perform frame sampling analysis on the continuous video frames collected in real time to obtain a recognition result of the first recognition of the target object based on the output of the miniature algorithm, and the recognition result includes the detection box coordinates, the target object identifier, the category to which the target object belongs, the time point and the first confidence level.
[0043] In a third aspect, an embodiment of the present application provides a cloud server, comprising a processor and a memory, wherein the processor is configured to run a computer program in the memory to execute the method of the first aspect and possible design methods thereof.
[0044] In a fourth aspect, an embodiment of the present application provides a camera device, comprising a camera, a memory and a processor, wherein the camera captures continuous video frames, the memory stores a computer program, and the processor is configured to run the computer program to execute the method of the second aspect and its possible design method.
[0045] In a fifth aspect, an embodiment of the present application provides a video miniature system, comprising a cloud server and a camera device, wherein the cloud server is configured as the method of the first aspect and its possible design methods, and the camera device is configured to execute the method of the second aspect and its possible design methods.
[0046] In a sixth aspect, an embodiment of the present application provides a storage medium, in which a computer program is stored, wherein the computer program is configured to execute the method of the first aspect and its possible design method, or execute the method of the second aspect and its possible design method when running.
[0047] Compared with the prior art, the embodiment of the present application provides a video epitome method, system and storage medium, in which the cloud server sends epitome configuration information to the camera device specified by the user according to the device identifier, and the epitome configuration information includes the epitome effective time period and the algorithm identifier. The cloud server receives the target segment sent by the camera device; the target segment is the camera device in the epitome effective time period, and the recognition result is obtained from the continuous video frames by the epitome algorithm corresponding to the algorithm identifier; the target segment corresponding to each recognition result is intercepted from the continuous video frame; and the multiple target segments are sent to the cloud server. The cloud server synthesizes the multiple target segments to generate a target epitome video. In this way, the camera device is no longer limited to the shooting function, but can automatically identify the content of interest to the user in the picture (i.e., the recognition result) and filter out the video segment corresponding to the content (i.e., the target segment) from the video frame through the epitome algorithms such as image recognition, motion detection, face recognition, and object tracking. In the subsequent video processing tasks, the cloud server only needs to further analyze a few key frames uploaded by the camera device, so the workload of the cloud server is greatly reduced, and the computing time and cost are also significantly reduced. The intelligent algorithm recognition of the camera equipment and the fragment synthesis of the cloud server work together to realize video miniature processing through the combination of end and cloud, which improves the efficiency and quality of video processing.
[0048] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0050] Figure 1 A flowchart of a video miniature method provided by an embodiment of the present application is shown;
[0051] Figure 2A hardware structure block diagram of a camera device provided in an embodiment of the present application is shown;
[0052] Figure 3 A schematic diagram of a video miniature system provided in an embodiment of the present application is shown;
[0053] Figure 4 A flowchart of a video miniature method provided by an embodiment of the present application is shown;
[0054] Figure 5 A flow chart of a method for analyzing a target fragment by a cloud server provided in an embodiment of the present application is shown;
[0055] Figure 6 A flowchart of another method for analyzing a target fragment by a cloud server provided in an embodiment of the present application is shown;
[0056] Figure 7 A flow chart of a material elimination method provided in an embodiment of the present application is shown;
[0057] Figure 8 A schematic diagram of a video content epitome system architecture provided by an embodiment of the present application is shown;
[0058] Fig. 9 A structural block diagram of a video miniature device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0059] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0060] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by people with general skills in the technical field to which this application belongs. The words "one", "a", "the", "these" and the like in this application do not indicate a quantitative limitation, and they may be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and apparatus, product or equipment comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or equipment. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" may mean: A exists alone, A and B exist at the same time, and B exists alone. Generally, the character " / " indicates that the objects associated with each other are in an "or" relationship. The terms "first", "second", "third", etc. in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.
[0061] With the continuous development of video surveillance technology, video cameras have become an indispensable part of users' homes. By performing miniaturization processing on the continuous video frames captured by the video camera and generating a short cloud playback video, the efficiency of users' video playback can be greatly improved. However, most of the current methods use cloud servers or specific playback terminals to process videos, ignoring the data processing capabilities of the video camera on the terminal side. Based on this, an embodiment of the present application provides a video miniaturization method, in which the video camera performs preliminary processing on the continuous video frames, and the cloud server performs a small amount of processing work based on the processing results of the video camera to generate a target miniature video.
[0062] For details, please refer to Figure 1 , Figure 1 A flowchart of a video miniature method provided in an embodiment of the present application is shown, and the method includes steps S101 to S105.
[0063] Step S101: The user sets the miniature configuration on the client.
[0064] Step S102: The platform (or cloud server) sends the algorithm plug-in to the camera device.
[0065] Step S103: The camera device reports the video clip and the recognition result to the platform.
[0066] Step S104: The platform analyzes and edits the video clips to obtain a target miniature video.
[0067] Step S105: The user views the target thumbnail video through the client.
[0068] It can be seen that in step S102, the camera device does part of the algorithm recognition work, sharing the computing pressure of the cloud server. The cloud server only needs to perform algorithm recognition on a few video frames, and then remove the video clips that do not meet the conditions (such as removing the video clips that are mistakenly detected by the camera device) based on the algorithm recognition results, and only synthesize the remaining video clips. In this embodiment, the cloud server does not need to analyze all the video frames captured by the camera device, so it effectively solves the problem of high cost of analyzing and synthesizing the miniature video on the pure cloud side.
[0069] In addition, the advantage of using the camera device to perform algorithm recognition is that there is no need to use a specific player dedicated to miniature processing and playing miniature videos. This is because the client of the camera device can display the target miniature video based on the real-time video, and the user can watch or share the target miniature video by opening the client. Compared with a specific player, the camera device is more versatile.
[0070] The method provided in the embodiment of the present application can be applied to the fields of intelligent security, smart home, vehicle monitoring, etc., and can be executed in a terminal, computer or similar computing system. Taking running on a camera device as an example, Figure 2 FIG. 1 shows a hardware structure block diagram of a camera device provided in an embodiment of the present application. Figure 2 As shown, the camera device may include one or more ( Figure 2 The image capturing device may further include a transmission device 203 for communication functions, an input / output device 204, and a camera 205.
[0071] It can be understood by those skilled in the art that Figure 2 The structure shown is only for illustration and does not limit the structure of the above-mentioned camera device. Figure 2 More or fewer components as shown, or with Figure 2 Different configurations are shown.
[0072] The processor 201 may include one or more processing units, for example, the processor 201 may include an application processor (AP), a graphics processor (GPU), an image signal processor (ISP), etc. Different processing units may be independent devices or integrated into one or more processors.
[0073] The memory 202 can be used to store computer programs, for example, software programs and modules of application software. The processor 201 executes various functional applications and data processing by running the computer programs stored in the memory 202, that is, to implement the above method. The memory 202 can be used to store data. The memory 202 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 202 may further include a memory remotely arranged relative to the processor 201, and these remote memories may be connected to the camera device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0074] The transmission device 203 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the camera device. In one example, the transmission device 203 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 203 can be a radio frequency (Radio Frequency, referred to as RF) module, which is used to communicate with the Internet wirelessly.
[0075] The camera device can realize the shooting function through ISP, camera 205, video codec, and GPU. Camera 205 is used to capture video; ISP is used to process the video collected by camera 205, such as skipping frames to select video frames, and distribute them to GPU, GPU recognizes the video frames through the epitome algorithm to obtain recognition results.
[0076] The embodiments of the present application can be applied to scenarios where multiple camera devices are combined with a cloud server. Figure 3 As shown, each camera device corresponds to a device identifier, such as multiple devices including: camera device A in room A, camera device B in room B. The user specifies one or more camera devices on the client, such as camera device A, and the cloud server sends a message to camera device A, indicating that camera device A is required to perform miniature processing on the video.
[0077] The following uses the video recording device having the above hardware structure as an example to illustrate the video miniaturization method provided by the embodiment of the present application. The method can realize the function that the end side performs a first analysis on the real-time image of the camera, and the cloud side performs a second analysis on the image reported by the end side. The method includes the following steps: Figure 4 Steps S401 to S404 are shown.
[0078] Step S401: The cloud server sends the miniature configuration information to the camera device according to the device identifier, where the miniature configuration information includes the miniature effective time period and the algorithm identifier.
[0079] Specifically, the camera device has a client for users, and the user can set the miniature on the client, such as specifying a camera device and setting the miniature effective time period. The miniature effective time period includes the miniature start time point and the miniature end time point, indicating that the video within this period is miniature processed.
[0080] In some embodiments, the miniature configuration information includes the duration of the miniature video, the effective time period of the miniature, and the algorithm identifier. Exemplarily, the user sets 9:00-17:00 on the client, and this time period represents the effective time period of the miniature, that is, the camera device performs miniature processing on the video within this time period. The user sets the pet to be followed on the client, and the cloud server determines that the camera device needs to call the pet detection algorithm, and the algorithm identifier is pet001 corresponding to the pet detection algorithm. The user sets 5 minutes on the client, and this duration represents the duration of the miniature video, that is, the video within the time period is miniaturized to a target miniature video with a length of 5 minutes. After the user completes the setting, the cloud server sends the miniature configuration information to the camera device.
[0081] Users can also set the content of interest on the client. For example, the content of interest refers to the target object of the miniature. For example, if the content of interest is pets, then the target object of the miniature is pets, and the video clips related to pets uploaded to the cloud server by the camera device; if the content of interest is falling, then the target object of the miniature is the falling action, and the video clips of the object (or specified user) falling detected by the camera device are uploaded to the cloud server. In short, different content of interest to users has different corresponding miniature algorithms.
[0082] The cloud server determines the epitome algorithm based on the content of interest set by the user, and sends the algorithm identifier corresponding to the epitome algorithm to the camera device. In some embodiments, the camera device runs the epitome algorithm after receiving the epitome configuration information. In some embodiments, the camera device loads the epitome algorithm at the epitome effective start point indicated by the epitome effective time period to execute the steps corresponding to the epitome algorithm, specifically referring to step S402.
[0083] Step S402: The camera device obtains recognition results from continuous video frames by using the algorithm to identify the corresponding reduction algorithm during the reduction effective time period.
[0084] Specifically, the reduction takes effect in a time period of [t1, t2]. The camera device calls the reduction algorithm at time t1 to start reducing the continuous video frames of the video, and ends the algorithm at time t2.
[0085] The epitome algorithm refers to an algorithm for extracting video clips from a video. The epitome algorithm is used to automatically extract and retain representative clips (such as key frames, scenes, or event clips) in the video by analyzing the structure of the video content, scene changes, event development and other information, thereby generating a simplified cloud playback video. The cloud playback video retains the key information of the original video while reducing the redundant parts. As an example, the epitome algorithm includes a target recognition algorithm and a behavior detection algorithm.
[0086] The miniature algorithm can be built into the camera device or stored in the algorithm warehouse. There can be multiple miniature algorithms, so in this step, the cloud server sends the algorithm identifier to the camera device so that the camera device can obtain the corresponding miniature algorithm from the memory or from the algorithm warehouse based on the algorithm identifier.
[0087] Step S403: The camera device captures the target segment corresponding to each recognition result from the continuous video frames.
[0088] Among them, the target segment corresponding to the recognition result can be a video frame in which the target object is recognized and a video composed of multiple video frames before and after it. It can be understood that different miniature algorithms and different recognition results will result in different corresponding target segments. For example, when the miniature algorithm is a pet detection algorithm, the recognition result is a pet detection frame in a video frame, and the corresponding target segment is a video frame in which the pet is detected and a video composed of multiple video frames before and after it. When the miniature algorithm is a fall detection algorithm, the recognition result is a detection frame of a fall action in multiple continuous video frames, and the corresponding target segment is a video composed of the multiple continuous video frames and the multiple video frames before and after them.
[0089] More specifically, the target segment is a video of about 2 seconds before and after the pet appearance event is detected, and the video can be an H264 or H265 code stream file.
[0090] In some embodiments, continuous video frames are acquired in real time by a camera device; the camera device extracts frames from the continuous video frames and performs miniature processing on the extracted video frames. Specifically, step S403 further includes: at the miniature start time point of the miniature effective time period, the miniature algorithm corresponding to the algorithm identifier is used to extract frames from the continuous video frames acquired in real time to obtain a recognition result of the first recognition of the target object based on the miniature algorithm output, and the recognition result includes the detection frame coordinates, the target object identifier, the category to which the target object belongs, the time point, and the first confidence. In this way, the number of video frames that the camera device needs to process and store can be reduced, thereby reducing the amount of calculation and storage space.
[0091] Step S404: The cloud server synthesizes multiple target segments to generate a target miniature video.
[0092] Specifically, the cloud server may combine multiple target segments in chronological order to obtain a target miniature video.
[0093] In some embodiments, the cloud server can further analyze multiple target segments. It is understood that the first analysis is performed in the camera device, which means that the camera device uses the epitome algorithm to identify the video frames it has collected. The second analysis is performed on the cloud server, which means that the cloud server uses the epitome algorithm to identify the target segments uploaded by the camera device.
[0094] In one case, the cloud server analyzes each target segment to ensure the accuracy of the final recognition result. In another case, a threshold condition is preset in the cloud server. If the recognition result does not meet the threshold condition, the cloud server will analyze the target segment corresponding to the recognition result, which can simplify the work of the cloud server.
[0095] It is worth mentioning that in this embodiment, the recognition accuracy on the cloud side is higher than that on the terminal side, so the cloud server analyzes the target segment to avoid the problem of misrecognition caused by the low recognition accuracy on the terminal side. This end-cloud combination not only solves the problem of large cloud-side computational complexity, but also avoids the problem of misrecognition caused by insufficient terminal-side computational accuracy, so the video miniature efficiency is higher.
[0096] The following describes the process of the cloud server analyzing the target fragment. Figure 5 A flowchart of a method for analyzing a target segment by a cloud server provided in an embodiment of the present application is shown. The method includes: Figure 5 Steps S501 to S502 are shown.
[0097] Step S501: Before synthesizing multiple target segments, if the recognition result does not meet the preset accuracy, the cloud server analyzes the target segment corresponding to the recognition result through a miniature algorithm to obtain an analysis result.
[0098] Step S502: If the analysis result meets the preset accuracy, the cloud server retains the target segment corresponding to the recognition result; if the analysis result does not meet the preset accuracy, the cloud server discards the target segment corresponding to the recognition result.
[0099] In this method, when the recognition result does not meet the preset accuracy, the cloud server schedules the algorithm plug-in of the miniature algorithm from the algorithm warehouse and runs the miniature algorithm to recognize the target segment. The computing power of the cloud server is higher than that of the camera device, so the detection accuracy of the analysis result obtained by the cloud server is higher than the detection accuracy of the recognition result output by the camera device. Therefore, based on the analysis result output by the cloud server, it is determined whether the target segment needs to be discarded. This can screen multiple target segments uploaded by the camera device to avoid misidentification.
[0100] In some embodiments, in the method shown in steps S501 to S502 above, the recognition result includes a first confidence level of the target object recognized by the camera device. Figure 6 A flowchart of another method for analyzing a target fragment by a cloud server provided in an embodiment of the present application is shown. The method includes Figure 6 Steps S601 to S602 are shown.
[0101] Step S601: If the first confidence is less than the confidence threshold preset by the cloud server, the target segment corresponding to the recognition result is recognized by the miniature algorithm to obtain an analysis result, which includes a second confidence of the target object recognized by the cloud server.
[0102] Step S602: If the second confidence is greater than the confidence threshold, the target segment corresponding to the recognition result is retained; if the second confidence is less than or equal to the confidence threshold, the target segment corresponding to the recognition result is discarded.
[0103] Specifically, the confidence level is used to evaluate the uncertainty of the algorithm output results. A high confidence level indicates that the confidence level of the constructed confidence interval that contains the true value of the population parameter is high, that is, the result is more accurate. Therefore, in this embodiment, the accuracy of the result is evaluated based on the first confidence level and the second confidence level. When the first confidence level is less than the confidence level threshold, it means that the target object identified by the camera device may not be accurate. At this time, the camera device may have misidentified the target object. Therefore, the cloud server performs a second analysis to obtain the second confidence level of the cloud server identifying the target object, and determines whether to discard the target segment based on the comparison result of the second confidence level and the confidence level threshold.
[0104] In the above two embodiments, the cloud server screens the multiple target segments uploaded by the camera device to obtain the videos corresponding to the recognition results with the recognition accuracy reaching the preset accuracy, and combines these videos into a miniature video. In some of the embodiments, the cloud server can determine whether it is necessary to discard the redundant target segments in combination with the duration of the video synthesized from the multiple target segments. Specifically, when the duration of the initial miniature video synthesized from the multiple target segments is greater than the target duration, or greater than the miniature video duration set by the user, it is considered that there are redundant videos in the multiple target segments, and then some of the target segments are discarded so that the duration of the synthesized miniature video is less than or equal to the target duration (or miniature video duration). Among them, the target duration is pre-set by the cloud server or the camera device, and the miniature video duration is set by the user on the user side according to needs.
[0105] The following description takes the duration of the miniature video as a threshold as an example. Specifically, the cloud server is configured to calculate the duration of the initial miniature video after the synthesis of multiple target segments. If the duration of the initial miniature video is greater than the duration of the miniature video, the target segments among the multiple target segments are eliminated based on a preset elimination strategy. The multiple target segments remaining after the elimination are synthesized to generate a target miniature video, and the duration of the target miniature video is equal to or less than the duration of the miniature video.
[0106] In some embodiments, a preset elimination strategy is used to make the material more evenly distributed in time. For example, if the number of first materials (target segments) obtained from the video frames collected by the camera device in the morning is greater than the number of second materials obtained from the video frames collected by the camera device in the afternoon, then some target segments are eliminated from the first material, so that the material obtained from the video frames collected in the morning and afternoon as a whole is more evenly distributed in time, and after the material is combined into the target miniature video, the video playback effect is better.
[0107] A specific implementation of the preset elimination strategy is given below: Figure 7 A flow chart of a material elimination method provided in an embodiment of the present application is shown, the method comprising: Figure 7 Steps S701 to S704 are shown.
[0108] Step S701: Determine the epitome time midpoint based on the epitome effective time period.
[0109] Step S702: Classify the multiple target segments according to the midpoints of the miniature time to obtain two sets of target segment sets.
[0110] Step S703: for the target segment set with more target segments in the two target segment sets, classify them according to the time midpoint of the target segment set, and repeat this step until there is only one target segment left in the target segment set.
[0111] Step S704: Eliminate the remaining target segments in the target segment set.
[0112] In step S701, the effective time period of the miniature is [t1, t2], and the midpoint of the miniature time is t3=(t1+t2) / 2. In step S702, all target segments in [t1, t3] are divided into the first group of target segment sets, and all target segments in [t3, t2] are divided into the second group of target segment sets. In step S703, the number of segments in the first group of target segment sets and the number of segments in the second group of target segment sets are compared, and the larger group (such as the first group of target segment sets) is selected for grouping again. During this grouping, the time midpoint of the group is no longer t3, but the time midpoint of [t1, t3], which is t4=(t1+t3) / 2. The first group of target segment sets is grouped with t4 as the demarcation point, and all target segments in [t1, t4] are divided into the first sub-group segment set, and all target segments in [t4, t3] are divided into the second sub-group segment set. Compare which group has a larger number of segments in the two sets, select the larger group, and further divide according to step S703 until only one target segment remains in the set. It can be understood that the target segment is a temporally redundant material, so the target segment is discarded (or eliminated). In some embodiments, after discarding the target segment, the remaining target segments are combined, and the duration of the obtained miniature video is less than or equal to the target duration, and the target segment is no longer eliminated. In other embodiments, after discarding the target segment, the remaining target segments are combined, and the duration of the obtained miniature video is greater than the target duration, then the remaining target segments are classified according to the time midpoint obtained by combining the remaining target segments to obtain two sets, and the set with the larger number of segments is further classified until there is only one target segment left in the target segment set, and the remaining target segments in the target segment set are eliminated. This is performed multiple times until the duration of the remaining target segment is less than or equal to the target duration.
[0113] pass Figure 7 In the illustrated embodiment, the cloud server controls the duration of the final video to ensure that it does not exceed the target duration. If the remaining material duration exceeds the target duration, the cloud server will continue to refine the material to ensure that the final synthesized video duration is appropriate, thereby meeting the preset duration requirement, which makes the video duration more in line with user needs. And by dividing the material segments according to time points, the material segments that are redundant in time are preferentially eliminated, which can ensure that the final miniature video can cover the entire time period more evenly, rather than some parts being too dense or some parts being too blank. This uniformity of the material in time avoids the problem of uneven rhythm of the video and helps to improve the viewing experience.
[0114] In some embodiments, the cloud server is further configured to receive the recognition results reported by the camera device; generate an analysis report based on the number of events, the high-frequency event time period, and the previous comparison data within the effective time period of the miniature analyzed from the recognition results. In addition, the above step S404 further includes: synthesizing multiple target clips, and packaging the synthesis results with the miniature background music to generate a target miniature video; and pushing a message to the client to instruct the user to view the target miniature video and the analysis report.
[0115] Through this embodiment, the cloud server pushes messages to the user client, and the user can quickly view the thumbnail video and related summary reports of the content of interest in the specified time period on the client, so as to understand the events that occurred on the day more quickly.
[0116] In summary, through the above steps S401 to S404, the cloud server sends the miniature configuration information to the camera device according to the device identifier, and the miniature configuration information includes the miniature effective time period and the algorithm identifier. During the miniature effective time period, the miniature algorithm corresponding to the algorithm identifier is used to identify the recognition result from the continuous video frames; the target segment corresponding to each recognition result is intercepted from the continuous video frames; and the multiple target segments are sent to the cloud server. The cloud server synthesizes the multiple target segments to generate the target miniature video. In this way, the camera device is no longer limited to the shooting function. The camera device automatically identifies the content of interest to the user in the picture (i.e., the recognition result) and the video segment corresponding to the content (i.e., the target segment) through the miniature algorithms such as image recognition, motion detection, face recognition, and object tracking. In the subsequent video processing tasks, the cloud server only needs to further analyze a few key frames uploaded by the camera device, so the workload of the cloud server is greatly reduced, and the computing time and cost are also significantly reduced. The intelligent algorithm recognition of the camera equipment and the fragment synthesis of the cloud server work together to realize video miniature processing through the combination of end and cloud, which greatly improves the efficiency and quality of video processing.
[0117] The method provided in the embodiment of the present application is further illustrated by a specific example below. Figure 8 A schematic diagram of a video content epitome system architecture provided by an embodiment of the present application is shown. Figure 8 As shown, the system includes a client, a video miniature server, a video cloud storage server, and an end-side AI camera.
[0118] Among them, the video miniature server is equivalent to the cloud server mentioned above; the end-side AI camera is equivalent to the camera device mentioned above; the video cloud server is used to store videos; the client refers to the application installed on the user's device, such as a video APP.
[0119] In the first step, the user selects a specific device through the client, sets the content to be focused on (or the algorithm to be run, such as the pet detection algorithm), the time period for the miniature video to take effect, the background music for the miniature video, and the duration of the miniature video. Figure 8 The settings in the micro-strategy.
[0120] In the second step, the video miniature server sends the miniature configuration to the end-side AI camera, and the end-side AI camera dynamically downloads the relevant algorithm plug-in (equivalent to Figure 8 The end-side AI camera determines whether to start the relevant algorithm plug-in based on the time period of the microcosm set by the user: when the time point of the microcosm set by the user is reached, the end-side AI camera starts the algorithm plug-in to analyze the real-time camera image. According to the device resources and algorithm accuracy requirements, the video can be analyzed by frame extraction (generally, for a 25-frame image, about 10 frames of data can be extracted for intelligent analysis), and the target detection result and video clip that appear for the first time are reported, and the target detection result that has not been lost is not reported repeatedly. The target detection result contains the detected target information, such as the detection frame coordinates, target tracking ID, target category, time point and confidence level. The video clip is an H264 or H265 stream file of about 2s before and after the event. This step corresponds to Figure 8 Detection event reporting and playback video upload.
[0121] Step 3: After receiving the algorithm detection results reported by the client and the cloud playback video clips, the video miniature server starts the relevant frame extraction service and the corresponding cloud-side algorithm (equivalent to Figure 8 The video epitome server determines whether to conduct further analysis based on the target confidence in the detection results reported by the end-side AI camera: If the detection target confidence does not reach the threshold set by the video epitome server, the frame extraction service extracts the video frames in the corresponding material video and sends them to the inference service for analysis. If the video epitome server still does not meet the accuracy requirements after analysis, the video epitome server discards the relevant data.
[0122] In the fourth step, when the epitome time is over, the video epitome server obtains video clips from the cloud playback material library that meets the conditions for processing. According to the algorithm results after filtering in the previous step, the reported events are summarized and a summary report of related events is generated. The report mainly includes the number of events in this period, the time period of high-frequency events, and the comparison report with previous periods. And according to the epitome video duration set by the user, the total duration of qualified materials is calculated. If the number and duration of material clips exceed the user-set duration, the following elimination strategy can be adopted to ensure the uniformity of material distribution: sort the materials by time, add them to the material list, calculate the middle time point of the epitome start time and end time, divide the nodes in the above material list into two parts according to the time point, continue to select the data on the side with a larger number to continue the above process until a group of materials remains in the interval, and delete the material node from the list. According to the above elimination strategy, the materials are eliminated to ensure that the final composite video duration meets the standard.
[0123] In the fifth step, the video miniature server synthesizes the final material files, sorts the material files, selects two adjacent material videos, splices them, and adds fade-in and fade-out filters. After the splicing is completed, the splicing result file is retained, and the result file is then spliced and merged with other materials in turn until the material list is empty. Finally, the result file is packaged with the audio material set by the user to generate a final miniature video of several minutes in length and stored in the cloud storage server. This step is equivalent to Figure 8 The video in is read and synthesized.
[0124] Step 6: The video thumbnail server pushes a message to the user client, so that the user can quickly view the thumbnail video of the content he / she is interested in during the specified time period and the related summary report. Figure 8 View in miniature video.
[0125] In summary, the end side generates material files based on the corresponding end-side algorithm according to user needs. Cloud testing only processes a small number of end-side material files, so the cost of video analysis miniaturization on the cloud side is greatly reduced. In addition, the cloud side will automatically eliminate redundant materials according to the uniformity of time distribution, so that the effect of miniaturized video is better when the duration is met. This application combines the end-side analysis capabilities of intelligent AI cameras, dynamically sends algorithm plug-ins to detect events of interest to users, and uploads relevant materials to the cloud side. The cloud side filters, edits and synthesizes the video materials reported by the end side, and generates a miniature compilation of a few minutes, which is convenient for users to quickly understand what happened in a day, eliminating the tedious steps of video playback.
[0126] The embodiment of the present application further provides a video epitome method, which is applied to a camera device. In the method, the camera device performs the following steps:
[0127] Receive the epitome configuration information sent by the cloud server according to the device identifier, where the epitome configuration information includes the epitome effective time period and the algorithm identifier.
[0128] During the effective period of the miniature, the miniature algorithm corresponding to the algorithm identification is used to process the continuous video frames to obtain the recognition results obtained from the continuous video frames output by the miniature algorithm, and the target segment corresponding to each recognition result is intercepted from the continuous video frames.
[0129] The multiple target segments are sent to the cloud server so that the cloud server can synthesize the multiple target segments to generate a target miniature video.
[0130] The embodiment of the present application also provides a video epitome method, which is applied to a cloud server. In the method, the cloud server performs the following steps:
[0131] According to the device identifier, the epitome configuration information is sent to the camera device specified by the user, and the epitome configuration information includes the epitome effective time period and the algorithm identifier. The target segment sent by the camera device is received, and the target segment is the camera device in the epitome effective time period, using the epitome algorithm corresponding to the algorithm identifier to identify the continuous video frames, and the video segment corresponding to each recognition result is intercepted from the continuous video frames according to the recognition result.
[0132] Receive multiple target segments uploaded by a camera device, synthesize the multiple target segments, and generate a target miniature video.
[0133] Fig. 9 A structural block diagram of a video miniature device provided in an embodiment of the present application is shown. Fig. 9 As shown, the device comprises:
[0134] A receiving module 91 is used to receive the epitome configuration information sent by the cloud server according to the device identifier, where the epitome configuration information includes the epitome effective time period and the algorithm identifier;
[0135] The processing module 92 is used to process the continuous video frames by using the corresponding epitome algorithm through the algorithm identification within the epitome effective time period, obtain the recognition results obtained from the continuous video frames output by the epitome algorithm, and intercept the target segment corresponding to each recognition result from the continuous video frames;
[0136] The sending module 93 is used to send the multiple target segments to the cloud server so that the cloud server can synthesize the multiple target segments to generate a target miniature video.
[0137] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0138] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and will not be repeated in this embodiment.
[0139] In addition, in combination with the method provided in the above embodiments, a storage medium may be provided in this embodiment to implement the method. The storage medium stores a computer program; when the computer program is executed by a processor, any one of the video miniature methods in the above embodiments is implemented.
[0140] The embodiment of the present application also provides a computer program product. When the computer program product is run on a computer, the computer executes each function or step executed by the processor in the above method embodiment.
[0141] It should be understood that the specific embodiments described herein are only used to explain the application, rather than to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of this application.
[0142] Obviously, the drawings are only some examples or embodiments of the present application. For ordinary technicians in the field, the present application can also be applied to other similar situations based on these drawings without creative work. In addition, it is understandable that although the work done in this development process may be complicated and lengthy, for ordinary technicians in the field, certain changes in design, manufacturing or production based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient content disclosed in this application.
[0143] The term "embodiment" in this application refers to a specific feature, structure or characteristic described in conjunction with the embodiment that can be included in at least one embodiment of the present application. The appearance of this phrase in various locations in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. It is clearly or implicitly understood by those of ordinary skill in the art that the embodiments described in this application can be combined with other embodiments without conflict.
[0144] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of patent protection. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the attached claims.
Claims
1. A video miniature method, characterized in that: Applied to a cloud server, the method comprises: According to the device identifier, the epitome configuration information is sent to the camera device specified by the user, wherein the epitome configuration information includes the epitome effective time period and the algorithm identifier; Receive the target segment sent by the camera device; the target segment is a video segment corresponding to each recognition result intercepted from the continuous video frames by the camera device during the time period when the epitome is effective, using the epitome algorithm corresponding to the algorithm identifier to identify the continuous video frames; wherein the epitome algorithm refers to an algorithm for intercepting video segments from a video; A target miniature video is generated by synthesizing a plurality of the target segments.
2. The video epitome method according to claim 1, characterized in that: Before synthesizing the plurality of target fragments, the method further comprises: Receiving a recognition result sent by the camera device; the recognition result includes a first confidence level of the target object recognized by the camera device; In the case where the recognition result does not meet the preset accuracy, the target segment corresponding to the recognition result is recognized by the epitome algorithm to obtain an analysis result; the analysis result includes a second confidence level of the target object recognized by the cloud server; If the analysis result meets the preset accuracy, the target segment corresponding to the recognition result is retained; if the analysis result does not meet the preset accuracy, the target segment corresponding to the recognition result is discarded.
3. The video epitome method according to claim 2, characterized in that: When the recognition result does not meet the preset accuracy, the target segment corresponding to the recognition result is recognized by the miniature algorithm to obtain an analysis result, including: If the first confidence level is less than a confidence level threshold preset by the cloud server, identifying the target segment corresponding to the recognition result by using the miniature algorithm to obtain the analysis result; If the analysis result satisfies the preset accuracy, retaining the target segment corresponding to the recognition result; if the analysis result does not satisfy the preset accuracy, discarding the target segment corresponding to the recognition result, includes: If the second confidence is greater than the confidence threshold, the target segment corresponding to the recognition result is retained; if the second confidence is less than or equal to the confidence threshold, the target segment corresponding to the recognition result is discarded.
4. The video epitome method according to any one of claims 1 to 3, characterized in that: The miniature configuration information also includes the duration of the miniature video; the method also includes: Calculating the duration of an initial miniature video after synthesizing the plurality of target segments; If the duration of the initial miniature video is greater than the duration of the miniature video, then based on a preset elimination strategy, the target segment among the multiple target segments is eliminated; The multiple target segments remaining after elimination are synthesized to generate a target miniature video, wherein the duration of the target miniature video is equal to or less than the duration of the miniature video.
5. The video epitome method according to claim 4, characterized in that: The step of eliminating a target segment from the plurality of target segments based on a preset elimination strategy includes: Based on the miniature effective time period, determining a miniature time midpoint; Classifying the multiple target segments according to the midpoints of the miniature time to obtain two sets of target segment sets; For a target segment set having a larger number of target segments in the two groups of target segment sets: if the number of target segments in the target segment set is greater than 1, the target segment set is classified according to the time midpoint of the target segment set until one target segment remains in the target segment set; and the remaining target segments in the target segment set are eliminated.
6. The video epitome method according to claim 1, characterized in that: The method further comprises: Determine the device identification of the camera device according to the user's selection operation on the camera device through the client, and determine the miniature algorithm according to the focus content set by the user; The miniature configuration information is generated according to the miniature effective time period, miniature background music and miniature video duration set by the user through the client, combined with the algorithm identifier corresponding to the miniature algorithm.
7. The video epitome method according to claim 6, characterized in that: The method further comprises: Receiving the recognition result reported by the camera device; Generate an analysis report based on the number of occurrences of various events, the time period of high-frequency events, and previous period comparison data within the effective time period of the epitome analyzed from the recognition results; The step of synthesizing the plurality of target segments to generate a target miniature video includes: Synthesizing a plurality of the target clips, and packaging the synthesis result with the miniature background music to generate a target miniature video; A message is pushed to the client to instruct the user to view the target thumbnail video and analysis report.
8. A video miniature method, characterized in that: Applied to a camera device, the method comprises: Receiving the epitome configuration information sent by the cloud server according to the device identifier, wherein the epitome configuration information includes the epitome effective time period and the algorithm identifier; During the effective period of the epitome, the epitome algorithm corresponding to the algorithm identifier is used to process the continuous video frames to obtain the recognition results output by the epitome algorithm from the continuous video frames, and the target segments corresponding to each recognition result are intercepted from the continuous video frames; wherein the epitome algorithm refers to an algorithm for intercepting video segments from a video; The multiple target segments are sent to the cloud server so that the cloud server synthesizes the multiple target segments to generate a target miniature video.
9. A video miniature system, characterized in that: The system includes a cloud server and a camera device, wherein the cloud server is configured to execute the video miniature method described in any one of claims 1 to original claim 7, and the camera device is configured to execute the video miniature method described in claim 8.
10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the video miniature method according to any one of claims 1 to original claim 7, or execute the video miniature method according to claim 8 when running.
Citation Information
Patent Citations
Imaging apparatus, server, control program therefor, computer readable recording medium which records the control program, event management system and control method
JP2008154100A
System and method for security management based on face recognition in CCTV environment
KR1020140089810A