A partition-based video slice storage method and apparatus
By analyzing real-time video timestamps and fixed video sizes using the YOLO algorithm, and combining video slice timestamps and thresholds to construct a video partitioning and mapping model, the problems of low access efficiency and slow retrieval speed in video slice storage are solved, achieving the effects of rapid location and reduced loss risk.
Patent Information
- Application Number
- CN202411553440.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-02
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2044-11-02
AI Technical Summary
In existing technologies, the storage of video slices in large-scale live video applications suffers from problems such as low access efficiency, slow retrieval speed, and insufficient slicing speed and access efficiency.
The YOLO algorithm, based on real-time video timestamps and fixed video size, is used to analyze and construct video slices. The video segmentation and mapping model is constructed by combining the timestamps and thresholds of the video slices. The video slice data is parsed and the storage location is determined by using message middleware.
It enables rapid location of video time points, allowing for faster identification of feature events, reducing the risk of video clip loss, and optimizing storage resource scheduling and usage.
Smart Images

Figure CN119629384B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of data storage technology, and particularly relates to a video slice storage method and device based on partition. BACKGROUND
[0002] In the prior art, video slices are used to improve video playback efficiency, facilitate video editing and processing, and facilitate network transmission. By splitting a video file into smaller segments, the video can be loaded and played on demand, thereby improving the loading speed and playback efficiency of the video. In addition, the slicing process can reduce the computational load and avoid network congestion caused by the large size of a single transmission file.
[0003] Video live streaming applications with a user base of over one million require storage and access of massive amounts of video slice data. Although distributed and data slicing technologies can be used for data processing, a single video slice index still affects the speed of data access. In addition, there is still much room for improvement in terms of slice speed, access efficiency, and access efficiency.
[0004] Therefore, it is necessary to provide a video slice storage method and device based on partition to solve the above problems. SUMMARY
[0005] The application aims to provide a video slice storage method and device based on partition to solve the technical problems of low access efficiency, slow retrieval speed, how to improve slice speed, and how to improve access efficiency in the video monitoring service of the prior art. The technical problems to be solved by the application are solved by the following technical solutions.
[0006] The first aspect of the application provides a video slice storage method based on partition, which comprises: based on the timestamp and fixed video size of real-time video, using a YOLO algorithm to analyze the video to be processed to construct video slices and determine the weight corresponding to each video slice; based on the timestamp and threshold of the video slice, constructing a video partition and video slice mapping model based on a timeline; using a message middleware to obtain a video message to be processed, and parsing video slice data from the video message; comparing the parsed video slice data and the timestamp of the video partition according to the constructed mapping model to determine the storage location.
[0007] The second aspect of the present application provides a video slice storage device based on partitioning, which adopts the video slice storage method based on partitioning in the first aspect of the present application. The video slice storage device comprises: a first construction module, which uses a YOLO algorithm to analyze a to-be-processed video to construct video slices and determine the weights corresponding to the video slices based on the timestamp of the real-time video and the fixed video size; a second construction module, which constructs a mapping model of video partitioning and video slices based on a timeline based on the timestamp of the video slices and a threshold value; an analysis processing module, which acquires a to-be-processed video message by using a message middleware and analyzes the video slice data from the video message; and a storage module, which compares the analyzed video slice data and the timestamp of the video partitioning to determine the storage location according to the constructed mapping model.
[0008] The third aspect of the present application provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs; and when the one or more programs are executed by the one or more processors, the one or more processors implement the video slice storage method based on partitioning in the first aspect of the present application.
[0009] The fourth aspect of the present application provides a computer readable medium having a computer program stored thereon, and the computer program is executed by a processor to implement the video slice storage method based on partitioning in the first aspect of the present application.
[0010] The embodiments of the present application have the following advantages:
[0011] Compared with the prior art, the present application uses a YOLO algorithm to analyze a to-be-processed video to construct video slices based on the timestamp of the real-time video and the fixed video size, constructs a mapping model of video partitioning and video slices based on a timeline based on the timestamp of the video slices and a threshold value, acquires a to-be-processed video message by using a message middleware, analyzes the video slice data from the video message, compares the analyzed video slice data and the timestamp of the video partitioning to determine the storage location, can realize fast positioning of the video time point, can locate the feature event faster, and reduces the loss risk of the video segment. The dynamic threshold value can effectively balance the overall weight of each partition, and ensures reasonable scheduling and use of storage resources.
[0012] In addition, by introducing the video slice partitioning mapping model, the data storage and access speed are accelerated. Based on the video slices with shorter time and fixed size, the video time point can be positioned quickly, especially the feature event can be positioned quickly, and the loss risk of the video segment is reduced. The dynamic threshold value can effectively balance the overall weight of each partition, and ensures reasonable scheduling and use of storage resources. BRIEF DESCRIPTION OF DRAWINGS
[0013] Figure 1 is a step flow chart of an example of the partition-based video slice storage method of the present application;
[0014] Figure 2 is a schematic diagram of video slice partitioning in the partition-based video slice storage method of the present application;
[0015] Figure 3 is a structural block diagram of the partition-based video slice storage device of the present application;
[0016] Figure 4 is a structural schematic diagram of an electronic device embodiment according to the present application;
[0017] Figure 5 is a structural schematic diagram of a computer readable medium embodiment according to the present application. DETAILED DESCRIPTION
[0018] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0019] The present application will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0020] In view of the above problems, the present application proposes a partition-based video slice storage method, which uses a YOLO algorithm to analyze a to-be-processed video based on timestamps of real-time videos and fixed video sizes to construct video slices, constructs a timeline-based video partitioning and video slice mapping model based on timestamps of the video slices and a threshold, uses a message middleware to obtain a to-be-processed video message, parses video slice data from the video message, and compares the parsed video slice data and timestamps of the video partitioning according to the constructed mapping model to determine a storage location, which can realize fast positioning of a video time point, reduce the risk of loss of video segments, and locate a feature event faster. The dynamic threshold can effectively balance the overall weight of each partitioning, ensuring reasonable scheduling and use of storage resources.
[0021] Embodiment 1
[0022] The content of the present application will be described in detail below with reference to Figure 1 , Figure 2 .
[0023] Figure 1 is a step flow chart of an example of the partition-based video slice storage method of the present application; Figure 2 is a schematic diagram of video slice partitioning in the partition-based video slice storage method of the present application;
[0024] As Figure 1As shown, in step S101, based on the timestamp of the real-time video and the fixed video size, the YOLO algorithm is used to analyze the to-be-processed video to construct video slices, and determine the weights corresponding to each video slice.
[0025] Specifically, the real-time video and the to-be-processed video are, for example, monitoring videos for monitoring indoor areas of a home, a home courtyard, and an internal site of a factory.
[0026] The to-be-processed video is analyzed, for example, as follows: Figure 2 As shown, the method specifically comprises the following steps:
[0027] Step S201: Extract key frame images in the to-be-processed video.
[0028] According to the human shape data, the animal shape features, and the article shape features, one or more key frame images in the to-be-processed video are extracted, the key frame images contain specified targets, and the specified targets include a person (for example, a specific user, specifically a child, an old person, etc.), an animal (for example, a cat, a dog, a rabbit), or a specific article (for example, a backpack, a book, a table lamp, etc.).
[0029] Step S202: Establish a target detection model using the YOLO algorithm, input one or more key frame images of the to-be-processed video into the established target detection model, and output the specified targets and the time points corresponding to each specified target, the specified targets being a person, an animal, or a specific article.
[0030] The target detection model is established using the YOLO algorithm.
[0031] One or more images (or a segment of video) labeled with the specified targets and the positions of each specified target relative to a predetermined point (for example, the center of a detection box or any vertex) are used to establish a training data set to train the target detection model. For example, a classification label is used to represent whether the to-be-processed video contains at least one specified target, so as to identify whether the to-be-processed video contains the specified target.
[0032] The target detection model comprises an input layer, a convolution layer, a pooling layer, and a classification layer.
[0033] The convolution operation is performed in the convolution layer of the target detection model. Specifically, the convolution operation uses the ReLU activation function to increase nonlinearity and uses batch normalization to reduce the bias of the data distribution of the to-be-processed video or the to-be-processed image (for example, one or more key frame images) containing the specified target.
[0034] The pooling operation is performed in the pooling layer of the target detection model, and the pooling operation is used to reduce the size of the feature map containing the specified target, so as to improve the detection efficiency and accuracy of the model. The present application uses maximum pooling and average pooling.
[0035] P = {f(i,j) / Σ(i,j)}
[0036] Where P is the pooled feature map containing the specified target, f(i,j) is the pixel value in the input feature map, and Σ(i,j) is the sum of all pixel values in the feature map (for example, max pooling will take the local maximum value instead of the average value).
[0037] Optionally, the established object detection model can be optimized by reducing the classification loss and localization loss:
[0038] Preferably, a deconvolution operation is used to upsample the feature map containing the specified target to recover the original size and resolution of the specified target, and is expressed by the following expression:
[0039] X=W^T*P*S^(-1)+b
[0040] Where X represents the feature map containing the specified target output after the deconvolution operation; W is the weight matrix for the deconvolution operation; P is the pooled feature map containing the specified target; S is the upsampling size; and b is the bias term.
[0041] By reducing classification loss and localization loss, the established target detection model can be optimized, which can effectively reduce the model prediction error.
[0042] When the trained object detection model is applied, one or more keyframes of the video to be processed (or a segment of the video to be processed at a specified time) are input, and the output is the specified target corresponding to the video to be processed and the time point corresponding to each specified target.
[0043] Next, when the specified target corresponding to the video to be processed and the time point corresponding to each specified target are output, the newly added video slice is determined.
[0044] By using the object detection model built with the YOLO v8 algorithm, the key video segments in the video to be processed are intelligently segmented. For example, when a specified target (such as a person or animal) appears in the video frame, the video is sliced, which can accurately segment the video to be processed.
[0045] Specifically, when the following specified targets appear in the video to be processed, it is defined as a feature event occurring in the video to be processed. The start time of the appearance of the following specified targets in the video to be processed is defined as the starting point of the feature event, and the time when the following specified targets disappear from the video to be processed, i.e., the end time of the appearance of the following specified targets in the video to be processed, is defined as the ending point of the feature event. The specified targets include people (e.g., specific users, such as children, the elderly, etc.), animals (e.g., cats, dogs, rabbits), or specific items (e.g., backpacks, books, desk lamps, etc.).
[0046] In one specific implementation, the target detection model inputs the video to be processed within a specified time period, i.e., multiple images to be processed, and outputs the specified target corresponding to the video to be processed, the start time of each specified target, and the end time.
[0047] Specifically, the start and end points of the feature events are used as the start and end points of the newly added video partitions, respectively.
[0048] It should be noted that the above is only an optional example and should not be construed as a limitation of the present invention.
[0049] Next, in step S102, a timeline-based mapping model of video partitions and video slices is constructed based on the timestamps and thresholds of the video slices.
[0050] Specifically, a fixed duration is allocated for slicing to form video slices, and the formed video slices are added to the video slice collection.
[0051] Based on the start time, end time, and dynamic threshold, video partitions are constructed to form a set of video partitions. Each video partition includes a partition weight and a maximum time span.
[0052] For constructing video slices, the following symbols are defined:
[0053] Tmax: Maximum slice duration (in seconds); T: Slice duration (in seconds);
[0054] Rmax: Maximum size of the slice; R(t): Size of the video segment from the start of the video to time t seconds; R.weight: Weight of the video slice.
[0055] Find a time T in a video segment such that the size of the video segment at time T does not exceed Rmax, while being as close as possible to Tmax. Define such a video segment as a video slice.
[0056] The video slice is represented as R(T), where T is represented as T = max{t | t ≤ Tmax and R(t) ≤ Rmax}.
[0057] Based on the timestamps and thresholds of video slices, a timeline-based mapping model for video partitions and video slices is constructed.
[0058] Based on the start and end times of video segments, a timeline-based mapping model is constructed between video sections and raw video slices. Each video section contains several raw video slices.
[0059] It should be noted that the mapping model can be regarded as a function that maps the set of video slices R to the set of video partitions S, while satisfying that the time span within each video partition does not exceed STmax, where STmax represents the maximum allowed time span of the video partition, i.e., the static threshold of the video partition, such as the maximum time span of the video partition being 1 hour.
[0060] Specifically, the set SR of video slices under video partition Si satisfies the following expression:
[0061]
[0062] Where Rj represents the j-th video slice, j is a positive integer, SR represents the set of video slices under video partition Si, i is a positive integer; Rj.start represents the start time of the j-th video slice; Si.start represents the start time of the i-th video partition to which the j-th video slice belongs; Rj.end represents the end time of the j-th video slice; Si.end represents the end time of the i-th video partition to which the j-th video slice belongs.
[0063] The video slice time span D(Si) satisfies the following expression:
[0064] D(Si)=max{Rj.end|Rj∈SR}-min{Rk.start|Rk∈SR)≤STmax
[0065] D(Si) represents the time span of the video partition, which is equal to the maximum end time of all video slices in the video partition minus the minimum start time of all video slices. Rj.end represents the end time of the j-th video slice, where Rj represents the j-th video slice; SR represents the set of video slices under video partition Si, where i is a positive integer; Rk.start represents the start time of the k-th video slice, where Rk represents the k-th video slice; STmax represents the maximum allowed time span of the video partition, i.e., the static threshold of the video partition, such as a maximum time span of 1 hour.
[0066] Based on the mapping model for constructing video sections and raw video slices, the location of the newly added video section, i.e., the storage location, is determined.
[0067] When a new video partition is determined, the average weight of the current video partitions is used as the dynamic threshold for partition division. A partition weight averaging function is employed to determine the weight of the new video partition, where the new video partition S... x The weight is less than or equal to the average weight of the current video partition:
[0068] S x (weight)≤Avg(weight)
[0069] Among them, S x (weight) indicates the newly added video partition S x The weights; Avg(weight) represents the average weight of the current video partition;
[0070] The average weight of all current video partitions:
[0071]
[0072] Where Avg(weight) represents the average weight of all current video partitions; S j (weight) is the weight function for the j-th video partition, m represents the number of all video partitions in the current storage system, and j is a positive integer, specifically 1, 2, ..., m.
[0073] Further, each video partition comprises multiple video slices, and the weight of the current video partition is calculated using the following expression:
[0074]
[0075] Among them, S 当 R represents the weight of the current video partition; i .weight represents the weight value of the i-th video slice in the current video partition, R i .factor represents the influence factor of the i-th video slice in the current video partition, n represents the number of video slices in the current video partition, and i is a positive integer, specifically 1, 2, ..., n.
[0076] By using the average weight of all current video partitions as the dynamic threshold for partition division, the overall weight of each partition can be effectively balanced, ensuring reasonable scheduling and use of storage resources.
[0077] It should be noted that the above is only an optional example and should not be construed as a limitation of the present invention.
[0078] Next, in step S103, the message middleware is used to obtain the video message to be processed, and the video slice data is parsed from the video message.
[0079] Specifically, the message middleware is used to obtain the video messages to be processed, and the video message type, start and end times of the video slices are parsed from the video messages.
[0080] For example, if the video to be processed is identified as a home indoor video, the analysis of the target detection model's output video message corresponding to the video to be processed will be performed. Specifically, this includes the designated target corresponding to the video to be processed, the start time of each designated target, and the end time. Furthermore, it can be determined whether the video message contains feature events.
[0081] When a video message to be processed contains feature events, the start and end times of the video slice are determined based on the feature events, or the start and end times of the video slice containing feature events are directly parsed out.
[0082] Furthermore, based on the time points corresponding to each specified target output, and according to the constraints of duration and video size, video slicing is performed to complete the addition of new video slices.
[0083] It should be noted that the above is only an optional example and should not be construed as a limitation of the present invention.
[0084] Next, in step S104, the timestamps of the parsed video slice data and video partitions are compared according to the constructed mapping model to determine the storage location.
[0085] Specifically, the storage system (e.g., the video slice storage device of Embodiment 2) includes a left partition and a right partition, wherein the start time of the video partition in the left partition is earlier than the start time of the newly added video slice, and the start time of the video partition in the right partition is later than the start time of the newly added video slice.
[0086] Based on the mapping model constructed in step S102, and based on the parsed video slice data and the timestamp of the video partition, a mapping judgment is made to determine the location of the newly added video slice and the location to be stored (also referred to as the "storage location").
[0087] When the start time and end time of a newly added video slice are both greater than the end time of the left partition, and there is no right partition, a new video partition is directly inserted, and a new video slice is directly inserted into the new video partition.
[0088] When the start time of a new video slice is between the start and end times of the left partition, and the end time of the new video slice is greater than the end time of the left partition, the end time of the left partition is updated to the end time of the new video slice, that is, the new video slice is mapped to the left partition.
[0089] When the start and end times of a newly added video slice are both between the start and end times of the left partition, the newly added video slice is directly mapped to the left partition.
[0090] In the case of a right partition, if the start time and end time of the newly added video slice are both greater than the end time of the left partition, and a right partition exists, it is further determined whether the start time of the right partition is greater than the end time of the newly added video slice. If the start time of the right partition is greater than the end time of the newly added video slice, a new video partition is added between the left and right partitions, and a new video slice is inserted into the new video partition.
[0091] When the start time of the right partition is between the start and end times of the newly added video slice, the start time of the right partition is updated to the start time of the newly added video slice, which is equivalent to mapping the newly added video slice to the right partition.
[0092] In one alternative implementation, video partitioning is performed based on characteristic events to determine new video partitions.
[0093] Specifically, the target detection model inputs the video to be processed within a specified time period, i.e., multiple images to be processed, and outputs the specified target corresponding to the video to be processed, the start time and end time corresponding to each specified target.
[0094] When the following specified targets appear in the video to be processed, it is defined as a feature event occurring in the video. The start time of the appearance of the specified targets in the video is defined as the start time of the feature event, and the time when the specified targets disappear from the video, i.e., the end time of the appearance of the specified targets in the video, is defined as the end time of the feature event. The specified targets include people (e.g., specific users, such as children, the elderly, etc.), animals (e.g., cats, dogs, rabbits), or specific items (e.g., backpacks, books, desk lamps, etc.), and the start and end times of the feature event are respectively used as the start and end times of the newly added video partition.
[0095] In one specific implementation, based on whether the output specified target is a person, animal or specific item, it is determined whether it contains a feature event. The video to be processed containing a feature event is divided into a video slice, and according to the mapping model constructed in step S102, a mapping judgment is performed to determine the location of the new video slice and the location to be stored.
[0096] In another specific implementation, based on whether the output specified target is a person, animal or specific item, it is determined whether it contains feature events. The video to be processed containing two or more feature events is divided into a video slice, and according to the mapping model constructed in step S102, a mapping judgment is performed to determine the location of the newly added video slice and the location to be stored.
[0097] Next, storage is performed according to the determined storage location, which enables rapid location of video time points and faster identification of feature events. At the same time, it reduces the risk of video segment loss. Dynamic thresholds can effectively balance the overall weight of each partition, ensuring reasonable scheduling and use of storage resources.
[0098] It should be noted that the above is only an optional example and should not be construed as a limitation of the present invention.
[0099] Furthermore, the accompanying drawings are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes shown in the drawings do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0100] Compared with existing technologies, this invention uses real-time video timestamps and fixed video size, employs the YOLO algorithm to analyze the video to be processed to construct video slices, and builds a timeline-based video partitioning and video slice mapping model based on the timestamps and thresholds of the video slices. It utilizes message middleware to obtain the video messages to be processed, parses video slice data from the video messages, and compares the parsed video slice data with the timestamps of the video partitions according to the constructed mapping model to determine the storage location. This enables rapid location of video time points, faster identification of feature events, and reduces the risk of video segment loss. Dynamic thresholds effectively balance the overall weight of each partition, ensuring reasonable scheduling and use of storage resources.
[0101] Furthermore, by introducing a video slice partitioning mapping model, data storage and access speeds are accelerated. Based on short-duration, fixed-size video slices, rapid location of video time points can be achieved, while reducing the risk of video segment loss and enabling faster identification of feature events. Dynamic thresholds effectively balance the overall weight of each partition, ensuring reasonable scheduling and use of storage resources.
[0102] Example 2
[0103] The following are embodiments of the apparatus of the present invention, which can be used to execute embodiments of the method of the present invention. For details not disclosed in the embodiments of the apparatus of the present invention, please refer to the embodiments of the method of the present invention.
[0104] Figure 3 This is a schematic diagram of an example of a partition-based video slice storage device according to the present invention.
[0105] Reference Figure 3 The second aspect of this disclosure provides a partition-based video slice storage device 300, which employs the partition-based video slice storage method described in the first aspect of this invention.
[0106] The partition-based video slice storage device 300 includes: a first construction module 310, a second construction module 320, a parsing and processing module 330, and a storage module 340.
[0107] In one specific implementation, the first construction module 310 analyzes the video to be processed using the YOLO algorithm based on the timestamp of the real-time video and a fixed video size to construct video slices and determine the weights corresponding to each video slice. The second construction module 320 constructs a timeline-based mapping model between video partitions and video slices based on the timestamps and thresholds of the video slices. The parsing and processing module 330 uses a message middleware to obtain the video message to be processed and parses the video slice data from the video message. The storage module 340 compares the parsed video slice data with the timestamps of the video partitions according to the constructed mapping model to determine the storage location.
[0108] According to the optional implementation method, key frame images are extracted from the video to be processed, and a target detection model is established using the YOLO algorithm. One or more key frame images of the video to be processed are input into the established target detection model, and a specified target and the time point corresponding to each specified target are output. The specified target is a person or an animal. When the specified target and the time point corresponding to each specified target are output, a new video slice is determined.
[0109] According to an optional implementation, when a new video partition is determined, the average weight of the current video partitions is used as the dynamic threshold for partition division. A partition weight average function is then used to determine the weight of the new video partition, where the new video partition S... x The weight is less than or equal to the average weight of the current video partition:
[0110] S x (weight)≤Avg(weight)
[0111] Among them, S x (weight) indicates the newly added video partition S x The weights; Avg(weight) represents the average weight of the current video partition.
[0112] The average weight of all current video partitions:
[0113]
[0114] Where Avg(weight) represents the average weight of all current video partitions; S j (weight) is the weight function for the j-th video partition, m represents the number of all video partitions in the current storage system, and j is a positive integer, specifically 1, 2, ..., m.
[0115] According to the optional implementation, each video partition includes multiple video slices, and the weight of the current video partition is calculated using the following expression:
[0116]
[0117] Among them, S 当 R represents the weight of the current video partition; i .weight represents the weight value of the i-th video slice in the current video partition, R i .factor represents the influence factor of the i-th video slice in the current video partition, n represents the number of video slices in the current video partition, and i is a positive integer, specifically 1, 2, ..., n.
[0118] Based on the constructed mapping model, the parsed video slice data and the timestamps of the video partitions are compared.
[0119] Specifically, the storage system includes a left partition and a right partition, wherein the start time of the video partition in the left partition is earlier than the start time of the newly added video slice, and the start time of the video partition in the right partition is later than the start time of the newly added video slice.
[0120] Based on the mapping model, a mapping judgment is performed to determine the position of the new video slice. Specifically, when the start time and end time of the new video slice are both greater than the end time of the left partition, and there is no right partition, a new video partition is directly inserted, and a new video slice is directly inserted into the new video partition. When the start time of the new video slice is between the start and end times of the left partition, and the end time of the new video slice is greater than the end time of the left partition, the end time of the left partition is updated to the end time of the new video slice, that is, the new video slice is mapped to the left partition. When the start time and end time of the new video slice are both between the start and end times of the left partition, the new video slice is directly mapped to the left partition.
[0121] According to the optional implementation, when the start time and end time of the newly added video slice are both greater than the end time of the left partition, and a right partition exists, it is further determined whether the start time of the right partition is greater than the end time of the newly added video slice. When the start time of the right partition is greater than the end time of the newly added video slice, a new video partition is added between the left partition and the right partition, and a new video slice is inserted into the new video partition. When the start time of the right partition is between the start time and the end time of the newly added video slice, the start time of the right partition is updated to the start time of the newly added video slice, which is equivalent to mapping the newly added video slice to the right partition.
[0122] According to an optional implementation, video partitioning is performed based on feature events to determine new video partitions; wherein, the video to be processed within a specified time period, i.e., multiple images to be processed, is input into the target detection model, and the model outputs the specified target corresponding to the video to be processed, the start time of each specified target, and the end time; when the following specified targets appear in the video to be processed, it is defined as the occurrence of a feature event in the video to be processed, the start time of the occurrence of the following specified targets in the video to be processed is defined as the starting point of the feature event, and the time when the following specified targets disappear in the video to be processed, i.e., the end time of the occurrence of the following specified targets in the video to be processed, is defined as the end point of the feature event: people, animals, or specific objects.
[0123] According to the optional implementation method, based on the time points corresponding to each specified target output, and according to the constraints of duration and video size, video slicing is performed to complete the addition of new video slices.
[0124] According to an optional implementation, the video slice storage method according to claim 1 is characterized in that, the step of constructing a timeline-based video partition and video slice mapping model based on the timestamp and threshold of the video slice includes: configuring a fixed duration for slicing to form video slices, and adding the formed video slices to a video slice set; constructing video partitions according to the start time, end time, and dynamic threshold to form a video partition set, wherein each video partition includes a partition weight and a configured maximum time span.
[0125] Figure 4 A schematic diagram of the structure of an electronic device according to an embodiment of the present invention.
[0126] like Figure 4 As shown, the electronic device is embodied in the form of a general-purpose computing device. There can be one or more processors working collaboratively. This invention also does not preclude distributed processing, meaning that processors can be distributed across different physical devices. The electronic device of this invention is not limited to a single entity, but can also be the sum of multiple physical devices.
[0127] The memory stores a computer-executable program, typically machine-readable code. The computer-readable program can be executed by the processor to enable the electronic device to perform the method of the present invention, or at least some steps of the method.
[0128] The memory includes volatile memory, such as random access memory (RAM) and / or cache memory, and may also be non-volatile memory, such as read-only memory (ROM).
[0129] Optionally, in this embodiment, the electronic device further includes an I / O interface for exchanging data with external devices. The I / O interface can represent one or more of several bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0130] It should be understood that Figure 4 The electronic device shown is merely one example of the present invention, and the electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as displays, and some electronic devices also include human-computer interaction elements such as buttons and keyboards. Any electronic device capable of executing a computer-readable program in memory to implement the method of the present invention or at least some steps of the method can be considered as an electronic device covered by the present invention.
[0131] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software, or by combining software with necessary hardware. Therefore, as... Figure 5 As shown, the technical solution according to the embodiments of the present invention can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.) or on a network, and includes several commands to cause a computing device (such as a personal computer, server, or network device, etc.) to execute the above-described method according to the embodiments of the present invention.
[0132] The software product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0133] The computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, capable of transmitting, propagating, or transmitting programs for use by or in connection with a command execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0134] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0135] The aforementioned computer-readable medium carries one or more programs (e.g., computer-executable programs) that, when executed by a device, cause the computer-readable medium to implement the methods of this disclosure.
[0136] Those skilled in the art will understand that the above modules can be distributed in the device as described in the embodiments, or they can be modified accordingly and placed in one or more devices that are unique to this embodiment. The modules in the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.
[0137] Through the description of the above embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions of the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, portable hard drive, etc.) or on a network, including several commands to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of the present invention.
[0138] Exemplary embodiments of the present invention have been specifically shown and described above. It should be understood that the present invention is not limited to the detailed structures, arrangements, or implementations described herein; rather, the present invention is intended to cover various modifications and equivalent arrangements contained within the spirit and scope of the appended claims.
Claims
1. A video slice storage method based on partitioning, characterized in that, The video slice storage method includes: Based on the timestamps of real-time video and a fixed video size, the YOLO algorithm is used to analyze the video to be processed to construct video slices, and the weights corresponding to each video slice are determined, including: Extract keyframe images from the video to be processed; A target detection model is established using the YOLO algorithm. One or more keyframe images of the video to be processed are input into the established target detection model, and the model outputs the specified target and the time point corresponding to each specified target. The specified target is a person, animal or specific object. When the output shows the specified target corresponding to the video to be processed and the time point corresponding to each specified target, a new video slice is determined; Based on the timestamps and dynamic thresholds of video slices, a timeline-based mapping model between video partitions and video slices is constructed. When a new video partition is determined, the average weight of the current video partition is used as the dynamic threshold for partition division. A partition weight averaging function is employed to determine the weight of the new video partition. The weight is less than or equal to the average weight of the current video partition: ; in, Indicates the addition of a new video partition The weights; This represents the average weight of all current video partitions; The average weight of all current video partitions: ; in, Let j be the weight function for the j-th video partition. This represents the number of video partitions in the current storage system, where j is a positive integer, specifically 1, 2, ... ; Each video partition consists of multiple video slices. The weight of the current video partition is calculated using the following expression: ; in, Indicates the weight of the current video partition; This represents the weight value of the i-th video slice in the current video partition. This represents the influence factor of the i-th video slice in the current video partition. This represents the number of video slices in the current video partition, where i is a positive integer, specifically 1, 2, ... ; Use a message broker to obtain video messages to be processed, and parse video slice data from the video messages; Based on the constructed mapping model, the timestamps of the parsed video slice data and video partitions are compared to determine the storage location.
2. The video slice storage method according to claim 1, characterized in that, Based on the constructed mapping model, the timestamps of the parsed video slice data and video partitions are compared to determine the storage location, specifically including: The storage system includes a left partition and a right partition. The start time of the video partition in the left partition is earlier than the start time of the newly added video slice, and the start time of the video partition in the right partition is later than the start time of the newly added video slice. Based on the mapping model, a mapping judgment is performed to determine the location of the new video slice, where... When the start time and end time of the newly added video slice are both greater than the end time of the left partition, and there is no right partition, directly insert a new video partition, and directly insert a new video slice into the new video partition. When the start time of a new video slice is between the start and end times of the left partition, and the end time of the new video slice is greater than the end time of the left partition, the end time of the left partition is updated to the end time of the new video slice, that is, the new video slice is mapped to the left partition. When the start and end times of a newly added video slice are both between the start and end times of the left partition, the newly added video slice is directly mapped to the left partition.
3. The video slice storage method according to claim 2, characterized in that, Further includes: When the start time and end time of the newly added video slice are both greater than the end time of the left partition, and a right partition exists, it is further determined whether the start time of the right partition is greater than the end time of the newly added video slice. If the start time of the right partition is greater than the end time of the newly added video slice, a new video partition is added between the left partition and the right partition, and a new video slice is inserted into the newly added video partition. When the start time of the right partition is between the start and end times of the newly added video slice, the start time of the right partition is updated to the start time of the newly added video slice, that is, the newly added video slice is mapped to the right partition.
4. The video slice storage method according to claim 1, characterized in that, Further includes: Video partitioning is performed based on characteristic events to determine new video partitions; among which... The target detection model takes a video (i.e., multiple images) within a specified time period as input and outputs a specified target corresponding to the video, along with the start and end times for each specified target. When the following specified targets appear in the video to be processed, it is defined as a feature event in the video to be processed. The start time of the appearance of the following specified targets in the video to be processed is defined as the start time of the feature event, and the time when the following specified targets in the video to be processed disappear, that is, the end time of the appearance of the following specified targets in the video to be processed, is defined as the end time of the feature event: people, animals or specific items.
5. A partition-based video slice storage device, characterized in that, The video slice storage method according to any one of claims 1 to 4 is wherein the video slice storage device comprises: The first construction module uses the YOLO algorithm to analyze the video to be processed based on the timestamp of the real-time video and the fixed video size to construct video slices and determine the weights corresponding to each video slice. The second building module constructs a timeline-based mapping model for video partitions and video slices based on the timestamps and thresholds of video slices. The parsing and processing module uses a message middleware to obtain the video message to be processed and parses the video slice data from the video message; The storage module compares the parsed video slice data with the timestamps of the video partitions based on the constructed mapping model to determine the storage location.
Citation Information
Patent Citations
Surveillance video slice storage method and surveillance video slice storage system
CN109361904A
Method and device for storing, retrieving and deleting embedded audio and video data, and memory
CN112558873A