Video storage method, device, electronic device and computer-readable medium

Through the same frequency extraction frames and image stitching of multi-directional video sets, combined with object recognition and cluster cluster center information generation model, the problem of inefficient video clip generation in the prior art is solved, and efficient and accurate video clip acquisition is achieved.

CN118827897BActive Publication Date: 2025-07-22ADDX (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410853534.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2025-07-22
Estimated Expiration
2044-06-28

AI Technical Summary

Technical Problem

When generating video clips corresponding to objects, relevant technical personnel need to detect one frame by one, resulting in inefficient cutting efficiency and waste of manpower and computing resources.

Method used

By obtaining multi-directional video sets, performing frame extraction and image stitching at the same frequency, using object identification information to generate models and cluster cluster center information generation models, generating image identification clusters and packaging processing, and storing them on the camera video storage end.

Benefits of technology

It realizes all-round acquisition of video clip object information, improves the cutting efficiency and accuracy, and reduces the waste of manpower and computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118827897B_ABST
    Figure CN118827897B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a video storage method, apparatus, electronic device, and computer-readable medium. A specific implementation of the method includes: obtaining a set of captured videos and at least one video segment object information; performing frame extraction at the same frequency on a first captured video, a second captured video, and a third captured video; splicing a first frame image sequence, a second frame image sequence, and a third frame image sequence in chronological order; inputting each spliced image into an object recognition information generation model to obtain an object recognition information sequence; setting corresponding image weights; generating at least one cluster center information; performing clustering processing to generate an image identification cluster set; performing packaging processing on the image identification cluster set, the spliced image sequence, and the image time sequence; storing the packaged file and the set of captured videos in a captured video storage end. This implementation can accurately and efficiently generate video segments corresponding to at least one video segment object information for the set of captured videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, and more particularly, to video storage methods, devices, electronic devices, and computer-readable media. Background Art

[0002] Currently, the technology of using a camera device to effectively monitor a target area has been widely applied in people's daily lives. For how to extract the video segment corresponding to an object from a captured video, the commonly adopted method is as follows: relevant technical personnel use video editing software to clip the video segment corresponding to the object.

[0003] However, when using the above method to generate the video segment corresponding to an object, the following technical problems often exist:

[0004] It is necessary for relevant technical personnel to detect frame by frame whether there is video content related to the object, and the situation of less clipping often occurs, resulting in low clipping efficiency and wasting a large amount of human resources and computer device resources.

[0005] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention

[0006] This summary of the disclosure is intended to introduce concepts in a brief form, which will be described in detail in the subsequent detailed description section. This summary of the disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.

[0007] Some embodiments of the present disclosure propose video storage methods, devices, electronic devices, and computer-readable media to solve one or more of the technical problems mentioned in the above background art section.

[0008] In a first aspect, some embodiments of the present disclosure provide a video storage method, including: obtaining a set of captured videos for a captured space and information of at least one video segment object, where the set of captured videos includes: a first captured video corresponding to a first azimuth camera device, a second captured video corresponding to a second azimuth camera device, and a third captured video corresponding to a third azimuth camera device; performing frame extraction at the same frequency on the first captured video, the second captured video, and the third captured video to generate a first frame image sequence, a second frame image sequence, and a third frame image sequence; splicing the first frame image sequence, the second frame image sequence, and the third frame image sequence in chronological order of the images to generate a spliced image sequence; inputting each spliced image in the spliced image sequence into a pre-trained object recognition information generation model to generate object recognition information and obtain an object recognition information sequence, where the object recognition information is recognition information associated with the information of the at least one video segment object; setting an image weight corresponding to each spliced image in the spliced image sequence according to the object recognition information sequence to obtain an image weight sequence; generating at least one cluster center information by using a cluster center information generation model according to the image weight sequence and the spliced image sequence; performing clustering processing on the spliced image sequence according to the at least one cluster center information to generate an image identification cluster set; performing packaging processing on the image identification cluster set, the spliced image sequence, and the image time sequence corresponding to the spliced image sequence to generate a packaged file; storing the packaged file and the set of captured videos at a captured video storage end for subsequent acquisition of video segments corresponding to the information of the video segment object.

[0009] Second aspect, some embodiments of the present disclosure provide a video storage device, including: an acquisition unit configured to acquire a set of captured videos for a captured space and information of at least one video segment object, wherein the set of captured videos includes: a first captured video corresponding to a first azimuth camera device, a second captured video corresponding to a second azimuth camera device, and a third captured video corresponding to a third azimuth camera device; a frame extraction unit configured to perform frame extraction processing at the same frequency on the first captured video, the second captured video, and the third captured video to generate a first frame image sequence, a second frame image sequence, and a third frame image sequence; a splicing unit configured to splice the first frame image sequence, the second frame image sequence, and the third frame image sequence in the same time order of the images to generate a spliced image sequence; an input unit configured to input each spliced image in the spliced image sequence into a pre-trained object recognition information generation model to generate object recognition information, obtaining an object recognition information sequence, wherein the object recognition information is recognition information associated with the information of the at least one video segment object; a setting unit configured to set an image weight corresponding to each spliced image in the spliced image sequence according to the object recognition information sequence, obtaining an image weight sequence; a generation unit configured to generate at least one cluster center information by using a cluster center information generation model according to the image weight sequence and the spliced image sequence; an execution unit configured to perform clustering processing on the spliced image sequence according to the at least one cluster center information to generate an image identification cluster set; a packaging unit configured to perform packaging processing on the image identification cluster set, the spliced image sequence, and the image time sequence corresponding to the spliced image sequence to generate a packaged file; a storage unit configured to store the packaged file and the set of captured videos at a captured video storage end for subsequent acquisition of video segments corresponding to the video segment object information.

[0010] Third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device having stored thereon one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation manner of the first aspect.

[0011] Fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having stored thereon a computer program, wherein when the program is executed by a processor, it implements the method described in any implementation manner of the first aspect.

[0012] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the video storage method of some embodiments of the present disclosure, it is possible to accurately and efficiently generate video segments corresponding to at least one video segment object information for a set of captured videos. Specifically, the reasons for the inaccurate and inefficient generation of video segments corresponding to relevant objects are as follows: Relevant technical personnel need to detect frame by frame whether there is video content related to the object, and the situation of less cropping often occurs, resulting in low cropping efficiency and wasting a large amount of human and computer device resources. Based on this, in the video storage method of some embodiments of the present disclosure, first, a set of captured videos for the captured space and at least one video segment object information are obtained. Among them, the above-mentioned set of captured videos includes: the first captured video corresponding to the first azimuth capturing device, the second captured video corresponding to the second azimuth capturing device, and the third captured video corresponding to the third azimuth capturing device. Here, through the multi-azimuth captured videos corresponding to the captured space, it is possible to determine the video segments corresponding to each video segment object information comprehensively and without dead angles in the subsequent process, realizing the comprehensive acquisition of the video content of the video segment objects. Then, the above-mentioned first captured video, the second captured video, and the third captured video are subjected to frame extraction at the same frequency to generate a first frame image sequence, a second frame image sequence, and a third frame image sequence, so as to facilitate the subsequent acquisition of video segments. Next, the above-mentioned first frame image sequence, the second frame image sequence, and the third frame image sequence are spliced in the order of the same time of the images to generate a spliced image sequence, so as to facilitate the aggregation of the image content at the same time in each azimuth and facilitate the more accurate determination of subsequent video segments. Then, each spliced image in the above-mentioned spliced image sequence is input into a pre-trained object recognition information generation model to accurately generate object recognition information and obtain an object recognition information sequence. Among them, the object recognition information is the recognition information associated with the above-mentioned at least one video segment object information. Here, the obtained object recognition information is used for the weight size of the image content of each subsequent spliced image, which can ensure that the subsequent acquisition of video segments is more accurate for the weight size of the image content. Furthermore, according to the above-mentioned object recognition information sequence, the image weights corresponding to each spliced image in the above-mentioned spliced image sequence can be accurately set to obtain an image weight sequence to determine the importance degree of the image content corresponding to each spliced image. Secondly, according to the above-mentioned image weight sequence and the spliced image sequence, using the clustering cluster center information generation model, at least one cluster center information can be accurately generated. Here, through the clustering cluster center information generation model, the preliminary determination of the cluster center before clustering can be realized, avoiding the occurrence of problems due to the large corresponding quantity of the input data set and the complex data calculation in clustering. In addition, the efficient determination of the cluster center of clustering can greatly shorten the clustering consumption time, improve the clustering efficiency, and shorten the clustering cycle. Among them, the clustering cluster center information generation model is a low-weight neural network model.Further, based on the above at least one cluster center information, perform clustering processing on the above spliced image sequence to efficiently generate an image identification cluster set. Next, package the above image identification cluster set, the above spliced image sequence, and the image time sequence corresponding to the above spliced image sequence to generate a packaged file for facilitating subsequent retrieval of the video segment corresponding to the object. Finally, store the above packaged file and the above camera video set in the camera video storage end for subsequent acquisition of the video segment corresponding to the object information of the video segment. In summary, first, the preliminary determination of at least one cluster center information can be realized through the clustering cluster center information generation model, which can greatly shorten the subsequent clustering cycle and improve the clustering efficiency and accuracy. In addition, by considering the feature of image weight, the accuracy of the subsequent video segment corresponding to the object to be acquired is further guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.

[0014] Figure 1 is a flowchart of some embodiments of a video storage method according to the present disclosure;

[0015] Figure 2 is a schematic structural diagram of some embodiments of a video storage device according to the present disclosure;

[0016] Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0018] It should also be noted that, for the sake of description, only the parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0019] It should be noted that the concepts such as "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0020] It should be noted that the modification of "one" and "multiple" mentioned in this disclosure is illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0021] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0022] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0023] Reference Figure 1 , which shows a flow 100 of some embodiments of a video storage method according to the present disclosure. The video storage method includes the following steps:

[0024] Step 101, obtaining a set of captured videos for a captured space and information of at least one video segment object.

[0025] In some embodiments, the execution subject of the above video storage method can obtain a set of captured videos for a captured space and information of at least one video segment object through a wired connection method or a wireless connection method. Among them, the captured space can be a spatial area captured by a first azimuth capturing device, a second azimuth capturing device, and a third azimuth capturing device. The first azimuth capturing device is a capturing device placed in the first azimuth. The second azimuth capturing device is a capturing device placed in the second azimuth. The third azimuth capturing device is a capturing device placed in the third azimuth. The first azimuth, the second azimuth, and the third azimuth are all different. The information of at least one video segment object can be the object information of a pre-determined main video segment object. The video segment object information can be the object information corresponding to a video object in the captured video. For example, the object information can be object identification information.

[0026] Step 102, performing same-frequency frame extraction processing on the above first captured video, the second captured video, and the third captured video to generate a first frame image sequence, a second frame image sequence, and a third frame image sequence.

[0027] In some embodiments, the above-mentioned execution entity may perform synchronous frame extraction processing on the above-mentioned first captured video, the above-mentioned second captured video, and the third captured video to generate a first frame image sequence, a second frame image sequence, and a third frame image sequence. The first frame image sequence is the image sequence obtained by frame extraction from the first captured video. The second frame image sequence is the image sequence obtained by frame extraction from the second captured video. The third frame image sequence is the image sequence obtained by frame extraction from the third captured video. The synchronous frame extraction processing may be performed according to the same frame extraction frequency.

[0028] Step 103, splice the above-mentioned first frame image sequence, the above-mentioned second frame image sequence, and the above-mentioned third frame image sequence in the chronological order of the images to generate a spliced image sequence.

[0029] In some embodiments, the above-mentioned execution entity may splice the above-mentioned first frame image sequence, the above-mentioned second frame image sequence, and the above-mentioned third frame image sequence in the chronological order of the images to generate a spliced image sequence. Among them, image splicing at the same time is to splice each image with the same video time point of the corresponding frame images.

[0030] In some optional implementation manners of some embodiments, the splicing of the above-mentioned first frame image sequence, the above-mentioned second frame image sequence, and the above-mentioned third frame image sequence in the chronological order of the images to generate a spliced image sequence may include the following steps:

[0031] The first step, for each first frame image in the above-mentioned first frame image sequence, perform the following first generation steps:

[0032] The first sub-step, determine the second frame image in the above-mentioned second frame image sequence that is time-synchronized with the above-mentioned first frame image as the second target frame image.

[0033] The second sub-step, determine the third frame image in the above-mentioned third frame image sequence that is time-synchronized with the above-mentioned first frame image as the third target frame image.

[0034] The third sub-step, perform camera calibration processing on the above-mentioned first frame image, the above-mentioned second target frame image, and the above-mentioned third target frame image to generate a first processed image, a second processed image, and a third processed image.

[0035] Here, due to the installation design and the differences between the imaging devices, there will be scaling (caused by inconsistent lens focal lengths), tilting (vertical rotation), and azimuth differences (horizontal rotation) between the frame images. Therefore, the physical differences need to be calibrated in advance to obtain images with good consistency for subsequent image splicing.

[0036] As an example, the above-mentioned execution entity can perform camera calibration processing on the above-mentioned first frame image, the above-mentioned second target frame image, and the above-mentioned third target frame image through relevant calibration algorithms to generate a first processed image, a second processed image, and a third processed image.

[0037] The fourth sub-step is to perform image coordinate transformation processing on the above-mentioned first processed image, the above-mentioned second processed image, and the above-mentioned third processed image to generate a first transformed image, a second transformed image, and a third transformed image.

[0038] The fifth sub-step is to perform image distortion correction on the above-mentioned first transformed image, the above-mentioned second transformed image, and the above-mentioned third transformed image to generate a first corrected image, a second corrected image, and a third corrected image. Here, image distortion correction is used to avoid distortion caused by the structure of the imaging device itself.

[0039] The sixth sub-step is to perform image projection transformation on the above-mentioned first corrected image, the above-mentioned second corrected image, and the above-mentioned third corrected image to generate a first projected image, a second projected image, and a third projected image.

[0040] The seventh sub-step is to perform matching point selection and calibration on the above-mentioned first projected image, the above-mentioned second projected image, and the above-mentioned third projected image to generate first feature point information, second feature point information, and third feature point information.

[0041] The eighth sub-step is to perform image stitching and fusion on the above-mentioned first projected image, the above-mentioned second projected image, and the above-mentioned third projected image according to the above-mentioned first feature point information, the above-mentioned second feature point information, and the above-mentioned third feature point information to generate a first stitched and fused image. Among them, image stitching and fusion include: registration and fusion. The purpose of registration is to register frame images into the same coordinate system according to the geometric motion model. Fusion is to synthesize the registered frame images into a large stitched image.

[0042] The ninth sub-step is to input the above-mentioned first frame image, the above-mentioned second target frame image, and the above-mentioned third target frame image into a stitching and fusion image generation model to generate a second stitched and fused image. Among them, the stitching and fusion image generation model can be a neural network model for generating stitched and fused images. In practice, the stitching and fusion image generation model can be a GAN network model.

[0043] The tenth sub-step is to input the above-mentioned first stitched and fused image and the above-mentioned second stitched and fused image into an end-to-end model to generate a third stitched and fused image. Among them, the end-to-end model can be a model of multiple cascaded first convolutional neural networks + multiple cascaded second convolutional neural networks. The multiple cascaded first convolutional neural networks are encoding models. The multiple cascaded second convolutional neural networks are decoding models.

[0044] The eleventh sub-step is to input the above-mentioned third spliced and fused image into a spliced and fused result classification model to generate a spliced and fused classification result. Among them, the above-mentioned spliced and fused classification result includes: a classification result indicating that the third spliced and fused image is not a spliced image, and a classification result indicating that the third spliced and fused image is a spliced image. The spliced and fused result classification model can be a neural network model that generates a spliced and fused classification result. The spliced and fused result classification model can be a discriminant model. In practice, the spliced and fused result classification model can be multiple cascaded convolutional neural network models.

[0045] The twelfth sub-step is to determine the above-mentioned third spliced and fused image as a spliced image in response to determining that the above-mentioned spliced and fused classification result represents a classification result indicating that the third spliced and fused image is not a spliced image.

[0046] Optionally, after the "tenth sub-step", the steps further include:

[0047] The first step is to perform the following second generation step for the third spliced and fused image:

[0048] The first sub-step is to extract the first image feature information and the first image main index information set corresponding to the above-mentioned third spliced and fused image. Among them, there is a one-to-one correspondence between the first image main index information in the first image main index information set and the image main indexes in the image main index set. The first image main index information is the index value corresponding to the image main index. The image main index set can be various pre-set image indexes. In practice, the image main index set can include but is not limited to at least one of the following: object position index, object wearing index. The first image feature information can represent the image content feature semantics of the third spliced and fused image.

[0049] As an example, the above-mentioned execution entity can use an image feature extraction model and an image main index information generation model to extract the first image feature information and the first image main index information set corresponding to the above-mentioned third spliced and fused image. The image main index information generation model can be a neural network model that generates image main index information.

[0050] The second sub-step is to extract the second image feature information and the second image main index information set corresponding to the above-mentioned first frame image. Details are not described again.

[0051] The third sub-step is to extract the third image feature information and the third image main index information set corresponding to the above-mentioned second target frame image. Details are not described again.

[0052] The fourth sub-step is to extract the fourth image feature information and the fourth image main index information set corresponding to the above-mentioned third target frame image. Details are not described again.

[0053] The fifth sub-step is to splice the second image feature information, the third image feature information, and the fourth image feature information to obtain spliced feature information.

[0054] The sixth sub-step is to input the spliced feature information into a convolutional neural network to generate compressed feature information. Among them, the vector dimension corresponding to the compressed feature information is the same as the vector dimension corresponding to the first image feature information. The convolutional neural network can be a convolutional network for feature compression of the spliced feature information.

[0055] The seventh sub-step is to determine the vector similarity between the first image feature information and the compressed feature information. The vector similarity can be a vector cosine distance value. That is, both the first image feature information and the compressed feature information can be information in vector form.

[0056] The eighth sub-step is to fuse the second image main index information set, the third image main index information set, and the fourth image main index information set to generate a fused index information set.

[0057] The ninth sub-step is to generate a first complete index information set for the fused index information set by using the target index knowledge graph. The target index knowledge graph can be a pre-set index knowledge graph for the corresponding frame image of the camera video.

[0058] As an example, first, the execution subject can determine the index set corresponding to the fused index information set. Then, determine the remaining index set in the target index knowledge graph that does not exist in the index set. Then, generate the remaining index information set corresponding to the remaining index set. Finally, fuse the remaining index information set and the fused index information set to generate a second complete index information set.

[0059] The tenth sub-step is to generate a second complete index information set for the first image main index information set by using the target index knowledge graph. The specific implementation method can refer to the generation of the first complete index information set.

[0060] The eleventh sub-step is to determine the index information difference between the first complete index information set and the second complete index information set.

[0061] The twelfth sub-step is to generate verification information indicating that the third spliced and fused image passes the verification in response to determining that the vector similarity is greater than the first value and the index information difference is less than the second value. The first value and the second value can be pre-set values.

[0062] In the second step, in response to determining that the above vector similarity is less than or equal to the first value and / or the above metric information difference is greater than or equal to the second value, perform image adjustment on the above third spliced and fused image to generate an adjusted image.

[0063] As an example, in response to determining that the above vector similarity is less than or equal to the first value and / or the above metric information difference is greater than or equal to the second value, sequentially select models whose corresponding network computation amount is greater than the corresponding model computation amount of the end-to-end model and whose corresponding model accuracy is greater than the corresponding model accuracy of the end-to-end model as candidate end-to-end models. Among them, there is an end-to-end model sequence, and each end-to-end model in the end-to-end model sequence is sorted according to the model accuracy and the model computation amount. Use the above candidate end-to-end models to generate the third spliced and fused image again, and determine the vector similarity and the metric information difference again, and loop sequentially until the vector similarity corresponding to the generated third spliced and fused image is greater than the first value and the above metric information difference is less than the second value.

[0064] In the third step, use the adjusted image as the third spliced and fused image, and continue to execute the above second generation step.

[0065] Step 104: Input each spliced image in the above spliced image sequence into a pre-trained object recognition information generation model to generate object recognition information, and obtain an object recognition information sequence.

[0066] In some embodiments, the above execution subject may input each spliced image in the above spliced image sequence into a pre-trained object recognition information generation model to generate object recognition information, and obtain an object recognition information sequence. Among them, the object recognition information is recognition information associated with the above at least one video segment object information. The object recognition information includes: at least one recognition result corresponding to the at least one video segment object information. There is a one-to-one correspondence between the video segment object information in the at least one video segment object information and the recognition result in the at least one recognition result. The recognition result is the result information of whether the spliced image has the content corresponding to the corresponding video segment object information. The object recognition information generation model may be a YOLO model.

[0067] Step 105: Set an image weight corresponding to each spliced image in the above spliced image sequence according to the above object recognition information sequence to obtain an image weight sequence.

[0068] In some embodiments, the above-mentioned execution entity may set an image weight corresponding to each spliced image in the above-mentioned spliced image sequence according to the above-mentioned object recognition information sequence, to obtain an image weight sequence. Wherein, the image weight may represent the content ratio of the object content corresponding to at least one video segment object information to the image content corresponding to the spliced image. The larger the image weight value, the larger the content ratio of the object content corresponding to at least one video segment object information.

[0069] As an example, for each spliced image in the above-mentioned spliced image sequence, first, determine the object recognition information corresponding to the spliced image. Then, according to the object recognition information, determine the content ratio of the object content corresponding to at least one video segment object information as the image weight.

[0070] Step 106, according to the above-mentioned image weight sequence and spliced image sequence, use the clustering cluster center information generation model to generate at least one cluster center information.

[0071] In some embodiments, the above-mentioned execution entity may use the clustering cluster center information generation model according to the above-mentioned image weight sequence and spliced image sequence to generate at least one cluster center information. Among them, the clustering cluster center information generation model may be a neural network model for generating clustering cluster center information. In practice, the clustering cluster center information generation model may be a feature extraction model + a multi-head attention mechanism model + at least one fully connected layer.

[0072] In some optional implementation manners of some embodiments, the above-mentioned using the clustering cluster center information generation model according to the above-mentioned image weight sequence and spliced image sequence to generate at least one cluster center information may include the following steps:

[0073] The first step is to input each spliced image in the above-mentioned spliced image sequence into the image feature extraction model included in the above-mentioned clustering cluster center information generation model to generate a spliced image feature information sequence. Wherein, the image feature extraction model may be a neural network model for extracting the semantic meaning of the image content.

[0074] The second step is to multiply the spliced image feature information in the above-mentioned spliced image feature information sequence by the corresponding image weight in the above-mentioned image weight sequence to generate multiplied feature information, to obtain a multiplied feature information sequence.

[0075] The third step is to determine the feature information distance between every two multiplied feature information in the above-mentioned multiplied feature information sequence to obtain a first feature information distance set.

[0076] The fourth step is to determine the distance distribution corresponding to the above-mentioned first feature information distance set. For example, the distance distribution may be a normal distribution or a Poisson distribution.

[0077] Step 5: Obtain a clustering cluster center information generation model corresponding to the above distance distribution. For each distribution, there is a corresponding clustering cluster center information generation model. For example, for a normal distribution, there is a corresponding clustering cluster center information generation model, and for a Poisson distribution, there is also a corresponding clustering cluster center information generation model.

[0078] Step 6: Input the above multiplication feature information sequence into the cluster center information generation model included in the above clustering cluster center information generation model to generate at least one candidate cluster center information. The cluster center information generation model can be a model for generating cluster center information.

[0079] Step 7: Set a cluster center interval distance range according to the above distance distribution.

[0080] As an example, the above execution entity can set a corresponding distribution abscissa intercept range according to different distance distributions to generate a corresponding cluster center interval distance range.

[0081] Step 8: Determine the second feature information distance between every two candidate cluster center information in the above at least one candidate cluster center information to obtain a second feature information distance set.

[0082] Step 9: In response to determining that each second feature information distance in the above second feature information distance set is within the above cluster center interval distance range, determine the above at least one candidate cluster center information as at least one cluster center information.

[0083] In some optional implementation manners of some embodiments, the above cluster center information generation model includes: a first encoding network model and at least one decoding network model. Among them, there is a one-to-one correspondence between the decoding network models in the above at least one decoding network model and the video segment object information in the above at least one video segment object information. The first encoding network model can be a convolutional neural network model for encoding connected in series in multiple layers. The decoding network models in the at least one decoding network model can also be convolutional neural network models for decoding connected in series in multiple layers.

[0084] Optionally, inputting the above multiplication feature information sequence into the cluster center information generation model included in the above clustering cluster center information generation model to generate at least one candidate cluster center information may include the following steps:

[0085] Step 1: Input each multiplication feature information in the above multiplication feature information sequence into the above first encoding network model to generate each encoded feature information.

[0086] In the second step, each of the above encoding feature information is input into the corresponding decoding network model among the above at least one decoding network model to output a decoding result, obtaining at least one candidate cluster center information.

[0087] Considering the problems of the above conventional solutions, in the face of the above technical problems: when performing image clustering through a conventional clustering algorithm, due to the high resolution of the image, not only the accuracy cannot be effectively guaranteed, but also the computational amount is not low. Combining the advantages / technical status owned, the following solution can be determined.

[0088] In some optional implementation manners of some embodiments, the above clustering cluster center information generation model is trained through the following steps:

[0089] In the first step, a training data set is obtained, where the training data in the above training data set is an image sequence.

[0090] In the second step, target training data is randomly selected from the training data set.

[0091] In the third step, for the target training data, the following training steps are performed:

[0092] In the first sub-step, each image in the image sequence corresponding to the above target training data is input into the initial image feature extraction model included in the initial clustering cluster center information generation model to generate initial image feature information, obtaining an image feature information sequence. Wherein, the initial clustering cluster center information generation model is a clustering center information generation model whose parameters have not been trained yet. The initial image feature extraction model can be an image feature extraction model whose parameters have not been trained yet.

[0093] In the second sub-step, each image feature information in the above image feature information sequence is multiplied by the corresponding image weight to generate image feature multiplication information, obtaining an image feature multiplication information sequence.

[0094] In the third sub-step, using at least one clustering algorithm, clustering processing is performed on each of the image feature multiplication information in the above image feature multiplication information sequence to generate at least one image feature multiplication information cluster set. Among them, the algorithm logics of the various clustering algorithms in the at least one clustering algorithm are different. For example, the at least one clustering algorithm may include: K-means algorithm, K-Means++ algorithm, DBSCAN algorithm.

[0095] In the fourth sub-step, for each image feature multiplication information cluster set in the at least one image feature multiplication information cluster set, the above image feature multiplication information cluster center set is determined.

[0096] The fifth sub-step is to perform clustering processing on each image feature multiplication information cluster center in the above at least one image feature multiplication information cluster center set to generate a cluster center cluster set.

[0097] The sixth sub-step is to perform weighted summation processing on each cluster center included in each cluster center cluster in the cluster center cluster set to generate a weighted summation cluster center, and obtain a weighted summation cluster center set as the cluster center label set corresponding to the image feature information sequence.

[0098] The seventh sub-step is to determine the third feature information distance set corresponding to the above image feature information sequence, where the generation of the third feature information distance set can refer to the generation of the first feature information distance set.

[0099] The eighth sub-step is to determine the distance distribution corresponding to the above third feature information distance set as the target distance distribution.

[0100] The ninth sub-step is to obtain a clustering cluster center information generation model corresponding to the above target distance distribution as the target initial clustering cluster center information generation model.

[0101] The tenth sub-step is to input the above image feature multiplication information sequence into the target initial cluster center information generation model included in the above target initial clustering cluster center information generation model to generate at least one initial candidate cluster center information.

[0102] The eleventh sub-step is to generate cluster center difference loss information between the above at least one initial candidate cluster center information and the above cluster center label set.

[0103] As an example, the above execution entity can generate cluster center difference loss information between the above at least one initial candidate cluster center information and the above cluster center label set through a cross-entropy loss function.

[0104] The twelfth sub-step is to, in response to determining that the cluster center difference loss information indicates that the loss information tends to converge, determine the initial clustering cluster center information generation model as the clustering cluster center information generation model. That is, the target initial cluster center information generation model is determined as the cluster center information generation model.

[0105] The fourth step is to, in response to determining that the cluster center difference loss information indicates that the loss information does not tend to converge, update the model parameters in the above initial clustering cluster center information generation model according to the above cluster center difference loss information to generate an updated clustering cluster center information generation model.

[0106] The fifth step is to remove the target training data from the training data set to obtain a post-removal training data set.

[0107] The sixth step is to select training data from the above post-removal training data set as candidate training data.

[0108] In the seventh step, use the candidate training data as the target training data, and update the cluster center information generation model as the initial updated cluster center information generation model, and continue to execute the training step.

[0109] The content in "In some optional implementation manners of some embodiments" above, as an inventive point of the present disclosure, solves the technical problem mentioned in the background art that "when performing clustering of images through a conventional clustering algorithm, due to the high resolution of the images, not only the accuracy cannot be effectively guaranteed, but also the calculation amount is not low." Based on this, in the present disclosure, preliminary clustering of image feature information is performed through multiple clustering algorithms. On this basis, re-clustering processing is performed on the obtained cluster centers of each image feature multiplication information. With the guarantee of multiple clustering algorithms, a corresponding cluster center label set can be generated accurately and quickly. Based on the accurate cluster center label set, accurate and efficient training of the initial cluster center information generation model can be realized, and a cluster center information generation model with more accurate output cluster center information can be obtained.

[0110] Step 107, perform clustering processing on the above-mentioned spliced image sequence according to the above at least one cluster center information to generate an image identification cluster set.

[0111] In some embodiments, the above-mentioned execution subject may perform clustering processing on the above-mentioned spliced image sequence according to the above at least one cluster center information to generate an image identification cluster set. Among them, there is a corresponding spliced image for the image identification. The main image content corresponding to each spliced image in the spliced image cluster corresponding to the image identification cluster in the image identification cluster set is for the same video segment object information. There is an object correspondence relationship between the image identification clusters in the image identification cluster set and the video segment object information in the at least one video segment object information.

[0112] In some optional implementation manners of some embodiments, the above-mentioned performing clustering processing on the above-mentioned spliced image sequence according to the above at least one cluster center information to generate an image identification cluster set may include the following steps:

[0113] In the first step, for each cluster center information in the above at least one cluster center information, screen out the spliced image feature information with the smallest vector distance from the above-mentioned spliced image feature information sequence to the above-mentioned cluster center information as the target spliced image feature information.

[0114] In the second step, use the obtained at least one target spliced image feature information as at least one initial cluster center, and perform clustering processing on the above-mentioned spliced image sequence to generate an image identification cluster set.

[0115] Step 108: Package the above image identification clusters, the above spliced image sequences, and the image time sequences corresponding to the above spliced image sequences to generate a packaged file.

[0116] In some embodiments, the above execution entity may package the above image identification clusters, the above spliced image sequences, and the image time sequences corresponding to the above spliced image sequences to generate a packaged file. Among them, there is a one-to-one correspondence between the spliced images in the spliced image sequence and the image times in the image time sequence. The image time may be the video time point corresponding to the spliced image.

[0117] Step 109: Store the above packaged file and the above camera video set in the camera video storage end for subsequent acquisition of video segments corresponding to video segment object information.

[0118] In some embodiments, the above execution entity may store the above packaged file and the above camera video set in the camera video storage end for subsequent acquisition of video segments corresponding to video segment object information. Among them, the camera video storage end may be a terminal for storing camera videos.

[0119] In some optional implementation manners of some embodiments, after step 109, the steps further include:

[0120] First step: In response to receiving a video segment acquisition request for target object information, obtain the above packaged file and the above camera video set from the above camera video storage end. Among them, the video segment acquisition request may be a request for acquiring a video segment corresponding to the target object information. The target object information may be an object for which a video segment is to be acquired.

[0121] Second step: Determine the image identification cluster corresponding to the above target object information from the image identification clusters included in the above packaged file as the target image identification cluster.

[0122] As an example, the above execution entity may determine the image identification cluster corresponding to the above target object information from the image identification clusters included in the above packaged file as the target image identification cluster by means of query.

[0123] Third step: Determine the sub-sequences of spliced images corresponding to each image identification in the above target image identification cluster.

[0124] Fourth step: Obtain a set of camera video segment sequences corresponding to the above sub-sequences of spliced images from the above camera video set.

[0125] As an example, the above execution entity may crop video segment sequences with the same corresponding time as the sub-sequences of spliced images from each camera video in the camera video set to obtain a set of camera video segment sequences.

[0126] Step 5: Send the above sequence set of camera video clips and the above subsequence of spliced images to the video clip acquisition terminal.

[0127] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the video storage method of some embodiments of the present disclosure, it is possible to accurately and efficiently generate video segments corresponding to at least one video segment object information for a set of captured videos. Specifically, the reasons for the inaccurate and inefficient generation of video segments corresponding to relevant objects are as follows: Relevant technical personnel need to detect frame by frame whether there is video content related to the object, and there are often cases of undercropping, resulting in low cropping efficiency and wasting a large amount of human and computer device resources. Based on this, in the video storage method of some embodiments of the present disclosure, first, a set of captured videos for the captured space and at least one video segment object information are obtained. Among them, the above-mentioned set of captured videos includes: the first captured video corresponding to the first azimuth camera device, the second captured video corresponding to the second azimuth camera device, and the third captured video corresponding to the third azimuth camera device. Here, through the multi-azimuth captured videos corresponding to the captured space, it is possible to determine the video segments corresponding to each video segment object information without dead angles in all directions, realizing the all-round acquisition of the video content of the video segment objects. Then, frame extraction processing is performed on the first captured video, the second captured video, and the third captured video at the same frequency to generate a first frame image sequence, a second frame image sequence, and a third frame image sequence, facilitating the subsequent acquisition of video segments. Next, the first frame image sequence, the second frame image sequence, and the third frame image sequence are spliced in the order of the same time of the images to generate a spliced image sequence, facilitating the aggregation of the image content at the same time in each azimuth and facilitating the more accurate determination of subsequent video segments. Then, each spliced image in the spliced image sequence is input into a pre-trained object recognition information generation model to accurately generate object recognition information, obtaining an object recognition information sequence. Among them, the object recognition information is the recognition information associated with the above-mentioned at least one video segment object information. Here, the obtained object recognition information is used for the weight size of the image content of each subsequent spliced image, which can ensure the more accurate acquisition of video segments for the weight size of the image content. Furthermore, according to the object recognition information sequence, the image weights corresponding to each spliced image in the spliced image sequence can be accurately set to obtain an image weight sequence to determine the importance degree of the image content corresponding to each spliced image. Secondly, according to the image weight sequence and the spliced image sequence, using a clustering cluster center information generation model, at least one cluster center information can be accurately generated. Here, through the clustering cluster center information generation model, the preliminary determination of the cluster center before clustering can be realized, avoiding the problems caused by the large number of input data sets and complex data calculations in clustering. In addition, the efficient determination of the cluster center of clustering can greatly shorten the clustering consumption time, improve the clustering efficiency, and shorten the clustering cycle. Among them, the clustering cluster center information generation model is a low-weight neural network model.Further, based on the above at least one cluster center information, perform clustering processing on the above-mentioned spliced image sequence to efficiently generate an image identification cluster set. Next, package the above image identification cluster set, the above spliced image sequence, and the image time sequence corresponding to the above spliced image sequence to generate a packaged file for facilitating subsequent retrieval of the video segment corresponding to the object. Finally, store the above packaged file and the above camera video set in the camera video storage end for subsequent acquisition of the video segment corresponding to the object information of the video segment. In summary, the preliminary determination of at least one cluster center information can be realized first through the clustering cluster center information generation model, which can greatly shorten the subsequent clustering cycle and improve the clustering efficiency and accuracy. In addition, the factor of image weight is also considered, further ensuring the accuracy of the subsequent video segment corresponding to the object to be obtained.

[0128] Further reference Figure 2 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a video storage device, and these device embodiments correspond to Figure 1 the method embodiments shown, and the video storage device can be specifically applied to various electronic devices.

[0129] As Figure 2As shown in the figure, a video storage device 200 includes: an acquisition unit 201, a frame extraction unit 202, a splicing unit 203, an input unit 204, a setting unit 205, a generation unit 206, an execution unit 207, a packaging unit 208, and a storage unit 209. Among them, the acquisition unit 201 is configured to acquire a set of video recordings of a camera space and information of at least one video clip object. Among them, the set of video recordings includes: a first video recording corresponding to a first azimuth camera device, a second video recording corresponding to a second azimuth camera device, and a third video recording corresponding to a third azimuth camera device; the frame extraction unit 202 is configured to perform frame extraction processing at the same frequency on the first video recording, the second video recording, and the third video recording to generate a first frame image sequence, a second frame image sequence, and a third frame image sequence; the splicing unit 203 is configured to splice the first frame image sequence, the second frame image sequence, and the third frame image sequence in the chronological order of the images to generate a spliced image sequence; the input unit 204 is configured to input each spliced image in the spliced image sequence into a pre-trained object recognition information generation model to generate object recognition information and obtain an object recognition information sequence, where the object recognition information is recognition information associated with the information of at least one video clip object; the setting unit 205 is configured to set an image weight corresponding to each spliced image in the spliced image sequence according to the object recognition information sequence to obtain an image weight sequence; the generation unit 206 is configured to generate at least one cluster center information by using a cluster center information generation model according to the image weight sequence and the spliced image sequence; the execution unit 207 is configured to perform clustering processing on the spliced image sequence according to the at least one cluster center information to generate an image identification cluster set; the packaging unit 208 is configured to perform packaging processing on the image identification cluster set, the spliced image sequence, and the image time sequence corresponding to the spliced image sequence to generate a packaged file; the storage unit 209 is configured to store the packaged file and the set of video recordings at a video recording storage end.

[0130] It can be understood that the various units described in the video storage device 200 correspond to the respective steps in the method described in the reference Figure 1 Therefore, the operations, features, and beneficial effects described above for the method also apply to the video storage device 200 and the units included therein, and will not be repeated here.

[0131] The following refers to Figure 3 , which shows a schematic structural diagram of an electronic device (for example, an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.

[0132] As Figure 3 shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0133] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included. Figure 3 Each block shown in

[0134] particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from a network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the methods of some embodiments of the present disclosure are performed.

[0135] It should be noted that in some embodiments of the present disclosure, the above-mentioned computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0136] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (“LAN”), wide area networks (“WAN”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.

[0137] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: obtain a set of captured video for a captured space and at least one video clip object information, wherein the set of captured video includes: a first captured video corresponding to a first azimuth capturing device, a second captured video corresponding to a second azimuth capturing device, and a third captured video corresponding to a third azimuth capturing device; perform synchronous frame extraction processing on the first captured video, the second captured video, and the third captured video to generate a first frame image sequence, a second frame image sequence, and a third frame image sequence; splice the first frame image sequence, the second frame image sequence, and the third frame image sequence in chronological order of the images to generate a spliced image sequence; input each spliced image in the spliced image sequence into a pre-trained object recognition information generation model to generate object recognition information, obtaining an object recognition information sequence, wherein the object recognition information is recognition information associated with the at least one video clip object information; according to the object recognition information sequence, set an image weight corresponding to each spliced image in the spliced image sequence to obtain an image weight sequence; according to the image weight sequence and the spliced image sequence, use a clustering cluster center information generation model to generate at least one cluster center information; according to the at least one cluster center information, perform clustering processing on the spliced image sequence to generate an image identification cluster set; perform a packaging process on the image identification cluster set, the spliced image sequence, and the image time sequence corresponding to the spliced image sequence to generate a packaged file; store the packaged file and the set of captured video in a captured video storage end for subsequent acquisition of video clips corresponding to the video clip object information.

[0138] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0140] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an acquisition unit, a frame extraction unit, a splicing unit, an input unit, a setting unit, a generation unit, an execution unit, a packaging unit, and a storage unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the acquisition unit can also be described as "the unit for acquiring a set of camera videos for a camera space and information on at least one video segment object".

[0141] The functions described above can be at least partially performed by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0142] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A video storage method, comprising: Obtaining a set of captured videos for a captured space and information of at least one video segment object, wherein the set of captured videos includes: a first captured video corresponding to a first azimuth capturing device, a second captured video corresponding to a second azimuth capturing device, and a third captured video corresponding to a third azimuth capturing device; Performing frame extraction at the same frequency on the first captured video, the second captured video, and the third captured video to generate a first sequence of frame images, a second sequence of frame images, and a third sequence of frame images; Splicing the first sequence of frame images, the second sequence of frame images, and the third sequence of frame images in chronological order of the images to generate a spliced sequence of images; Inputting each spliced image in the spliced sequence of images into a pre-trained object recognition information generation model to generate object recognition information, obtaining a sequence of object recognition information, wherein the object recognition information is recognition information associated with the information of the at least one video segment object; Setting an image weight corresponding to each spliced image in the spliced sequence of images according to the sequence of object recognition information to obtain a sequence of image weights; Generating at least one cluster center information by using a cluster center information generation model according to the sequence of image weights and the spliced sequence of images; Performing clustering processing on the spliced sequence of images according to the at least one cluster center information to generate a set of image identification clusters; Packaging the set of image identification clusters, the spliced sequence of images, and the image time sequence corresponding to the spliced sequence of images to generate a packaged file; Storing the packaged file and the set of captured videos in a captured video storage terminal for subsequent acquisition of video segments corresponding to the information of the video segment object.

2. The method according to claim 1, wherein, The method further includes: In response to receiving a video segment acquisition request for target object information, obtaining the packaged file and the set of captured videos from the captured video storage terminal; Determining, according to the set of image identification clusters included in the packaged file, the image identification cluster corresponding to the target object information as a target image identification cluster; Determining a subsequence of spliced images corresponding to each image identification in the target image identification cluster; Obtaining a set of sequences of captured video segments corresponding to the subsequence of spliced images from the set of captured videos; Sending the set of sequences of captured video segments and the subsequence of spliced images to a video segment acquisition terminal.

3. The method according to claim 2, wherein The splicing the first sequence of frame images, the second sequence of frame images, and the third sequence of frame images in chronological order of the images to generate a spliced sequence of images includes: For each first frame image in the first sequence of frame images, performing the following first generation step: Determining a second target frame image in the second sequence of frame images that is time-synchronized with the first frame image as the second target frame image; Determining a third target frame image in the third sequence of frame images that is time-synchronized with the first frame image as the third target frame image; Performing camera calibration processing on the first frame image, the second target frame image, and the third target frame image to generate a first processed image, a second processed image, and a third processed image; Perform image coordinate transformation processing on the first processed image, the second processed image, and the third processed image to generate a first transformed image, a second transformed image, and a third transformed image; Perform image distortion correction on the first transformed image, the second transformed image, and the third transformed image to generate a first corrected image, a second corrected image, and a third corrected image; Perform image projection transformation on the first corrected image, the second corrected image, and the third corrected image to generate a first projected image, a second projected image, and a third projected image; Perform matching point selection and calibration on the first projected image, the second projected image, and the third projected image to generate first feature point information, second feature point information, and third feature point information; According to the first feature point information, the second feature point information, and the third feature point information, perform image stitching and fusion on the first projected image, the second projected image, and the third projected image to generate a first stitched and fused image; Input the first frame image, the second target frame image, and the third target frame image into a stitching and fusion image generation model to generate a second stitched and fused image; Input the first stitched and fused image and the second stitched and fused image into an end-to-end model to generate a third stitched and fused image; Input the third stitched and fused image into a stitching and fusion result classification model to generate a stitching and fusion classification result, where the stitching and fusion classification result includes: a classification result indicating that the third stitched and fused image is not a stitched image, and a classification result indicating that the third stitched and fused image is a stitched image; In response to determining that the stitching and fusion classification result indicates that the third stitched and fused image is not a stitched image, determine the third stitched and fused image as a stitched image.

4. The method according to claim 3, wherein After the step of inputting the first stitched and fused image and the second stitched and fused image into an end-to-end model to generate a third stitched and fused image, the method further includes: For the third stitched and fused image, perform the following second generation steps: Extract the first image feature information and the first set of image main index information corresponding to the third stitched and fused image; Extract the second image feature information and the second set of image main index information corresponding to the first frame image; Extract the third image feature information and the third set of image main index information corresponding to the second target frame image; Extract the fourth image feature information and the fourth set of image main index information corresponding to the third target frame image; Perform feature information stitching on the second image feature information, the third image feature information, and the fourth image feature information to obtain stitched feature information; Input the stitched feature information into a convolutional neural network to generate compressed feature information, where the vector dimension corresponding to the compressed feature information is the same as the vector dimension corresponding to the first image feature information; Determine the vector similarity between the first image feature information and the compressed feature information; Perform index information fusion on the second set of image main index information, the third set of image main index information, and the fourth set of image main index information to generate a fused set of index information; Generate a first complete index information set for the fused index information set by using the target index knowledge graph; Generate a second complete index information set for the first image main index information set by using the target index knowledge graph; Determine the index information difference between the first complete index information set and the second complete index information set; In response to determining that the vector similarity is greater than a first value and the index information difference is less than a second value, generate verification information indicating that the third spliced and fused image passes the verification; In response to determining that the vector similarity is less than or equal to the first value and / or the index information difference is greater than or equal to the second value, perform image adjustment on the third spliced and fused image to generate an adjusted image; Use the adjusted image as the third spliced and fused image, and continue to execute the second generation step.

5. The method according to claim 4, wherein, The generating of at least one cluster center information by using the clustering cluster center information generation model according to the image weight sequence and the spliced image sequence includes: Input each spliced image in the spliced image sequence into the image feature extraction model included in the clustering cluster center information generation model to generate a spliced image feature information sequence; Multiply the spliced image feature information in the spliced image feature information sequence by the corresponding image weight in the image weight sequence to generate multiplied feature information, and obtain a multiplied feature information sequence; Determine the feature information distance between every two multiplied feature information in the multiplied feature information sequence to obtain a first feature information distance set; Determine the distance distribution corresponding to the first feature information distance set; Obtain the clustering cluster center information generation model corresponding to the distance distribution; Input the multiplied feature information sequence into the cluster center information generation model included in the clustering cluster center information generation model to generate at least one candidate cluster center information; Set a cluster center interval distance range according to the distance distribution; Determine the second feature information distance between every two candidate cluster center information in the at least one candidate cluster center information to obtain a second feature information distance set; In response to determining that each second feature information distance in the second feature information distance set is within the cluster center interval distance range, determine the at least one candidate cluster center information as at least one cluster center information.

6. The method according to claim 5, wherein The cluster center information generation model includes: a first encoding network model and at least one decoding network model, wherein there is a one-to-one correspondence between the decoding network models in the at least one decoding network model and the video segment object information in the at least one video segment object information; and Inputting the multiplied feature information sequence into the cluster center information generation model included in the clustering cluster center information generation model to generate at least one candidate cluster center information includes: Input each multiplied feature information in the multiplied feature information sequence into the first encoding network model to generate respective encoded feature information; Input each encoded feature information in the respective encoded feature information into the corresponding decoding network model in the at least one decoding network model to output a decoding result, and obtain at least one candidate cluster center information.

7. The method according to claim 6, wherein Performing clustering processing on the spliced image sequence according to the at least one cluster center information to generate an image identification cluster set includes: For each piece of cluster center information in the at least one piece of cluster center information, screening out the spliced image feature information with the smallest vector distance from the spliced image feature information sequence to the cluster center information as the target spliced image feature information; Performing clustering processing on the spliced image sequence with the obtained at least one piece of target spliced image feature information as at least one initial cluster center to generate an image identification cluster set.

8. A video storage device includes: An acquisition unit configured to acquire a set of camera videos for a camera space and at least one video segment object information, where the set of camera videos includes: a first camera video corresponding to a first azimuth camera device, a second camera video corresponding to a second azimuth camera device, and a third camera video corresponding to a third azimuth camera device; A frame extraction unit configured to perform synchronous frame extraction processing on the first camera video, the second camera video, and the third camera video to generate a first frame image sequence, a second frame image sequence, and a third frame image sequence; A splicing unit configured to splice the first frame image sequence, the second frame image sequence, and the third frame image sequence in the same time order of the images to generate a spliced image sequence; An input unit configured to input each spliced image in the spliced image sequence into a pre-trained object recognition information generation model to generate object recognition information and obtain an object recognition information sequence, where the object recognition information is recognition information associated with the at least one video segment object information; A setting unit configured to set an image weight corresponding to each spliced image in the spliced image sequence according to the object recognition information sequence to obtain an image weight sequence; A generation unit configured to generate at least one piece of cluster center information by using a clustering cluster center information generation model according to the image weight sequence and the spliced image sequence; An execution unit configured to perform clustering processing on the spliced image sequence according to the at least one piece of cluster center information to generate an image identification cluster set; A packaging unit configured to perform packaging processing on the image identification cluster set, the spliced image sequence, and the image time sequence corresponding to the spliced image sequence to generate a packaged file; A storage unit configured to store the packaged file and the set of camera videos in a camera video storage end.

9. An electronic device includes: One or more processors; A storage device having stored thereon one or more programs, When the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method according to any one of claims 1-7.

10. A computer-readable medium having a computer program stored thereon, wherein, The program, when executed by the processor, implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-angle free viewing angle video data generation method and device, medium and server

    CN111669567A

  • Vehicle-related information sending method, device and equipment and computer readable medium

    CN118247744A