Video acquisition method, device, electronic device and computer-readable medium
The method uses multiple camera angles and two-stage clustering to efficiently and accurately extract video segments of target objects by leveraging machine learning algorithms, addressing inefficiencies and inaccuracies in manual video slicing.
Patent Information
- Application Number
- CN202411206550.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-08-30
AI Technical Summary
In the prior art, artificial video slicing is inefficient and video clips are not complete enough, making it difficult to accurately obtain video information of the target object.
By acquiring multiple target videos, shooting at different directions using multiple camera devices, video frame extraction processing is performed to generate frame sequences, and the first and second clustering algorithms are used to cluster and label processing of frame vectors to generate a video clip set.
It realizes efficient and accurate acquisition of video clip sets corresponding to target objects from multiple target videos, improving the accuracy and efficiency of video clip acquisition.
Smart Images

Figure CN118945422B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and more particularly, to methods, devices, electronic devices, and computer-readable media for video acquisition. Background Art
[0002] Currently, camera videos play a crucial monitoring role in people's daily lives. For the acquisition of relevant videos of a target object, the commonly adopted method is to manually slice the captured video to obtain a video segment including the target object.
[0003] However, when the above method is used to obtain the video segment corresponding to the target object, the following technical problems often exist:
[0004] First, the efficiency of manual video slicing is low, and the obtained video segment may not be complete enough.
[0005] Second, it is crucial to accurately determine the corresponding object recognition information for at least one object vector under the same vector cluster.
[0006] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to ordinary skilled artisans in the country. Summary of the Invention
[0007] The content part of the present disclosure is used to introduce the inventive concept in a brief form, and these inventive concepts will be described in detail in the following detailed implementation part. The content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0008] Some embodiments of the present disclosure propose methods, devices, electronic devices, and computer-readable media for video acquisition to solve one or more of the technical problems mentioned in the above background art section.
[0009] In a first aspect, some embodiments of the present disclosure provide a video acquisition method, including: acquiring a plurality of target videos for a target space, where a plurality of camera devices corresponding to the plurality of target videos are in different orientations in the target space; performing video frame extraction processing on each of the plurality of target videos to generate a video frame sequence, obtaining a plurality of video frame sequences; generating a plurality of frame vector sequences for the plurality of video frame sequences; using a first clustering algorithm to perform clustering processing on each frame vector in the plurality of frame vector sequences to generate a first clustering result; according to the first clustering result, tagging each frame vector in the plurality of frame vector sequences with a feature label to generate a plurality of first frame vector sequences after adding labels; using a second clustering algorithm to perform clustering processing on each first frame vector in the plurality of first frame vector sequences to generate a second clustering result; in response to receiving a video acquisition request for a target object and the second clustering result passing verification, according to the second clustering result, acquiring a video segment set corresponding to the target object from the plurality of target videos.
[0010] In a second aspect, some embodiments of the present disclosure provide a video acquisition device, including: a first acquisition unit configured to acquire a plurality of target videos for a target space, where a plurality of camera devices corresponding to the plurality of target videos are in different orientations in the target space; a frame extraction unit configured to perform video frame extraction processing on each of the plurality of target videos to generate a video frame sequence, obtaining a plurality of video frame sequences; a first generation unit configured to generate a plurality of frame vector sequences for the plurality of video frame sequences; a first execution unit configured to use a first clustering algorithm to perform clustering processing on each frame vector in the plurality of frame vector sequences to generate a first clustering result; a tagging unit configured to, according to the first clustering result, tag each frame vector in the plurality of frame vector sequences with a feature label to generate a plurality of first frame vector sequences after adding labels; a second execution unit configured to use a second clustering algorithm to perform clustering processing on each first frame vector in the plurality of first frame vector sequences to generate a second clustering result; a second acquisition unit configured to, in response to receiving a video acquisition request for a target object and the second clustering result passing verification, according to the second clustering result, acquire a video segment set corresponding to the target object from the plurality of target videos.
[0011] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method described in any implementation manner of the first aspect.
[0012] Fourthly, some embodiments of the present disclosure provide a computer-readable medium storing a computer program, wherein when the program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0013] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: The video acquisition method according to some embodiments of the present disclosure can accurately and efficiently acquire a set of video segments related to a target object from multiple target videos. Specifically, the reasons for the inaccurate and inefficient acquisition of relevant video segments are as follows: The efficiency of manually slicing videos is low, and the obtained video segments may not be complete. Based on this, the video acquisition method according to some embodiments of the present disclosure first acquires multiple target videos for a target space as the video source of the target object to subsequently acquire video segments for the target object. Among them, the multiple camera devices corresponding to the above-mentioned multiple target videos are in different orientations in the above-mentioned target space. Then, video frame extraction is performed on each of the above-mentioned multiple target videos to generate a video frame sequence, and multiple video frame sequences are obtained to facilitate the determination of video segments from the frame perspective, ensuring the accuracy of subsequent video segments. Next, multiple frame vector sequences for the above-mentioned multiple video frame sequences can be accurately generated for subsequent clustering processing. Then, using the first clustering algorithm, clustering processing can be accurately performed on each frame vector in the above-mentioned multiple frame vector sequences to generate a first clustering result. Here, through the first clustering algorithm, vectors with a certain feature can be clustered together to facilitate the subsequent determination of corresponding video segments for the target object. Furthermore, according to the above-mentioned first clustering result, each frame vector in the above-mentioned multiple frame vector sequences is labeled with a feature to generate multiple first frame vector sequences after adding labels, facilitating subsequent second clustering processing and improving the accuracy of subsequent video segment acquisition. Further, using the second clustering algorithm, clustering processing can be accurately performed on each first frame vector in the above-mentioned multiple first frame vector sequences to generate a second clustering result to obtain a label for each video frame, facilitating subsequent video segment acquisition. Finally, in response to receiving a video acquisition request for a target object and the above-mentioned second clustering result passing the verification, according to the above-mentioned second clustering result, a set of video segments corresponding to the above-mentioned target object can be accurately acquired from the above-mentioned multiple target videos. In summary, through two-stage clustering processing, a label set corresponding to each frame image in multiple target videos can be obtained. Based on this, after receiving a video acquisition request for a target object and the above-mentioned second clustering result passing the verification, a set of video segments corresponding to the above-mentioned target object can be efficiently and accurately acquired from the above-mentioned multiple target videos. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In conjunction with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the elements and components are not necessarily drawn to scale.
[0015] Figure 1 is a flowchart of some embodiments of a video acquisition method according to the present disclosure;
[0016] Figure 2 is a schematic structural diagram of some embodiments of a video acquisition device according to the present disclosure;
[0017] Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Specific Embodiments
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not used to limit the protection scope of the present disclosure.
[0019] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0020] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules, or units.
[0021] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] The present disclosure will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0024] Reference Figure 1, which shows the process 100 of some embodiments of the video acquisition method according to the present disclosure. The video acquisition method includes the following steps:
[0025] Step 101, acquire multiple target videos for a target space.
[0026] In some embodiments, the execution subject of the above video acquisition method may acquire multiple target videos for the target space through a wired connection method or a wireless connection method. Among them, the target space may be the space where video shooting is performed. For example, the target space may be a factory area. There is a one-to-one shooting relationship between the target videos in the multiple target videos and the camera devices in the multiple camera devices. And the multiple camera devices corresponding to the multiple target videos are in different orientations of the above target space.
[0027] Step 102, perform video frame extraction processing on each of the multiple target videos to generate a video frame sequence, and obtain multiple video frame sequences.
[0028] In some embodiments, the execution subject may perform video frame extraction processing on each of the multiple target videos to generate a video frame sequence, and obtain multiple video frame sequences.
[0029] As an example, the execution subject may set the frame extraction duration, and perform video frame extraction processing on each of the multiple target videos according to the frame extraction duration to generate a video frame sequence, and obtain multiple video frame sequences.
[0030] Step 103, generate multiple frame vector sequences for the multiple video frame sequences.
[0031] In some embodiments, the execution subject may generate multiple frame vector sequences for the multiple video frame sequences. A frame vector may be information in the form of a vector corresponding to at least one video frame.
[0032] In some optional implementation manners of some embodiments, the generation of multiple frame vector sequences for the multiple video frame sequences may include the following steps:
[0033] The first step, acquire the frame extraction frequency information corresponding to the multiple target videos. Among them, the frame extraction frequency information may be the frequency size of frame extraction for the target video. The frame extraction frequency information may be to extract one frame every 0.5 seconds.
[0034] The second step, generate the number of merged frame features according to the frame extraction frequency information.
[0035] As an example, the above-mentioned execution entity may multiply the number of frame extraction frequencies by a preset value to generate the number of frame feature combinations. Among them, the preset value may be a preset value. For example, the preset value may be the value 2. The preset value may be obtained from relevant experimental data.
[0036] In the third step, according to the above-mentioned number of frame feature combinations, at least two adjacent video frames in the above-mentioned multiple video frame sequences are grouped to generate video frame groups, and multiple video frame group sequences are obtained. Among them, the number of video frames included in the video frame group is the same as the above-mentioned number of frame feature combinations. For example, the video frame sequence includes: the first video frame, the second video frame, the third video frame, the fourth video frame, the fifth video frame, the sixth video frame, the seventh video frame, and the eighth video frame. The number of frame feature combinations may be 4. Then the corresponding first video frame group may include: the first video frame, the second video frame, the third video frame, and the fourth video frame. The corresponding first video frame group may include: the fifth video frame, the sixth video frame, the seventh video frame, and the eighth video frame.
[0037] In the fourth step, for each video frame group in the above-mentioned multiple video frame group sequences, the following first generation step is performed:
[0038] In the first sub-step, each video frame in the above-mentioned video frame group is input into a pre-trained image feature extraction model to generate video frame feature information, and a video frame feature information group is obtained. Among them, the image feature extraction model may be a neural network model for extracting image feature information. In practice, the image feature extraction model may be a convolutional neural network model connected in series in multiple layers. The image feature extraction model may be trained based on a conventional model training method.
[0039] In the second sub-step, combined feature information for the above-mentioned video frame feature information group is generated as the frame vector corresponding to the above-mentioned video frame group.
[0040] As an example, according to the chronological order, the above-mentioned execution entity may sequentially splice each video frame feature information in the above-mentioned video frame feature information group to generate a spliced vector as the frame vector corresponding to the above-mentioned video frame group.
[0041] Step 104, using the first clustering algorithm, perform clustering processing on each frame vector in the above-mentioned multiple frame vector sequences to generate a first clustering result.
[0042] In some embodiments, the above-mentioned execution entity may use the first clustering algorithm to perform clustering processing on each frame vector in the above-mentioned multiple frame vector sequences to generate a first clustering result. Among them, the first clustering algorithm may be but is not limited to at least one of the following: FCM clustering algorithm, Kmeans clustering algorithm. The first clustering result may include: a frame vector cluster set and a cluster label corresponding to each frame vector cluster.
[0043] In some alternative implementations of some embodiments, the above-mentioned clustering process for each frame vector in the above-mentioned multiple frame vector sequences by using the first clustering algorithm to generate a first clustering result may include the following steps:
[0044] First step, for each frame vector in the above-mentioned multiple frame vector sequences, extract the object vector corresponding to the frame vector.
[0045] As an example, the above-mentioned execution entity may input the frame vector into the object vector generation model included in the object information recognition model to generate an object vector. The object information recognition model may be a neural network for generating object recognition information. The object recognition information may include: object type and object position. The object vector generation model may be a neural network model for generating object vectors. The object vector may be information in the form of a vector representing the semantics related to the object in the frame image. In practice, the object vector generation model may be a multi-layer cascaded residual model.
[0046] Second step, use the first clustering algorithm to perform clustering processing on each object vector in the obtained multiple object vector sequences to generate an object vector cluster set. Among them, each object vector cluster in the object vector cluster set has at least one corresponding object information.
[0047] Third step, for each object vector cluster in the above-mentioned object vector cluster set, perform the following second generation steps:
[0048] First sub-step, determine at least one object vector whose corresponding vector distance is less than the target value with the cluster center of the above-mentioned object vector cluster as the center. Among them, the vector distance may be the cosine vector distance. The target value may be a pre-set value. For example, the target value may be 3.2.
[0049] Second sub-step, determine the distance between the above-mentioned at least one object vector and the above-mentioned cluster center to obtain at least one distance. Among them, there is a one-to-one correspondence between the object vector in the at least one object vector and the distance in the at least one distance. The distance in the at least one distance is the cosine distance.
[0050] Third sub-step, input the above-mentioned at least one object vector and the above-mentioned at least one distance into the object recognition information generation model to generate object recognition information. Among them, the object recognition information generation model may be a neural network model for generating object recognition information. In practice, the object recognition information generation model may be the YOLO model. The object recognition information may include: object identifier.
[0051] Fourth sub-step, determine the above-mentioned object recognition information as the cluster label corresponding to the above-mentioned object vector cluster.
[0052] Fourth, generate a first clustering result according to the obtained object vector clusters and the corresponding cluster label set.
[0053] As an example, the above-mentioned execution entity determines the correspondence between the object vector clusters and the cluster label set as the first clustering result.
[0054] Optionally, input the above at least one object vector and the above at least one distance into an object recognition information generation model to generate object recognition information, including the following steps:
[0055] First, input each object vector in the above at least one object vector into the above object recognition information generation model to generate initial object recognition information, obtaining at least one initial object recognition information. Among them, the initial object recognition information includes: an object category probability set corresponding to the object category set, where there is a one-to-one correspondence between the object categories in the object category set and the object category probabilities in the object category probability set. The object category probability can represent the probability situation that the object category corresponding to the object vector is the corresponding object category. The larger the object category probability, the greater the probability situation for the corresponding object category.
[0056] Second, for each object vector in the above at least one object vector, determine the distance weight corresponding to the above object vector to obtain at least one distance weight. Among them, the distance weight can represent the importance degree of the corresponding features of the corresponding vector.
[0057] As an example, the above-mentioned execution entity can determine the distance weight corresponding to the above object vector by querying the association table to obtain at least one distance weight. Among them, the association table can represent the correspondence between the distance and the distance weight.
[0058] Third, for each object category in at least one object category, perform the following first information generation step:
[0059] Sub-step 1, determine at least one object category probability corresponding to the above object category, where there is a one-to-one correspondence between the object category probabilities in the at least one object category probability and the object vectors in the at least one object vector.
[0060] Sub-step 2, perform corresponding weighted summation on the above at least one object category probability and at least one distance weight to generate a weighted summation value.
[0061] Fourth, screen out the object categories corresponding to the weighted summation values that are among the top target numbers from at least one object category to obtain an object category set.
[0062] Fifth, sort each object category in the object category set according to the weighted summation value to obtain an object category sequence.
[0063] Step 6: Determine the object vector clusters corresponding to the at least one object vector above as the target object vector clusters.
[0064] Step 7: Determine the category vectors corresponding to each object category in the object category set above to obtain a category vector set.
[0065] Step 8: For each category vector in the category vector set, perform the following second information generation steps:
[0066] First sub-step: Determine the vector distances between the target object vectors in the target object vector clusters and the category vectors above to obtain a vector distance set.
[0067] Second sub-step: According to the object category positions in the object category sequence, determine the object category positions corresponding to the category vectors as the target object category positions.
[0068] Third sub-step: Determine the pre-set distance weights corresponding to the above target object category positions as the target distance weights.
[0069] Fourth sub-step: Screen out the vector distances in the vector distance set whose corresponding vector distances are less than the preset vector distance as the target vector distances to obtain at least one target vector distance.
[0070] Fifth sub-step: Add up the respective target vector distances in the at least one target vector distance above to obtain a summed value.
[0071] Sixth sub-step: Multiply the summed value by the target distance weight to obtain a multiplied value.
[0072] Step 9: Screen out the object category with the largest corresponding multiplied distance from the object category set as the object recognition information.
[0073] The content in the above "Optionally" is an inventive point of the present disclosure, which solves the technical problem mentioned in the background art "How to accurately determine the corresponding object recognition information for at least one object vector under the same vector cluster is crucial." Based on this, in the present disclosure, first, through the object category probability, the most likely object category sequence can be determined, and then, on this basis, through the vector distance between the object category and the object vector, the object recognition information can be accurately determined.
[0074] Optionally, the above object recognition information generation model includes: an attention mechanism model based on a first number of convolutional layers, an attention mechanism model based on a second number of convolutional layers, an attention mechanism model based on a third number of convolutional layers, and an output layer. Among them, the above first number is greater than the above second number, the above second number is greater than the above third number, the output weight corresponding to the attention mechanism model based on the first number of convolutional layers is greater than the output weight corresponding to the attention mechanism model based on the second number of convolutional layers, the output weight corresponding to the attention mechanism model based on the second number of convolutional layers is greater than the output weight corresponding to the attention mechanism model based on the third number of convolutional layers, the vector level corresponding to the attention mechanism model based on the first number of convolutional layers is the first level, the vector level corresponding to the attention mechanism model based on the second number of convolutional layers is the second level, and the vector level corresponding to the attention mechanism model based on the third number of convolutional layers is the third level. For example, the first number can be 20. The second number can be 15. The third number can be 10. The attention mechanism model based on the first number of convolutional layers may include: 20 serially connected convolutional layers + a multi-head attention mechanism model. The attention mechanism model based on the second number of convolutional layers may include: 15 serially connected convolutional layers + a multi-head attention mechanism model. The attention mechanism model based on the third number of convolutional layers may include: 10 serially connected convolutional layers + a multi-head attention mechanism model. For example, the output weight corresponding to the attention mechanism model based on the first number of convolutional layers can be 0.5. The output weight corresponding to the attention mechanism model based on the second number of convolutional layers can be 0.3. The output weight corresponding to the attention mechanism model based on the third number of convolutional layers can be 0.2. The output layer can be a fully connected layer.
[0075] Optionally, the step of inputting the at least one object vector and the at least one distance into the object recognition information generation model to generate object recognition information may include the following steps:
[0076] First step, according to the at least one distance, perform vector level division on the at least one object vector to generate an object vector group set. Among them, each object vector group has a corresponding vector level, and each vector level has a corresponding distance range. For the vector levels of the first level, the second level, and the third level, the corresponding object vector group sets include: the object vector group corresponding to the first level, the object vector group corresponding to the second level, and the object vector group corresponding to the third level.
[0077] Second step, for each object vector group in the object vector group set, perform the following fifth generation step:
[0078] First sub-step, for each object vector in the object vector group, perform the following sixth generation step:
[0079] Sub-step 1: Determine at least one object vector that takes the above object vector as the center and has a vector distance less than a predetermined distance from the above object vector as at least one target object vector. For example, the predetermined distance can be 1.5.
[0080] Sub-step 2: Generate a labeled object vector corresponding to the cluster label of the above object vector cluster included in the above object vector according to the above at least one target object vector and the above object vector.
[0081] As an example, first, the above execution entity can perform vector fusion on at least one target object vector and the object vector to generate a vector set. Then, input the above vector set into an object feature extraction model to generate an object feature vector as the labeled object vector.
[0082] The second sub-step: Determine the vector level corresponding to the above object vector group as the target vector level.
[0083] The third sub-step: In response to determining that the above target vector level is the above first level, input the labeled object vector group corresponding to the above object vector group into the above attention mechanism model based on the first number of convolutional layers to generate a first attention vector.
[0084] The fourth sub-step: Multiply the above first attention vector by the output weight corresponding to the above attention mechanism model based on the first number of convolutional layers to obtain a multiplication result.
[0085] The fifth sub-step: In response to determining that the above target vector level is the above second level, input the labeled object vector group corresponding to the above object vector group into the above attention mechanism model based on the second number of convolutional layers to generate a second attention vector.
[0086] The sixth sub-step: Multiply the above second attention vector by the output weight corresponding to the above attention mechanism model based on the second number of convolutional layers to obtain a multiplication result.
[0087] The seventh sub-step: In response to determining that the above target vector level is the above third level, input the labeled object vector group corresponding to the above object vector group into the above attention mechanism model based on the third number of convolutional layers to generate a third attention vector.
[0088] The eighth sub-step: Multiply the above third attention vector by the output weight corresponding to the above attention mechanism model based on the third number of convolutional layers to obtain a multiplication result.
[0089] The third step: Add up each multiplication result in the obtained multiplication result set to obtain an addition result.
[0090] In the fourth step, input the above added result into the above output layer to generate the above object recognition information.
[0091] Step 105: According to the above first clustering result, label each frame vector in the above multiple frame vector sequences with feature labels to generate multiple first frame vector sequences with added labels.
[0092] In some embodiments, the above execution entity may label each frame vector in the above multiple frame vector sequences with feature labels according to the above first clustering result to generate multiple first frame vector sequences with added labels. Among them, the label corresponding to the added first frame vector is the cluster label corresponding to the corresponding object vector cluster.
[0093] Step 106: Use the second clustering algorithm to perform clustering processing on each first frame vector in the above multiple first frame vector sequences to generate a second clustering result.
[0094] In some embodiments, the above execution entity may use the second clustering algorithm to perform clustering processing on each first frame vector in the above multiple first frame vector sequences to generate a second clustering result. The second clustering result may include: a set of frame vector clusters after re-clustering and the cluster label corresponding to each re-clustered frame vector cluster. In practice, the second clustering algorithm may be the same clustering algorithm as the first clustering algorithm or a different clustering algorithm. For example, the second clustering algorithm may be, but is not limited to, at least one of the following: FCM clustering algorithm, Kmeans clustering algorithm.
[0095] In some optional implementation manners of some embodiments, after step 106, the steps further include:
[0096] First step: Use the above second clustering algorithm to perform clustering processing on each first frame vector in the above multiple first frame vector sequences to generate a third clustering result.
[0097] Second step: According to the above third clustering result, label each first frame vector in the above multiple first frame vector sequences with feature labels to generate multiple second frame vector sequences with added labels.
[0098] As an example, first, the above execution entity may determine the set of cluster centers corresponding to the third clustering result. Then, determine the feature label corresponding to each cluster center in the above set of cluster centers to obtain a set of feature labels. According to the set of feature labels, label each first frame vector in the above multiple first frame vector sequences with feature labels to generate multiple second frame vector sequences with added labels.
[0099] Step 3: Using the above first clustering algorithm, perform clustering processing on each of the second frame vector sequences among the multiple second frame vector sequences to generate a fourth clustering result.
[0100] Among them, the feature labels corresponding to each cluster in the cluster set corresponding to the fourth clustering result are the same. The same certain target feature label corresponding to each cluster is obtained by tagging features.
[0101] Step 4: Using the third clustering algorithm, perform clustering processing on each of the frame vectors among the multiple frame vector sequences to generate a fifth clustering result. Among them, the above third clustering algorithm is a clustering algorithm different from the above first clustering algorithm and the above second clustering algorithm. For example, the third clustering algorithm can be the DBSCAN algorithm.
[0102] Step 5: According to the above fourth clustering result and the above fifth clustering result, perform result verification on the above second clustering result to generate a verification result.
[0103] As an example, the above execution entity can determine whether there is a large difference from the clustering situation corresponding to the second clustering result according to the clustering situation corresponding to the fourth clustering result and the clustering situation corresponding to the fifth clustering result. If the difference is large, a verification result indicating non-pass of verification is generated. If the difference is small, a verification result indicating pass of verification is generated.
[0104] Step 107: In response to receiving a video acquisition request for a target object and the above second clustering result passing verification, according to the above second clustering result, obtain a video segment set corresponding to the target object from the above multiple target videos.
[0105] In some embodiments, in response to receiving a video acquisition request for a target object and the above second clustering result passing verification, the above execution entity can obtain a video segment set corresponding to the target object from the above multiple target videos according to the above second clustering result. Among them, the display objects corresponding to the video segments in the video segment set include the target object. The video acquisition request can be a request to acquire video segments corresponding to the target object.
[0106] In some optional implementation manners of some embodiments, after step 107, the steps further include:
[0107] Step 1: According to the time information corresponding to each video segment in the above video segment set, perform time alignment processing on each video segment in the above video segment set to generate an aligned video. The video frames in the aligned video are arranged in chronological order.
[0108] Second step, set video tags for the above-mentioned sorted video. Among them, the video tags can be object tags corresponding to the target object. For example, the video tag is "video clip of the target object".
[0109] Third step, store the above-mentioned sorted video and the above-mentioned video tags correspondingly in the target server. Among them, the target server can be a pre-deployed server.
[0110] In some optional implementation manners of some embodiments, the above-mentioned obtaining the video clip set corresponding to the target object from the above-mentioned multiple target videos according to the above-mentioned second clustering result may include the following steps:
[0111] First step, according to the above-mentioned second clustering result, label each first frame vector in the above-mentioned multiple first frame vector sequences after adding tags to generate multiple third frame vector sequences after adding tags. Among them, the feature tags labeled on the third frame vectors are the clustering tags corresponding to the second clustering result.
[0112] Second step, for each third frame vector in the above-mentioned multiple third frame vector sequences, perform the following fourth generation steps:
[0113] First sub-step, determine the video frame group corresponding to the above-mentioned third frame vector as the target video frame group.
[0114] Second sub-step, determine at least one feature tag corresponding to the above-mentioned third frame vector as at least one feature tag corresponding to the above-mentioned target video frame group.
[0115] Third step, screen out the video frames corresponding to the feature tags that are the object tags corresponding to the target object from the above-mentioned multiple target videos to obtain multiple video frame sets.
[0116] Fourth step, integrate each video frame in the above-mentioned multiple video frame sets to generate a video clip set.
[0117] As an example, the above-mentioned execution subject may integrate each video frame in the above-mentioned multiple video frame sets according to the time sequence to generate a video clip set.
[0118] The above-described various embodiments of the present disclosure have the following beneficial effects: Through the video acquisition method of some embodiments of the present disclosure, a video clip set related to a target object can be accurately and efficiently acquired from multiple target videos. Specifically, the reasons for the inaccurate and inefficient acquisition of relevant video clips are as follows: The efficiency of manually slicing videos is low, and the obtained video clips may not be complete. Based on this, in the video acquisition method of some embodiments of the present disclosure, first, multiple target videos for a target space are acquired as the video sources of the target object for subsequent acquisition of video clips for the target object. Among them, the multiple camera devices corresponding to the above multiple target videos are in different orientations in the above target space. Then, video frame extraction processing is performed on each of the above multiple target videos to generate a video frame sequence, and multiple video frame sequences are obtained to facilitate the determination of video clips from the frame perspective and ensure the accuracy of subsequent video clips. Next, multiple frame vector sequences for the above multiple video frame sequences can be accurately generated for subsequent clustering processing. Then, using the first clustering algorithm, clustering processing can be accurately performed on each frame vector in the above multiple frame vector sequences to generate a first clustering result. Here, through the first clustering algorithm, vectors with a certain feature can be clustered together to facilitate the subsequent determination of corresponding video clips for the target object. Furthermore, according to the above first clustering result, feature labels are added to each frame vector in the above multiple frame vector sequences to generate multiple first frame vector sequences with labels added, to facilitate subsequent second clustering processing and improve the accuracy of subsequent video clip acquisition. Further, using the second clustering algorithm, clustering processing can be accurately performed on each first frame vector in the above multiple first frame vector sequences to generate a second clustering result to obtain the label for each video frame, facilitating subsequent video clip acquisition. Finally, in response to receiving a video acquisition request for a target object and the above second clustering result passing the verification, according to the above second clustering result, a video clip set corresponding to the above target object can be accurately acquired from the above multiple target videos. In summary, through two-stage clustering processing, a label set corresponding to each frame image in multiple target videos can be obtained. Based on this, after receiving a video acquisition request for a target object and the above second clustering result passing the verification, a video clip set corresponding to the above target object can be efficiently and accurately acquired from the above multiple target videos.
[0119] Further referring to Figure 2 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a video acquisition device. These device embodiments correspond to Figure 1 the method embodiments shown, and the video acquisition device can be specifically applied to various electronic devices.
[0120] As shown in Figure 2As shown in the figure, a video acquisition device 200 includes: a first acquisition unit 201, a frame extraction unit 202, a first generation unit 203, a first execution unit 204, a tagging unit 205, a second execution unit 206, and a second acquisition unit 207. Among them, the first acquisition unit 201 is configured to acquire a plurality of target videos for a target space, where a plurality of camera devices corresponding to the plurality of target videos are in different orientations of the target space; the frame extraction unit 202 is configured to perform video frame extraction processing on each of the plurality of target videos to generate a video frame sequence, obtaining a plurality of video frame sequences; the first generation unit 203 is configured to generate a plurality of frame vector sequences for the plurality of video frame sequences; the first execution unit 204 is configured to use a first clustering algorithm to perform clustering processing on each frame vector in the plurality of frame vector sequences to generate a first clustering result; the tagging unit 205 is configured to, according to the first clustering result, tag each frame vector in the plurality of frame vector sequences with feature tags to generate a plurality of first frame vector sequences after adding tags; the second execution unit 206 is configured to use a second clustering algorithm to perform clustering processing on each first frame vector in the plurality of first frame vector sequences to generate a second clustering result; the second acquisition unit 207 is configured to, in response to receiving a video acquisition request for a target object and the second clustering result passing verification, according to the second clustering result, acquire a set of video segments corresponding to the target object from the plurality of target videos.
[0121] It can be understood that the units described in the video acquisition device 200 correspond to the respective steps in the method described in the reference Figure 1 description. Therefore, the operations, features, and beneficial effects described above for the method also apply to the video acquisition device 200 and the units included therein, and will not be elaborated here.
[0122] Next, refer to Figure 3 , which shows a schematic structural diagram of an electronic device (for example, an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.
[0123] As Figure 3As shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in the read-only memory (ROM) 302 or a program loaded from the storage device 308 into the random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.
[0124] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 an electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had. Figure 3 Each block shown in the figure may represent one device or, as needed, multiple devices.
[0125] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the methods of some embodiments of the present disclosure are performed.
[0126] It should be noted that in some embodiments of the present disclosure, the above-mentioned computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0127] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.
[0128] The above computer-readable medium may be included in the above electronic device; or it may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain a plurality of target videos for a target space, wherein a plurality of camera devices corresponding to the plurality of target videos are in different orientations in the target space; perform video frame extraction processing on each of the plurality of target videos to generate a video frame sequence, and obtain a plurality of video frame sequences; generate a plurality of frame vector sequences for the plurality of video frame sequences; use a first clustering algorithm to perform clustering processing on each frame vector in the plurality of frame vector sequences to generate a first clustering result; according to the first clustering result, label each frame vector in the plurality of frame vector sequences to generate a plurality of first frame vector sequences after adding labels; use a second clustering algorithm to perform clustering processing on each first frame vector in the plurality of first frame vector sequences to generate a second clustering result; in response to receiving a video acquisition request for a target object and the second clustering result passing verification, according to the second clustering result, obtain a set of video segments corresponding to the target object from the plurality of target videos.
[0129] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0130] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0131] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a first acquisition unit, a frame extraction unit, a first generation unit, a first execution unit, a tagging unit, a second execution unit, and a second acquisition unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring multiple target videos for a target space".
[0132] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0133] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the embodiments of the present disclosure.
Claims
1. A video acquisition method, comprising: Acquiring a plurality of target videos for a target space, wherein a plurality of camera devices corresponding to the plurality of target videos are in different orientations of the target space; Performing video frame extraction processing on each of the plurality of target videos to generate a video frame sequence, and obtaining a plurality of video frame sequences; Generating a plurality of frame vector sequences for the plurality of video frame sequences; Using a first clustering algorithm to perform clustering processing on each frame vector in the plurality of frame vector sequences to generate a first clustering result; According to the first clustering result, tagging each frame vector in the plurality of frame vector sequences with feature tags to generate a plurality of first frame vector sequences after adding tags; Using a second clustering algorithm to perform clustering processing on each first frame vector in the plurality of first frame vector sequences to generate a second clustering result; In response to receiving a video acquisition request for a target object and the second clustering result passing verification, according to the second clustering result, obtaining a video segment set corresponding to the target object from the plurality of target videos.
2. The method according to claim 1, wherein, The method further comprises: Performing time alignment processing on each video segment in the video segment set according to the time information corresponding to each video segment in the video segment set to generate an aligned video; Setting a video tag for the aligned video; Correspondingly storing the aligned video and the video tag in a target server.
3. The method according to claim 2, wherein, The generating a plurality of frame vector sequences for the plurality of video frame sequences includes: Obtaining the frame extraction frequency information corresponding to the plurality of target videos; Generating a frame feature merging number according to the frame extraction frequency information; According to the frame feature merging number, performing grouping processing on at least two adjacent video frames in the plurality of video frame sequences to generate video frame groups, and obtaining a plurality of video frame group sequences, wherein the number of video frames included in a video frame group is the same as the frame feature merging number; For each video frame group in the plurality of video frame group sequences, performing the following first generation step: Inputting each video frame in the video frame group into a pre-trained image feature extraction model to generate video frame feature information, and obtaining a video frame feature information group; Generating combined feature information for the video frame feature information group as the frame vector corresponding to the video frame group.
4. The method according to claim 3, wherein, The using a first clustering algorithm to perform clustering processing on each frame vector in the plurality of frame vector sequences to generate a first clustering result includes: For each frame vector in the plurality of frame vector sequences, extracting the object vector corresponding to the frame vector; Using a first clustering algorithm to perform clustering processing on each object vector in the obtained plurality of object vector sequences to generate an object vector cluster set; For each object vector cluster in the object vector cluster set, performing the following second generation step: Determining at least one object vector whose corresponding vector distance is less than a target value with the cluster center corresponding to the object vector cluster as the center; Determining the distance between the at least one object vector and the cluster center to obtain at least one distance; Input the at least one object vector and the at least one distance into an object recognition information generation model to generate object recognition information; Determine the object recognition information as the cluster label corresponding to the object vector cluster; generate a first clustering result according to the obtained object vector cluster set and the corresponding cluster label set.
5. The method according to claim 4, wherein After performing the clustering process on each first frame vector in the multiple first frame vector sequences by using the second clustering algorithm to generate a second clustering result, the method further includes: Perform the clustering process on each first frame vector in the multiple first frame vector sequences by using the second clustering algorithm to generate a third clustering result; Perform feature tagging on each first frame vector in the multiple first frame vector sequences after adding labels according to the third clustering result to generate multiple second frame vector sequences after adding labels; Perform the clustering process on each second frame vector in the multiple second frame vector sequences by using the first clustering algorithm to generate a fourth clustering result; Perform the clustering process on each frame vector in the multiple frame vector sequences by using a fusion clustering algorithm to generate a fifth clustering result, where the fusion clustering algorithm is a clustering algorithm different from the first clustering algorithm and the second clustering algorithm; Perform result verification on the second clustering result according to the fourth clustering result and the fifth clustering result to generate a verification result.
6. The method according to claim 5, wherein, The obtaining, according to the second clustering result, a video clip set corresponding to the target object from the multiple target videos includes: Perform feature tagging on each first frame vector in the multiple first frame vector sequences after adding labels according to the second clustering result to generate multiple third frame vector sequences after adding labels; For each third frame vector in the multiple third frame vector sequences, perform the following fourth generation step: Determine the video frame group corresponding to the third frame vector as the target video frame group; Determine at least one feature label corresponding to the third frame vector as at least one feature label corresponding to the target video frame group; Screen out the video frames from the multiple target videos whose corresponding feature labels are the object labels corresponding to the target object to obtain multiple video frame sets; Integrate each video frame in the multiple video frame sets to generate a video clip set.
7. The method according to claim 6, wherein, The object recognition information generation model includes: an attention mechanism model based on a first number of convolutional layers, an attention mechanism model based on a second number of convolutional layers, an attention mechanism model based on a third number of convolutional layers, and an output layer, where the first number is greater than the second number, the second number is greater than the third number, the output weight corresponding to the attention mechanism model based on the first number of convolutional layers is greater than the output weight corresponding to the attention mechanism model based on the second number of convolutional layers, the output weight corresponding to the attention mechanism model based on the second number of convolutional layers is greater than the output weight corresponding to the attention mechanism model based on the third number of convolutional layers, the vector level corresponding to the attention mechanism model based on the first number of convolutional layers is the first level, the vector level corresponding to the attention mechanism model based on the second number of convolutional layers is the second level, and the vector level corresponding to the attention mechanism model based on the third number of convolutional layers is the third level; and Inputting the at least one object vector and the at least one distance into the object recognition information generation model to generate object recognition information includes: Performing vector level division on the at least one object vector according to the at least one distance to generate a set of object vector groups, where each object vector group has a corresponding vector level, and each vector level has a corresponding distance range; For each object vector group in the set of object vector groups, perform the following fifth generation step: For each object vector in the object vector group, perform the following sixth generation step: Determine at least one object vector centered on the object vector and having a vector distance less than a predetermined distance from the object vector as at least one target object vector; Generating, according to the at least one target object vector and the object vector, a labeled object vector included in the object vector corresponding to the cluster label corresponding to the object vector cluster; Determine the vector level corresponding to the object vector group as the target vector level; In response to determining that the target vector level is the first level, input the labeled object vector group corresponding to the object vector group into the attention mechanism model based on the first number of convolutional layers to generate a first attention vector; Multiply the first attention vector by the output weight corresponding to the attention mechanism model based on the first number of convolutional layers to obtain a multiplication result; In response to determining that the target vector level is the second level, input the labeled object vector group corresponding to the object vector group into the attention mechanism model based on the second number of convolutional layers to generate a second attention vector; Multiply the second attention vector by the output weight corresponding to the attention mechanism model based on the second number of convolutional layers to obtain a multiplication result; In response to determining that the target vector level is the third level, input the labeled object vector group corresponding to the object vector group into the attention mechanism model based on the third number of convolutional layers to generate a third attention vector; Multiply the third attention vector by the output weights corresponding to the attention mechanism model based on the third number of convolutional layers to obtain a multiplication result; Add up the respective multiplication results in the obtained multiplication result set to obtain an addition result; Input the addition result into the output layer to generate the object recognition information.
8. A video acquisition device, comprising: A first acquisition unit configured to acquire a plurality of target videos for a target space, wherein a plurality of camera devices corresponding to the plurality of target videos are in different orientations in the target space; A frame extraction unit configured to perform video frame extraction processing on each of the plurality of target videos to generate a video frame sequence, obtaining a plurality of video frame sequences; A first generation unit configured to generate a plurality of frame vector sequences for the plurality of video frame sequences; A first execution unit configured to perform clustering processing on each frame vector in the plurality of frame vector sequences by using a first clustering algorithm to generate a first clustering result; A tagging unit configured to tag feature labels for each frame vector in the plurality of frame vector sequences according to the first clustering result to generate a plurality of first frame vector sequences after adding labels; A second execution unit configured to perform clustering processing on each first frame vector in the plurality of first frame vector sequences by using a second clustering algorithm to generate a second clustering result; A second acquisition unit configured to, in response to receiving a video acquisition request for a target object and the second clustering result passing verification, acquire a video clip set corresponding to the target object from the plurality of target videos according to the second clustering result.
9. An electronic device, comprising: One or more processors; A storage device having stored thereon one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.
10. A computer-readable medium having a computer program stored thereon, wherein, The program, when executed by a processor, implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Object clustering method and device, computer readable medium and electronic equipment
CN111667018A
Subtitle tracking method and device and electronic equipment
CN112954455A