Model optimization method and device
By utilizing the weight optimization of K target cluster centers and similar video frame evaluation in the global feature extraction model, the accuracy of global feature extraction of video frames and the model's anti-attack capability are improved, making it suitable for video duplication detection and copyright protection.
Patent Information
- Application Number
- CN202210526858.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-12
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2042-05-12
AI Technical Summary
Existing technologies have shortcomings in improving the performance and accuracy of global feature extraction models for video frames, especially when facing attacks such as black-edged frosted glass in videos.
The trained global feature extraction model is used to cluster the local features of video frames into K target cluster centers. K target cluster centers with the same weight are used for feature transformation. When the influence weight of some cluster centers is zero, similar video frames are recalled for evaluation. Finally, the model is optimized to reduce negative impacts and improve model performance and accuracy.
It improves the accuracy of extracting global features of video frames and enhances the model's ability to resist attacks on videos. It is suitable for application scenarios such as video duplication detection and copyright protection.
Smart Images

Figure CN117115489B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a model optimization method and device. Background Art
[0002] When extracting global features from video frames, a trained global feature extraction model is typically used to cluster the local features of the video frames into multiple target cluster centers. These multiple target cluster centers are model parameters of the trained global feature extraction model. The trained global feature extraction model can be obtained by training the global feature extraction model. Specifically, multiple reference cluster centers within the model parameters of the global feature extraction model can be optimized and adjusted to obtain multiple target cluster centers within the model parameters of the trained global feature extraction model. Improving the performance of the trained global feature extraction model and the accuracy of the global features extracted from video frames are current research hotspots. Summary of the Invention
[0003] The embodiments of the present application provide a model optimization method, apparatus, device, storage medium, and computer program product, which can improve model performance and increase the accuracy of global features extracted from video frames.
[0004] In one aspect, an embodiment of the present application provides a model optimization method, comprising:
[0005] For any optimized video frame in the optimized sample set, multiple local features of the optimized video frame are clustered towards K target cluster centers using the trained global feature extraction model to obtain K primary global features of the optimized video frame; the influence weight of each target cluster center in the K target cluster centers is the same; K is a positive integer;
[0006] Traversing the K target cluster centers, determining the influence weight of the t-th target cluster center currently traversed among the K target cluster centers to be zero, obtaining the updated influence weights of the K target cluster centers, and converting the K primary global features of any optimized video frame based on the updated influence weights of the K target cluster centers to obtain the global features of any optimized video frame; until the global features of each optimized video frame in the optimized sample set are obtained; t is a positive integer less than or equal to K;
[0007] Based on the global features of each optimized video frame, recall similar video frames similar to each optimized video frame from multiple candidate video frames in the database, and add each similar video frame to the tth similar video frame group;
[0008] The t-th similar video frame group is evaluated using a preset evaluation rule to obtain an evaluation value of the t-th similar video frame group, until evaluation values of K similar video frame groups are obtained; the evaluation value of the t-th similar video frame group is used to characterize the performance of the trained global feature extraction model when the influence weight of the t-th target cluster center is zero, and the K similar video frame groups correspond to the K target cluster centers one-to-one;
[0009] Comparing the K evaluation values, determining Z evaluation values from the K evaluation values, and determining the target cluster centers corresponding to the Z evaluation values; Z is a natural number less than K;
[0010] The influence weights of the Z target cluster centers corresponding to the trained global feature extraction model are optimized to zero to obtain an optimized global feature extraction model.
[0011] In one aspect, an embodiment of the present application provides a model optimization device, comprising:
[0012] a processing unit configured to cluster, for any optimized video frame in the optimized sample set, a plurality of local features of the any optimized video frame toward K target clustering centers using a trained global feature extraction model to obtain K primary global features of the any optimized video frame; wherein the influence weight of each target clustering center in the K target clustering centers is the same; and K is a positive integer;
[0013] The processing unit is further configured to traverse the K target cluster centers, determine the influence weight of the t-th target cluster center currently traversed among the K target cluster centers to be zero, obtain the updated influence weights of the K target cluster centers, and perform conversion processing on the K primary global features of any optimized video frame based on the updated influence weights of the K target cluster centers to obtain the global features of any optimized video frame; until the global features of each optimized video frame in the optimized sample set are obtained; t is a positive integer less than or equal to K;
[0014] The processing unit is further configured to recall similar video frames that are similar to the respective optimized video frames from a plurality of candidate video frames in a database based on the global features of the respective optimized video frames, and add the respective similar video frames to a t-th similar video frame group;
[0015] an evaluation unit, configured to evaluate the t-th similar video frame group using a preset evaluation rule to obtain an evaluation value for the t-th similar video frame group, until evaluation values for K similar video frame groups are obtained; the evaluation value for the t-th similar video frame group is used to characterize the performance of the trained global feature extraction model when the influence weight of the t-th target cluster center is zero, and the K similar video frame groups correspond one-to-one to the K target cluster centers;
[0016] The evaluation unit is further configured to compare the K evaluation values, determine Z evaluation values from the K evaluation values, and determine target cluster centers corresponding to the Z evaluation values; Z is a natural number smaller than K;
[0017] The optimization unit is used to optimize the influence weights of the Z target cluster centers corresponding to the trained global feature extraction model to zero, so as to obtain an optimized global feature extraction model.
[0018] In one aspect, an embodiment of the present application provides an electronic device, characterized in that the electronic device includes an input interface and an output interface, and further includes:
[0019] a processor adapted to implement one or more instructions; and
[0020] A computer storage medium storing one or more instructions, wherein the one or more instructions are suitable for being loaded by the processor and executing the above-mentioned model optimization method.
[0021] On the one hand, an embodiment of the present application provides a computer storage medium, characterized in that computer program instructions are stored in the computer storage medium, and when the computer program instructions are executed by a processor, they are used to execute the above-mentioned model optimization method.
[0022] On the one hand, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer storage medium; a processor of an electronic device reads the computer program from the computer storage medium, and the processor executes the computer program, so that the electronic device performs the above-mentioned model optimization method.
[0023] In an embodiment of the present application, the trained global feature extraction model can cluster multiple local features of the video frame to K target clustering centers to obtain K primary global features of the video frame, and convert the K primary global features of the video frame based on the influence weights of the K target clustering centers to obtain the global features of the video frame, so as to realize the extraction of the global features of the video frame, wherein the influence weights of each target clustering center in the K target clustering centers are the same; the model optimization method proposed in the embodiment of the present application points out that the trained global feature extraction model can be used to extract the global features of each optimized video frame in the optimized sample set when the influence weight of the tth target clustering center in the K target clustering centers is zero, and based on the global features of each optimized video frame, similar video frames similar to each optimized video frame are recalled from multiple candidate video frames in the database, and each similar video frame is added to the tth similar In the video frame group; the preset evaluation rules are used to evaluate the t-th similar video frame group to obtain the evaluation value of the t-th similar video frame group, until the evaluation values of K similar video frame groups are obtained; and then based on the size relationship of the K evaluation values, Z target clustering centers that have a greater impact on the model performance of the trained global feature extraction model can be determined from the K target clustering centers; the influence weights of the Z target clustering centers corresponding to the trained global feature extraction model are optimized to zero to obtain the optimized global feature extraction model, so that when the optimized global feature extraction model extracts the global features of the video frame, the negative impact of the local features of the video frame that are clustered to the Z target clustering centers on the global features of the video frame can be reduced, thereby improving the model performance and improving the accuracy of the global features of the extracted video frame, and can well deal with attacks such as black-edged frosted glass implemented on the video. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0025] Figure 1 Schematic diagram of the structure of an optimized global feature extraction model provided in an embodiment of the present application;
[0026] Figure 2 This is a flow chart of a model optimization method provided in an embodiment of the present application;
[0027] Figure 3 This is a schematic diagram of clustering local features of optimized video frames provided by an embodiment of the present application;
[0028] Figure 4 1 is a schematic diagram of extracting and optimizing global features of a video frame through a trained global feature extraction model provided by an embodiment of the present application;
[0029] Figure 5 Schematic diagram of clustering different target cluster centers among K target cluster centers provided in an embodiment of the present application;
[0030] Figure 6 This is a flowchart of a training process of a global feature extraction model provided in an embodiment of the present application;
[0031] Figure 7 This is a schematic structural diagram of a model optimization device provided in an embodiment of the present application;
[0032] Figure 8 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0034] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0035] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision (CV), speech processing, natural language processing, and machine learning (ML) / deep learning (DL).
[0036] Machine learning is a multidisciplinary field, encompassing probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to imbue computers with intelligence. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0037] Based on the machine learning technology mentioned above, an embodiment of the present application provides a model optimization scheme, which can optimize the trained global feature extraction model so that the optimized global feature extraction model has improved performance compared with the trained global feature extraction model before optimization; the trained global feature extraction model can cluster multiple local features of the video frame to K target clustering centers to obtain K primary global features of the video frame, and convert the K primary global features of the video frame based on the influence weights of the K target clustering centers to obtain the global features of the video frame, so as to realize the extraction of the global features of the video frame, wherein the influence weights of each target clustering center in the K target clustering centers are the same; the model optimization scheme proposed in the present application can extract the global features of each optimized video frame in the optimized sample set through the trained global feature extraction model when the influence weight of the tth target clustering center in the K target clustering centers is zero, and extract the global features of each optimized video frame based on the global weights of each optimized video frame. Local features, recall similar video frames that are similar to each optimized video frame from multiple candidate video frames in the database, and add each similar video frame to the tth similar video frame group; use the preset evaluation rule to evaluate the tth similar video frame group to obtain the evaluation value of the tth similar video frame group, until the evaluation values of K similar video frame groups are obtained; then the K evaluation values can be compared, Z evaluation values can be determined from the K evaluation values, and the target cluster centers corresponding to the Z evaluation values can be determined; the influence weights of the Z target cluster centers corresponding to the trained global feature extraction model are optimized to zero to obtain the optimized global feature extraction model, wherein the evaluation value of the tth similar video frame group is used to characterize the performance of the trained global feature extraction model when the influence weight of the tth target cluster center is zero, the K similar video frame groups correspond to the K target cluster centers one-to-one, K is a positive integer, and K can be set according to specific needs, t is a positive integer less than or equal to K, and Z is a natural number less than K.
[0038] The optimized global feature extraction model obtained by optimizing the trained global feature extraction model based on the above-mentioned model optimization scheme can be used to extract the global features of video frames in any video, and the extracted global features of the video frames are used as the video fingerprint of the video; the video fingerprint of the video can then be used in application scenarios such as video duplication detection and video copyright protection. For example, in the application scenario of video duplication detection, the global features of the video frames in the video to be checked for duplicates can be extracted. By comparing the global features of the video frames in the video to be checked for duplicates with the global features of the video frames in each video in the video library, it is determined whether the video to be checked for duplicates exists in the video library. For another example, in the application scenario of video copyright protection, the global features of the video frames in the video to be identified for copyright can be extracted. By comparing the global features of the video frames in the video to be identified for copyrights with the global features of the video frames in each copyrighted video on the target platform, it is determined whether the video to be identified for copyright is an infringing video.
[0039] In a specific implementation, the model optimization scheme proposed in this application can be executed by an electronic device, which can be a terminal device or a server; the terminal device here may include but is not limited to: computers, smart phones, tablets, laptops, intelligent voice interaction devices, smart home appliances, vehicle terminals, smart wearable devices, etc.; the server here can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0040] In one embodiment, the trained global feature extraction model can be optimized based on the above-mentioned model optimization scheme to obtain an optimized global feature extraction model. The trained global feature extraction model can be obtained based on the training of the global feature extraction model. That is, the global feature extraction model can be optimized and trained to obtain a trained global feature extraction model, and then the trained global feature extraction model can be optimized to obtain an optimized global feature extraction model. The global feature extraction model, the trained global feature extraction model, and the optimized global feature extraction model have the same structure but different model parameters. The model parameters of the trained global feature extraction model are obtained by optimizing and adjusting the model parameters of the global feature extraction model, and the model parameters of the optimized global feature extraction model are obtained by optimizing and adjusting the model parameters of the trained global feature extraction model. See. Figure 1, is a schematic structural diagram of an optimized global feature extraction model provided in an embodiment of the present application; Figure 1 The optimized global feature extraction model shown can include a local feature extraction module, a feature clustering module and a global feature generation module; wherein the local feature extraction module can be used to extract local features of a video frame, the feature clustering module can be used to cluster the local features of a video frame to obtain K primary global features of the video frame, and the global feature generation module can be used to convert the K primary global features of the video frame to obtain the global features of the video frame. That is, when extracting the global features of a video frame through the optimized global feature extraction model, the local features of the video frame can be extracted through the local feature extraction module in the optimized global feature extraction model; then, the local features of the video frame can be clustered through the feature clustering module in the optimized global feature extraction model to obtain K primary global features of the video frame; then, the K primary global features of the video frame can be converted through the global feature generation module in the optimized global feature extraction model to obtain the global features of the video frame.
[0041] Optionally, the local feature extraction module in the optimized global feature extraction model can be a convolutional neural network (CNN), the Inception network in the convolutional neural network, the residual network (ResNet), the visual geometry group network (VGG), and a local feature extraction network with more accurate local features (such as ASLFeat and D2Net, a local feature extraction network based on feature point detection), etc., which can perform local feature extraction, generally including a downsampling layer, an upsampling layer, a convolution layer and a nonlinear activation layer; in actual applications, the specific structure of the local feature extraction module can be flexibly selected according to specific needs; for the sake of convenience of explanation, the embodiments of this application will be explained later with the local feature extraction module as a convolutional neural network. The feature clustering module can be a related structure of a local aggregated descriptor vector implemented based on a network (i.e., NetVLAD Layer); the feature clustering module in the optimized global feature extraction model can include a convolution layer, a probability prediction layer, a local aggregated descriptor vector core (Vector of Locally Aggregated Descriptors core, VLAD core), a first regularization layer (i.e., intra-normalization layer), and a second regularization layer (i.e., L2 normalization layer), wherein the convolution layer can be a 1*1 convolution layer, and the probability prediction layer can be a softmax function layer.
[0042] Optionally, the global feature generation module may include multiple dimensionality reduction submodules and a normalization submodule; multiple dimensionality reduction submodules may be used to perform dimensionality reduction processing, and the normalization submodule may be used to normalize the dimensionality reduction processing results output by the multiple dimensionality reduction submodules to obtain the global features of the video frame. The introduction of the normalization submodule may ensure the distinguishability of the global features of the video frame. Optionally, the dimensionality reduction submodule may be a multilayer perceptron (MLP), an autoencoder network (AutoencoderNet), a principal component analysis network (PCANet), and the like; for ease of explanation, the embodiment of the present application will be subsequently introduced as a multilayer perceptron using the dimensionality reduction submodule, wherein the multilayer perceptron may be composed of a fully connected layer, a normalization layer, and an activation function layer, wherein the normalization layer may be a Batch Norm layer, and the activation function layer may be a ReLU layer; the normalization submodule may be a NORM layer. Among them, multiple dimensionality reduction sub-modules are connected in sequence, and among the multiple dimensionality reduction sub-modules, except for the penultimate dimensionality reduction sub-module, the input of the latter two correspondingly connected dimensionality reduction sub-modules is the output of the previous dimensionality reduction sub-module; the input of the penultimate dimensionality reduction sub-module is the weighted processing result between the output of the second-to-last dimensionality reduction sub-module and the output of the first dimensionality reduction sub-module; that is, a jump connection is made between the first dimensionality reduction sub-module and the penultimate dimensionality reduction sub-module in the global feature generation module (i.e., a shortcut connection is used); for example, if the output of the first dimensionality reduction sub-module is represented by y1, the corresponding weight is represented by h1, the output of the second-to-last dimensionality reduction sub-module is represented by y2, and the corresponding weight is represented by h2, then the input of the penultimate dimensionality reduction sub-module can be expressed as h1*y1+h2*y2. In the global feature generation module, the gradient divergence problem can be avoided by connecting the first dimensionality reduction submodule with the last dimensionality reduction submodule with a shortcut. The number of dimensionality reduction submodules in the global feature generation module can be set according to specific needs. The more dimensionality reduction submodules there are, the more hidden layer neurons there are, and the better the quality of the global features of the generated video frames. The fewer dimensionality reduction submodules there are, the faster the global features of the video frames are generated. Therefore, the number of dimensionality reduction submodules can be set based on different needs. For example, the number of dimensionality reduction submodules can be set to 4 or 5. Figure 1 The number of dimensionality reduction submodules in the optimized global feature extraction model shown is 4.
[0043] It is particularly important to note that in the specific implementation of this application, user-related data is involved, for example, when the video frames are video frames in a video generated by the user, when the embodiments of this application are applied to specific products or technologies, user permission or consent must be obtained, and the collection, use and processing of relevant data must comply with local laws, regulations and standards.
[0044] Based on the above model optimization solution, the present application embodiment provides a model optimization method. Figure 2 , which is a flow chart of a model optimization method provided in an embodiment of the present application. Figure 2 The model optimization method shown can be executed by an electronic device. Figure 2 The model optimization method shown may include the following steps:
[0045] S201, for any optimized video frame in the optimized sample set, cluster multiple local features of any optimized video frame to K target clustering centers through the trained global feature extraction model to obtain K primary global features of any optimized video frame.
[0046] Among them, the trained global feature extraction model can be obtained based on the training of the global feature extraction model. The relevant process of optimizing the global feature extraction model to obtain the trained global feature extraction model will be introduced in the subsequent embodiments and will not be expanded here. The trained global feature extraction model has the same structure as the global feature extraction model, but different model parameters. The model parameters of the trained global feature extraction model are obtained by optimizing and adjusting the model parameters of the global feature extraction model; the model parameters of the global feature extraction model include K reference cluster centers and the influence weight of each reference cluster center in the K reference cluster centers, K is a positive integer, and K can be set according to specific needs. The model parameters of the trained global feature extraction model include K target cluster centers and the influence weight of each target cluster center in the K target cluster centers; wherein, the K target cluster centers are obtained by optimizing and adjusting the K reference cluster centers, and the influence weight of each target cluster center is the same as the influence weight of each reference cluster center, that is, In the process of optimizing and training the global feature extraction model to obtain the trained global feature extraction model, the influence weights of each reference cluster center will not be optimized and adjusted; further, the influence weights of each reference cluster center are the same, and the influence weights of each reference cluster center can be set according to specific needs. The influence weights of each target cluster center in the K target cluster centers corresponding to the trained global feature extraction model are the same. The embodiment of the present application does not limit the influence weights of each reference cluster center and the influence weights of each target cluster center corresponding to the trained global feature extraction model; generally speaking, the influence weight of each reference cluster center can be set to 1. For the sake of convenience, the embodiment of the present application will be described later when the influence weight of each reference cluster center is 1.
[0047] In one embodiment, the optimized sample set includes one or more optimized video frames, and the optimized sample set is used to optimize the trained global feature extraction model. The optimized video frames included in the optimized sample set are different from the sample video frames included in the training sample set used to optimize the global feature extraction model.
[0048] In one embodiment, for any optimized video frame in the optimized sample set, the electronic device clusters multiple local features of any optimized video frame to K target cluster centers through a trained global feature extraction model to obtain K primary global features of any optimized video frame. This may include: performing texture feature extraction processing on any optimized video frame through a trained global feature extraction model to obtain multiple texture features of any optimized video frame, and using the multiple texture features of any optimized video frame as local features of any optimized video frame; based on the allocation probability of each local feature of any optimized video frame to each target cluster center in the K target cluster centers, and the residual between each local feature of any optimized video frame and each target cluster center, clustering each local feature of any optimized video frame to obtain K primary global features of any optimized video frame.
[0049] In one embodiment, the electronic device performs texture feature extraction processing on any optimized video frame through a trained global feature extraction model to obtain multiple texture features of any optimized video frame, and the related process of using the multiple texture features of any optimized video frame as local features of any optimized video frame can be implemented by a local feature extraction module in the trained global feature extraction model; when the local feature extraction module is a convolutional neural network, the multiple texture features of any optimized video frame obtained can be: when the convolutional neural network performs convolution processing on the any optimized video frame, the feature map (Feature Map) output by the shallow convolution layer of the convolutional neural network. Because in the related field of video fingerprints, more attention is paid to the texture features of video frames than to the high-level semantic features of video frames, and because the semantic feature expression ability of convolutional neural networks increases with the increase of the number of network layers, but the texture feature expression ability decreases with the increase of the number of network layers, the features extracted by the shallow convolution layer in the convolutional neural network can capture richer texture information in the video frame. When the convolutional neural network performs convolution processing on the video frame, the feature map output by the shallow convolution layer of the convolutional neural network can be used as the texture feature of the video frame. When the global features of the video frame are extracted based on the texture features of the video frame, the problem of weak discrimination ability of high-level semantic features can be effectively avoided. Furthermore, if the convolutional neural network performs convolution processing on the video frame, the feature map output by the shallow convolution layer of the convolutional neural network is W H*D-dimensional feature maps (i.e., W*H*DFeature Maps), then the W H*D-dimensional feature maps can be used as the N (i.e., W*H) D-dimensional texture features of the video frame obtained by the convolutional neural network, and then the N texture features of the video frame are used as the local features of the video frame, that is, N local features of the video frame are obtained, where W and H are both positive integers.
[0050] Furthermore, the electronic device can cluster the local features of any optimized video frame based on the probability of each local feature of any optimized video frame being assigned to each target cluster center in K target cluster centers, and the residuals between each local feature of any optimized video frame and each target cluster center through the feature clustering module in the trained global feature extraction model, so as to obtain K primary global features of any optimized video frame; in a specific implementation, the electronic device can perform convolution processing on each local feature of any optimized video frame to obtain the convolution results corresponding to each local feature of any optimized video frame; based on any optimized video frame, the local features of any optimized video frame can be clustered to obtain the K primary global features ... The convolution results corresponding to the local features of the optimized video frame predict the distribution probability of the local features of any optimized video frame to each target cluster center; for any target cluster center among the target cluster centers, based on the distribution probability of each local feature of any optimized video frame to any target cluster center, the residuals between each local feature of any optimized video frame and any target cluster center are weighted summed up to obtain the primary global features corresponding to any target cluster center; until the primary global features corresponding to each target cluster center are obtained, they are used as the K primary global features of any optimized video frame.
[0051] In one embodiment, in the process of clustering the local features of any optimized video frame by the feature clustering module (i.e., NetVLAD) in the trained global feature extraction model to obtain K primary global features of any optimized video frame, the probability of each local feature of any optimized video frame being assigned to each target cluster center can be obtained by performing exponential normalization processing on the convolution results corresponding to each local feature of the optimized video frame through the probability prediction layer (i.e., softmax layer) in the feature clustering module; in the process of predicting the allocation probability based on each local feature of any optimized video frame, the texture complexity of each local feature of any optimized video frame can be fully referenced, that is, the texture complexity of each local feature of any optimized video frame can be referenced. The texture complexity of each local feature of the video frame is used to determine the probability of each local feature of the optimized video frame being assigned to each target cluster center. Based on the allocation probability, the local features of the optimized video frame can be clustered into different target cluster centers. Then, based on the global feature generation module in the trained global feature extraction model, in the process of converting the K primary global features obtained by the feature clustering module to obtain global features, the primary global features obtained by clustering local features with higher texture complexity are processed with higher weight parameters, and the primary global features obtained by clustering local features with lower texture complexity are processed with lower weight parameters to improve the discrimination of the obtained global features. Moreover, in the process of clustering the local features of the optimized video frame by the feature clustering module (i.e., NetVLAD) in the trained global feature extraction model to obtain the K primary global features of the optimized video frame, the residual between the local features of the optimized video frame and the target cluster centers can be used to erase the feature distribution differences of the optimized video frame itself, and only retain the distribution differences between the local features of the optimized video frame and the target cluster centers. See Figure 3 , which is a schematic diagram of clustering local features of an optimized video frame provided in an embodiment of the present application, wherein the local features of an optimized video frame are shown as marked 301, and a local feature of the optimized video frame can be shown as marked 302. If there are two target clustering centers, one target clustering center is shown as marked 303, and the other target clustering center is shown as marked 304, then the two primary global features of the optimized video frame obtained by clustering can be shown as marked 305; it can be seen that when clustering the local features of the optimized video frame, the clustering process is independent of the position of the local features, so it can cope with various spatial attacks; and similar local features tend to be assigned to the cluster subspace corresponding to the same target cluster center, making the aggregation more compact; during clustering, the differences between local features can be fully explored through the residual, so that the differences between features are fully retained.
[0052] If the N texture features of the optimized video frame are extracted by the local feature extraction module in the trained global feature extraction model, the vector dimension of each texture feature is D-dimensional, and the N local features of the optimized video frame are used as the vector dimension of each local feature is D-dimensional; if the i-th local feature of the N local features of the optimized video frame is x i Indicates that the jth eigenvalue in the i-th local feature of the optimized video frame is x i (j) indicates that the kth target cluster center among the K target cluster centers is represented by c k Indicates that the jth eigenvalue in the kth target cluster center is c k (j) represents, where j∈D and k∈K; then the jth eigenvalue of the primary global feature corresponding to the kth target cluster center can be expressed by the following formula 1:
[0053]
[0054] Among them, in the primary global features corresponding to the k-th target cluster center, the j-th eigenvalue is represented by V(j, k), represents the convolution result corresponding to the i-th local feature, Represents the convolution result corresponding to the i-th local feature among N local features, Express Sum from k' to K, x i (j)-c k (j) represents the residual between the jth eigenvalue in the i-th local feature and the jth eigenvalue in the k-th target cluster center; w k , b k and c k is the model parameter of the trained global feature extraction model, which is obtained by optimizing and adjusting the relevant model parameters in the global feature extraction model. k =2αc k , b k =-α||c k || 2 , α is a hyperparameter.
[0055] In one embodiment, the convolution layer in the feature clustering module in the trained global feature extraction model can be used to perform convolution processing on the various local features of the optimized video frame; the probability prediction layer (i.e., softmax layer) in the feature clustering module can be used to predict the allocation probability of the various local features of the optimized video frame to each target cluster center; and the VLAD core in the feature clustering module can be used to determine the K primary global features of the optimized video frame based on the allocation probability and the residual. Furthermore, for the K primary global features of the optimized video frame output by the VLAD core, the first regularization layer in the feature clustering module can be used to perform regularization processing on each of the K primary global features obtained to erase the absolute size of the residual in the cluster and only retain the distribution of the residual; further, the output of the first regularization layer can be L2 regularized as a whole through the second regularization layer to obtain K primary global features after quadratic regularization (that is, the K*D-dimensional feature vector output by the first regularization layer is L2 regularized); it is worth noting that in the subsequent process of converting the K primary global features of the optimized video frame to obtain the global features of the optimized video frame, the K primary global features of the optimized video frame should be the K primary global features of the optimized video frame after quadratic regularization.
[0056] S202, traverse K target cluster centers, determine the influence weight of the tth target cluster center currently traversed among the K target cluster centers to be zero, obtain the updated influence weights of the K target cluster centers, and convert the K primary global features of any optimized video frame based on the updated influence weights of the K target cluster centers to obtain the global features of any optimized video frame; until the global features of each optimized video frame in the optimized sample set are obtained.
[0057] In one embodiment, the electronic device traverses K target cluster centers, determines the influence weight of the tth target cluster center currently traversed among the K target cluster centers to be zero, and in the process of obtaining the updated influence weights of the K target cluster centers, t is a positive integer less than or equal to K; if the influence weight of each target cluster center in the K target cluster centers is 1, K is 5, and t is 1, then the updated influence weights of the K target cluster centers are: 0, 1, 1, 1, 1 respectively.
[0058] In one embodiment, the electronic device converts K primary global features of any optimized video frame based on the updated influence weights of the K target cluster centers to obtain the global features of any optimized video frame, which may include: weighting the K primary global features of any optimized video frame based on the updated influence weights of the K target cluster centers to obtain K weighted global features of any optimized video frame; performing dimensionality reduction processing on the K weighted global features of any optimized video frame to obtain the dimensionality reduction processing results corresponding to any optimized video frame; and normalizing the dimensionality reduction processing results corresponding to any optimized video frame to obtain the global features of any optimized video frame. Among them, since the updated influence weights of the K target cluster centers include the updated influence weight of the t-th target cluster center to be 0, the K primary global features of the optimized video frame are weighted based on the updated influence weights of the K target cluster centers, and the t-th weighted global feature among the K weighted global features of the optimized video frame is 0; further, the global features of the optimized video frame obtained based on the K weighted global features of the optimized video frame are the global features obtained without including the t-th primary global features of the optimized video frame. In a specific implementation, multiple dimensionality reduction submodules in the global feature generation module in the trained global feature extraction model can be called to perform dimensionality reduction processing on the K weighted global features of the optimized video frame to obtain the dimensionality reduction processing results corresponding to the optimized video frame; the normalization submodule in the global feature generation module can be called to perform normalization processing on the dimensionality reduction processing results corresponding to the optimized video frame to obtain the global features of the optimized video frame; wherein, the weight parameters in the global feature generation module in the trained global feature extraction model can be obtained based on optimizing and adjusting the corresponding weight parameters in the model parameters of the global feature extraction model, and the global feature generation module in the trained global feature extraction model obtained after optimization and adjustment is used to convert the K primary global features obtained by the feature clustering module. In the process of obtaining global features, the primary global features obtained by clustering local features with higher texture complexity can be processed with higher weight parameters, and the primary global features obtained by clustering local features with lower texture complexity can be processed with lower weight parameters to improve the discrimination of the obtained global features; that is, specifically, when converting the K weighted global features obtained based on K primary global features to obtain global features, the weighted global features corresponding to the primary global features obtained by clustering local features with higher texture complexity can be processed with higher weight parameters, and the weighted global features corresponding to the primary global features obtained by clustering local features with lower texture complexity can be processed with lower weight parameters to improve the discrimination of the obtained global features.
[0059] See also Figure 4, which is a schematic diagram of an embodiment of the present application for extracting global features of an optimized video frame through a trained global feature extraction model; a local feature extraction module in the trained global feature extraction model can be used to perform texture feature extraction on the optimized video frame to obtain multiple texture features of the optimized video frame, and the multiple texture features of the optimized video frame are used as local features of the optimized video frame; a feature clustering module in the trained global feature extraction model is used to allocate the probability of each target cluster center in K target cluster centers based on each local feature of the optimized video frame, and the probability of each local feature of the optimized video frame being allocated to each target cluster center in K target cluster centers is used to allocate the probability of each local feature of the optimized video frame to each target cluster center in K target cluster centers. The residuals between the local features of the optimized video frame and the target cluster centers are used to cluster the local features of the optimized video frame to obtain K primary global features of the optimized video frame; the global feature generation module in the trained global feature extraction model performs weighted processing on the K primary global features of the optimized video frame based on the updated influence weights of the K target cluster centers to obtain K weighted global features of the optimized video frame; the K weighted global features of the optimized video frame are subjected to dimensionality reduction processing to obtain the dimensionality reduction processing results corresponding to the optimized video frame; the dimensionality reduction processing results corresponding to the optimized video frame are normalized to obtain the global features of the optimized video frame.
[0060] S203 , based on the global features of each optimized video frame, recall similar video frames that are similar to each optimized video frame from multiple candidate video frames in the database, and add each similar video frame to the tth similar video frame group.
[0061] In one embodiment, the database can be a commonly used database in related fields such as image duplication checking and video fingerprinting, or a database constructed by a technician; the electronic device recalls similar video frames similar to each optimized video frame from multiple candidate video frames in the database based on the global features of each optimized video frame, which can include: extracting the global features of each candidate video frame in the database using a trained global feature extraction model when the influence weight of the t-th target cluster center among the K target cluster centers is zero; for any optimized video frame among the optimized video frames, determining the candidate video frame with the greatest similarity between the global features and the global features of any optimized video frame as a similar video frame similar to any optimized video frame; until a similar video frame similar to each optimized video frame is obtained. The process of extracting the global features of each candidate video frame in the database using the trained global feature extraction model when the influence weight of the t-th target cluster center among the K target cluster centers is zero is similar to the process of extracting the global features of the optimized video frame using the trained global feature extraction model when the influence weight of the t-th target cluster center among the K target cluster centers is zero, and will not be repeated here.
[0062] Furthermore, for any optimized video frame among the optimized video frames, when the electronic device determines the candidate video frame with the greatest similarity between its global features and the global features of any optimized video frame as a similar video frame similar to any optimized video frame, it can first determine the similarity between the global features of the optimized video frame and the global features of each candidate video frame, and determine the candidate video frame indicated by the maximum similarity as a similar video frame similar to the optimized video frame; wherein the similarity between the global features of the optimized video frame and the global features of each candidate video frame can be the cosine similarity between the global features of the optimized video frame and the global features of each candidate video frame, or can be the similarity determined based on the feature distance between the global features of the optimized video frame and the global features of each candidate video frame, for example, the feature distance can be Euclidean distance, Manhattan distance, etc., which is not limited in the embodiments of the present application.
[0063] S204 , evaluating the t-th similar video frame group using a preset evaluation rule to obtain an evaluation value of the t-th similar video frame group, until evaluation values of K similar video frame groups are obtained.
[0064] Among them, the preset evaluation rules can be set according to specific needs; the evaluation value of the t-th similar video frame group can be used to characterize the performance of the trained global feature extraction model when the influence weight of the t-th target cluster center is zero, and the obtained K similar video frame groups correspond one to one with the K target cluster centers, that is, the evaluation value of the t-th similar video frame group in the K similar video frame groups is determined when the influence weight of the t-th target cluster center in the K target cluster centers is zero. The evaluation value of the t-th similar video frame group can be used to characterize the performance of the trained global feature extraction model when the influence weight of the t-th target cluster center is zero, that is, the evaluation value of the t-th similar video frame group can be used to characterize the performance of the trained global feature extraction model without including the t-th primary global feature obtained by clustering to the t-th target cluster center, and can be used to characterize the influence of the t-th target cluster center on the performance of the trained global feature extraction model.
[0065] In one embodiment, the preset evaluation rules can be set according to specific needs, and are used to determine the specific method of the evaluation value of the similar video frame group; if the evaluation value is the accuracy rate, the preset evaluation rules are used to indicate the specific method of determining the accuracy rate of the similar video frame group. Optionally, the evaluation value may include but is not limited to one or more of the following: accuracy rate, recall rate, F1 value (F1score); when the evaluation value includes two or more, one can be selected from the multiple evaluation values as the final evaluation value, which is used to compare K evaluation values in the subsequent process, determine Z evaluation values from the K evaluation values, and determine the target cluster center corresponding to the Z evaluation values. For example, when the preset evaluation rules are used to evaluate the t-th similar video frame group, the accuracy rate (Accuracy), recall rate ( When the Precision and Recall rates are calculated, the obtained accuracy of the t-th similar video frame group can be used as the evaluation value of the t-th similar video frame group; optionally, multiple evaluation values can be processed to obtain a final evaluation value, which is used to compare K evaluation values in a subsequent process, determine Z evaluation values from the K evaluation values, and determine the target cluster centers corresponding to the Z evaluation values. For example, the F1 value determined based on the precision and recall rate of the t-th similar video frame group can be used as the evaluation value of the t-th similar video frame group.
[0066] S205 , comparing the K evaluation values, determining Z evaluation values from the K evaluation values, and determining target cluster centers corresponding to the Z evaluation values.
[0067] Among them, the determined Z evaluation values are greater than or equal to the remaining evaluation values excluding the Z evaluation values among the K evaluation values, and Z is a natural number less than K. That is to say, when the K evaluation values are arranged from large to small, the first Z evaluation values can be selected from the K evaluation values. Since the evaluation value of the t-th similar video frame group can be used to characterize the performance of the trained global feature extraction model without including the t-th primary global feature obtained by clustering to the t-th target cluster center, in other words, the evaluation value of the t-th similar video frame group can be used to characterize the influence of the t-th target cluster center on the performance of the trained global feature extraction model, when the evaluation value is larger, it represents that the performance of the trained global feature extraction model after removing the t-th target cluster center is better, that is, when the t-th primary global feature obtained by clustering to the t-th target cluster center is removed, the performance of the obtained global feature is better, that is, it can be represented that the local features clustered to the t-th target cluster center have a greater negative impact on the obtained global features. Therefore, when the K evaluation values are arranged from large to small, the influence weights of the target cluster centers corresponding to the first Z evaluation values are optimized to zero. This is so that when the global feature extraction process is subsequently performed on the video frame, the negative impact of the local features of the video frame that are clustered towards the Z target cluster centers on the global features of the video frame is reduced, thereby improving the accuracy of extracting the global features of the video frame. Generally speaking, among the local features of the video frame, the local features that are clustered towards the Z target cluster centers determined among the K target cluster centers are mostly local features extracted from areas with less texture information, such as background areas, black edge areas, or frosted glass areas. Therefore, optimizing the influence weights of the Z target cluster centers to zero can reduce the negative impact of the local features of the video frame that are clustered towards the Z target cluster centers on the global features of the video frame, and can effectively deal with attacks such as black edge frosted glass implemented on the video.
[0068] See also Figure 5 , is a schematic diagram of clustering different target cluster centers among K target cluster centers provided in an embodiment of the present application, such as Figure 5 The mark 501 shows a video frame, the image shown by the mark 502 is a visualization representation of the K primary global features obtained by clustering the video frame to each of the K target cluster centers, and the image shown by the mark 503 is a visualization representation of the Z primary global features obtained by clustering the video frame to the Z target cluster centers determined among the K target cluster centers, wherein, among the various local features of the video frame, the local features clustered to the Z target cluster centers determined among the K target cluster centers are mostly local features extracted from the background area.
[0069] S206 , optimizing the influence weights of the Z target cluster centers corresponding to the trained global feature extraction model to zero, thereby obtaining an optimized global feature extraction model.
[0070] In one embodiment, the model parameters of the optimized global feature extraction model include: K target cluster centers obtained by optimizing and adjusting the K reference cluster centers in the model parameters of the global feature extraction model, and target influence weights of the K target cluster centers obtained by optimizing the influence weights of the K target cluster centers corresponding to the trained global feature extraction model; in the model parameters of the optimized global feature extraction model, the target influence weights of the Z target cluster centers among the K target cluster centers are zero, and the rest are the same as the model parameters of the trained global feature extraction model. In a feasible real-time method, the relevant weight parameters for processing the Z primary global features of the video frame obtained by clustering the local features of the video frame to the Z target cluster centers in the first dimensionality reduction submodule of the global feature generation module of the optimized global feature extraction model can be set to zero. When the dimensionality reduction submodule is a multi-layer perceptron, the relevant weight parameters for processing the Z primary global features of the video frame in the fully connected layer of the first multi-layer perceptron can be set to zero.
[0071] In one embodiment, an optimized global feature extraction model can be used to extract global features from video frames in any video and use the extracted global features as the video fingerprint of the video. This video fingerprint can then be used in application scenarios such as video duplication checking and video copyright protection. In a specific implementation, an electronic device can obtain a target video from which the video fingerprint is to be extracted; invoke the optimized global feature extraction model to extract global features from the target video frame based on the target influence weights of K target cluster centers, the probability of each local feature of a target video frame being assigned to each of the K target cluster centers, and the residuals between each local feature of the target video frame and each target cluster center; and use the global features of the target video frame as the video fingerprint of the target video. The target video can be any legitimate video. For example, in the application scenario of video duplication checking, the target video can be the video to be checked for duplication; in the application scenario of video copyright protection, the target video can be the video to be copyrighted. Optionally, the target video frame in the target video can be any video frame in the target video, a key frame in the target video, or a video frame determined based on preset rules.
[0072] In a specific implementation, when extracting the global features of a target video frame, the electronic device can perform texture feature extraction processing on the target video frame through an optimized global feature extraction model to obtain multiple texture features of the target video frame, and use the multiple texture features of the target video frame as local features of the target video frame; based on the allocation probability of each local feature of the target video frame to each target cluster center in K target cluster centers, and the residuals between each local feature of the target video frame and each target cluster center, the local features of the target video frame are clustered to obtain K primary global features of the target video frame; based on the target influence weights of the K target cluster centers, the K primary global features of the target video frame are weighted to obtain K weighted global features of the target video frame; the K weighted global features of the target video frame are dimensionality reduced to obtain the dimensionality reduced processing results corresponding to the target video frame; the dimensionality reduced processing results corresponding to the target video frame are normalized to obtain the global features of the target video frame.
[0073] In an embodiment of the present application, the trained global feature extraction model can cluster multiple local features of the video frame to K target clustering centers to obtain K primary global features of the video frame, and convert the K primary global features of the video frame based on the influence weights of the K target clustering centers to obtain the global features of the video frame, so as to realize the extraction of the global features of the video frame, wherein the influence weights of each target clustering center in the K target clustering centers are the same; the model optimization method proposed in the embodiment of the present application points out that the trained global feature extraction model can be used to extract the global features of each optimized video frame in the optimized sample set when the influence weight of the tth target clustering center in the K target clustering centers is zero, and based on the global features of each optimized video frame, similar video frames similar to each optimized video frame are recalled from multiple candidate video frames in the database, and each similar video frame is added to the tth similar In the video frame group; the preset evaluation rules are used to evaluate the t-th similar video frame group to obtain the evaluation value of the t-th similar video frame group, until the evaluation values of K similar video frame groups are obtained; and then based on the size relationship of the K evaluation values, Z target clustering centers that have a greater impact on the model performance of the trained global feature extraction model can be determined from the K target clustering centers; the influence weights of the Z target clustering centers corresponding to the trained global feature extraction model are optimized to zero to obtain the optimized global feature extraction model, so that when the optimized global feature extraction model extracts the global features of the video frame, the negative impact of the local features of the video frame that are clustered to the Z target clustering centers on the global features of the video frame can be reduced, thereby improving the model performance and improving the accuracy of the global features of the extracted video frame, and can well deal with attacks such as black-edged frosted glass implemented on the video.
[0074] Based on the relevant embodiments of the above-mentioned model optimization method, the embodiment of the present application introduces the relevant process of optimizing the global feature extraction model to obtain the trained global feature extraction model; see Figure 6 , is a flowchart of a training process of a global feature extraction model provided in an embodiment of the present application. The training process of the global feature extraction model can be performed by any electronic device and may specifically include the following steps:
[0075] S601, obtaining a training sample set; the training sample set includes a plurality of sample video frames.
[0076] The sample video frames included in the training sample set are different from the optimized video frames included in the optimization sample set.
[0077] In one embodiment, the sample video frames included in the training sample set can be extracted from one or more reference videos; the training sample set includes one or more training sample subsets, and any training sample subset in the training sample set includes three sample video frames, and the three sample video frames included in any training sample subset are respectively a reference sample video frame, a similar sample video frame and a difference sample video frame; the similar sample video frames included in any training sample subset are similar to the reference sample video frames, and the difference sample video frames included are not similar to the reference sample video frames; the multiple sample video frames included in the training sample set can be extracted from one or more reference videos. In a specific implementation, the electronic device obtains the training sample set, which may include: obtaining one or more reference videos, and segmenting each reference video based on the video scenes included in each reference video in the one or more reference videos to obtain each video scene One or more corresponding reference video clips; for any reference video clip corresponding to any video scene in each video scene, two reference video frames are selected from each reference video frame of any reference video clip to construct a reference video frame pair, and the reference video frame pair is added to the video frame pair set corresponding to any video scene; until the video frame pair set corresponding to each video scene is obtained; based on the video frame pair set corresponding to each video scene, one or more training sample subsets are constructed; wherein, the reference sample video frame and the similar sample video frame in any training sample subset of the one or more training sample subsets are respectively two reference video frames in a reference video frame pair, and the difference sample video frame in any training sample subset is selected from other video frame pair sets that are different from the video frame pair set to which the reference sample video frame in any training sample subset belongs; a training sample set containing one or more training sample subsets is constructed.
[0078] In one embodiment, when an electronic device acquires a training sample set, it may perform scene detection on each reference video after acquiring one or more reference videos; and based on the video scenes contained in each detected reference video, it may segment each reference video to obtain one or more reference video segments corresponding to each video scene. When constructing a training sample subset, the electronic device may select a reference video frame pair from a video frame pair set corresponding to a video scene, and use the two reference video frames in the selected reference video frame pair as reference sample video frames and similar sample video frames, respectively. It may also select a reference video frame from a reference video frame pair set corresponding to another video scene as a difference sample video frame, and then construct a training sample subset based on the reference sample video frame, the similar sample video frame, and the difference sample video frame. Optionally, the difference sample video frame selected from the video frame pair set corresponding to another video scene may belong to a different reference video than the determined reference sample video frame and the similar sample video frame. The model training method proposed in the embodiment of the present application can be trained in a semi-supervised training manner. Compared with the training method that requires a training set with category label information, the semi-supervised training method can quickly construct a training sample subset from multiple unlabeled reference videos, and further construct a training sample set; further optionally, a self-supervised training method can also be adopted to obtain a training sample set through data enhancement (data argument).
[0079] In one embodiment, for any reference video segment corresponding to any video scene in each video scene, the electronic device selects two reference video frames from each reference video frame of any reference video segment to construct a reference video frame pair, which may include: performing a scale-invariant feature transformation on each reference video frame of any reference video segment to obtain reference features of each reference video frame; constructing an initial reference video frame pair based on any two reference video frames whose similarity between corresponding reference features in each reference video frame is greater than a preset similarity threshold; determining the interval distance between two reference video frames in each initial reference video frame pair in any reference video segment; and selecting the initial reference video frame pair corresponding to the largest interval distance from each initial reference video frame pair as the reference video frame pair.
[0080] Among them, the reference features of the reference video frame are the SIFT features obtained after performing a scale-invariant feature transform (SIFT) on the reference video frame. Optionally, the similarity between the reference features can be represented by the cosine similarity between the reference features, or can be determined based on the feature distance between the reference features (such as Euclidean distance, Manhattan distance), which is not limited in the embodiment of the present application; the preset similarity threshold can be set based on specific needs. Optionally, the distance between the two video frames in the video to which the two video frames belong can be represented by the number of video frames between the two video frames in the video to which the two video frames belong, that is, the number of reference video frames between the two reference video frames in each initial reference video frame pair in any reference video clip can be determined as the distance between the two reference video frames in each initial reference video frame pair in any reference video clip. For example, if there is a reference video segment A, the initial reference video frame pairs corresponding to the reference video segment A are determined to be: initial reference video frame pair A1, initial reference video frame pair A2, and initial reference video frame pair A3. If the number of reference video frames between the two reference video frames in the initial reference video frame pair A1 in the reference video segment A is 6 frames, the number of reference video frames between the two reference video frames in the initial reference video frame pair A2 in the reference video segment A is 4 frames, and the number of reference video frames between the two reference video frames in the initial reference video frame pair A3 in the reference video segment A is 12 frames, then the initial reference video frame pair A3 can be used as the reference video frame pair determined from the reference video segment A.
[0081] In one embodiment, in order to improve the robustness of the trained global feature extraction model, attacks such as cropping, rotation, scaling and filtering can be randomly added to the sample video frames included in the training sample set, that is, the sample video frames included in the training sample set can be randomly selected, and one or more operations such as cropping, rotation, scaling and filtering can be performed on them to obtain an updated training sample set, and then the global feature extraction model can be optimized and trained based on the updated training sample set to obtain the trained global feature extraction model.
[0082] S602: Initialize model parameters of the global feature extraction model.
[0083] In one embodiment, the model parameters of the global feature extraction model include K reference cluster centers and the influence weight of each reference cluster center in the K reference cluster centers. The influence weight of each reference cluster center is the same as the influence weight of each target cluster center; that is, the K reference cluster centers will be continuously optimized and adjusted during the training process of the global feature extraction model, but the influence weight of each reference cluster center will not be adjusted; further, the influence weight of each reference cluster center is the same, and the influence weight of each reference cluster center can be set according to specific needs. The embodiment of the present application does not limit the influence weight of each reference cluster center; generally speaking, the influence weight of each reference cluster center can be set to 1. For the sake of convenience, the embodiment of the present application will be described later when the influence weight of each reference cluster center is 1.
[0084] In one feasible embodiment, the model parameters of the global feature extraction model can be randomly initialized. In another feasible embodiment, the model parameters of the global feature extraction model can be initialized based on the distribution of each local feature of the sample video frames included in the training sample set, specifically the K reference cluster centers in the model parameters of the global feature extraction model are initialized. In a specific implementation, the electronic device can select at least one sample video frame from multiple sample video frames as the target sample video frame, and perform texture feature extraction processing on each target sample video frame in the at least one target sample video frame to obtain multiple texture features of each target sample video frame, and use the multiple texture features of each target sample video frame as the local features of each target sample video frame; based on the feature distance between each local feature of each target sample video frame and each initial cluster center in the K initial cluster centers, iteratively update each initial cluster center to obtain K iterative cluster centers; initialize the K iterative cluster centers as the K reference cluster centers in the model parameters of the global feature extraction model. The K initial cluster centers can be randomly initialized, and the vector dimension of an initial cluster center is the same as the vector dimension of a local feature of the target sample video frame.
[0085] In a specific implementation, the electronic device can perform texture feature extraction processing on each target sample video frame in at least one target sample video frame through the local feature extraction module in the global feature extraction model, obtain multiple texture features of each target sample video frame, and use the multiple texture features of each target sample video frame as the local features of each target sample video frame. This process is similar to the related process of extracting the local features of the optimized video frame through the trained global feature extraction model in the above S201, and will not be repeated here. The electronic device iteratively updates each initial cluster center based on the feature distances between each local feature of each target sample video frame and each initial cluster center in the K initial cluster centers. When obtaining the K iterative cluster centers, it can be implemented based on the K-means algorithm (i.e., the K-means algorithm); that is, the electronic device can calculate the feature distances between each local feature of each target sample video frame and each initial cluster center, and assign each local feature of each target sample video frame to the cluster corresponding to the initial cluster center with the shortest corresponding feature distance, until each local feature of all target sample video frames is assigned to the K clusters; based on the local features included in each cluster, the initial cluster center corresponding to each cluster is updated; based on the updated initial cluster centers, the above process is repeated to iterate the initial cluster centers until convergence, and the K updated initial cluster centers at the time of convergence are used as the K reference cluster centers. Optionally, the feature distance between any local feature of any target sample video frame and any initial cluster center can be the Euclidean distance, cosine distance, etc. between the local feature of the target sample video frame and the initial cluster center, which is not limited in the embodiments of the present application. Optionally, when the initial cluster center corresponding to each cluster is updated based on the local features included in each cluster, each updated initial cluster center can be the centroid of each cluster determined based on the local features included in each cluster. Using K iterative cluster centers instead of the random initialization of the K reference cluster centers in the model parameters of the global feature extraction model can reduce the optimized distance of the K target cluster centers obtained by optimizing and adjusting the K reference cluster centers, thereby reducing the training time of the global feature extraction model and increasing the training efficiency.
[0086] S603, extracting global features of each sample video frame in the training sample set through the global feature extraction model, and optimizing and training the global feature extraction model based on the global features of each sample video frame to optimize and adjust the model parameters of the global feature extraction model to obtain a trained global feature extraction model.
[0087] Among them, the global features of any sample video frame in the training sample set are obtained based on the influence weights of each reference cluster center, the probability of each local feature of any sample video frame being assigned to each reference cluster center in K reference cluster centers, and the residuals between each local feature of any sample video frame and each reference cluster center.
[0088] In a specific implementation, the electronic device can use a global feature extraction model to extract the global features of the reference sample video frames, similar sample video frames and difference sample video frames in the training sample subset in the training sample set; based on the global features of the reference sample video frames, the global features of the similar sample video frames and the global features of the difference sample video frames, the global feature extraction model is optimized and trained to optimize and adjust the model parameters of the global feature extraction model to obtain a trained global feature extraction model.
[0089] In one embodiment, the relevant processes for extracting global features of reference sample video frames, similar sample video frames, and difference sample video frames in a training sample subset are similar. The embodiments of the present application are introduced with the extraction of global features of reference sample video frames. In a specific implementation, the electronic device can perform texture feature extraction processing on the reference sample video frame through a global feature extraction model to obtain multiple texture features of the reference sample video frame, and use the multiple texture features of the reference sample video frame as local features of the reference sample video frame; based on the allocation probability of each local feature of the reference sample video frame being assigned to each reference cluster center, and the residual between each local feature of the reference sample video frame and each reference cluster center, the local features of the reference sample video frame are clustered to obtain K primary global features of the reference sample video frame; the K primary global features of the reference sample video frame are converted to obtain the global features of the reference sample video frame.
[0090] If N texture features of any sample video frame are extracted by the local feature extraction module in the global feature extraction model, the vector dimension of each texture feature is D-dimensional, and the N local features of the sample video frame are used as the vector dimension of each local feature is D-dimensional; if the i-th local feature of the N local features of the sample video frame is represented by x′ i Indicates that the jth feature value in the i-th local feature of the sample video frame is x′ i (j) indicates that the kth reference cluster center among the K reference cluster centers is c′ k Indicates that the j-th eigenvalue in the k-th reference cluster center is c′ k (j) represents, where j∈D and k∈K; then the jth eigenvalue of the primary global feature corresponding to the kth reference cluster center can be expressed by the following formula 2:
[0091]
[0092] Among them, in the primary global features corresponding to the k-th reference cluster center, the j-th eigenvalue is represented by V′(j, k), represents the convolution result corresponding to the i-th local feature, Represents the convolution result corresponding to the i-th local feature among N local features, k′∈K, Express Sum from k' to K, x' i (j)-c′ k (j) represents the residual between the jth eigenvalue in the i-th local feature and the jth eigenvalue in the k-th reference cluster center; w′ k , b′ k and c′ k is the model parameter that needs to be optimized and adjusted in the global feature extraction model, where w′ k =2αc′ k , b′ k =-α||c′ k || 2 , α is a hyperparameter.
[0093] In one embodiment, the convolution layer in the feature clustering module in the global feature extraction model can be used to perform convolution processing on the local features of the reference sample video frame; the probability prediction layer in the feature clustering module can be used to predict the local features of the reference sample video frame based on the convolution results corresponding to the local features of the reference sample video frame, and the allocation probability of the local features to the reference cluster centers can be predicted; and the VLAD core in the feature clustering module can be used to determine the K primary global features of the reference sample video frame based on the allocation probability and the residual. Furthermore, for the K primary global features of the reference sample video frame output by the VLAD core, the first regularization layer in the feature clustering module can be used to perform regularization processing on each of the K primary global features obtained to erase the absolute size of the residual in the cluster and only retain the distribution of the residual; further, the output of the first regularization layer can be L2 regularized as a whole through the second regularization layer to obtain K primary global features after quadratic regularization (that is, the K*D-dimensional feature vector output by the first regularization layer is L2 regularized); it is worth noting that in the subsequent process of converting the K primary global features of the reference video frame to obtain the global features of the reference video frame, the K primary global features of the reference video frame should be the K primary global features of the reference sample video frame after quadratic regularization.
[0094] In one embodiment, since the influence weight of each reference cluster center is 1, the electronic device can call the global feature generation module in the global feature extraction model to convert the K primary global features of the reference sample video frame to obtain the global features of the reference sample video frame. In a specific implementation, multiple dimensionality reduction sub-modules in the feature generation module can be called to perform dimensionality reduction processing on the K primary global features of the reference sample video frame to obtain the dimensionality reduction processing results corresponding to the reference sample video frame; and call the normalization sub-module to normalize the dimensionality reduction processing results corresponding to the reference sample video frame to obtain the global features of the reference sample video frame.
[0095] In one embodiment, the electronic device optimizes and trains the global feature extraction model based on the global features of the reference sample video frame, the global features of the similar sample video frame, and the global features of the difference sample video frame to optimize and adjust the model parameters of the global feature extraction model to obtain the trained global feature extraction model, which may include: determining the global feature distance between the reference sample video frame and the similar sample video frame, and the global feature distance between the reference sample video frame and the difference sample video frame based on the global features of the reference sample video frame, the global features of the similar sample video frame, and the global features of the difference sample video frame; determining the model loss value of the global feature extraction model based on the global feature distance between the reference sample video frame and the similar sample video frame, and the global feature distance between the reference sample video frame and the difference sample video frame; optimizing and adjusting the model parameters of the global feature extraction model in the direction of reducing the model loss value of the global feature extraction model to obtain the trained global feature extraction model.
[0096] In one embodiment, the electronic device can determine the global feature distance between the reference sample video frame and the similar sample video frame, as well as the global feature distance between the reference sample video frame and the difference sample video frame based on the global features of the reference sample video frame, the global features of the similar sample video frame, and the global features of the difference sample video frame; determine the model loss value of the global feature extraction model based on the global feature distance between the reference sample video frame and the similar sample video frame, as well as the global feature distance between the reference sample video frame and the difference sample video frame; optimize and adjust the model parameters of the global feature extraction model in the direction of reducing the model loss value of the global feature extraction model to obtain a trained global feature extraction model. Specifically, the feature distance between the global features of the reference sample video frame and the global features of the similar sample video frame can be determined as the global feature distance between the reference sample video frame and the similar sample video frame, and the feature distance between the global features of the reference sample video frame and the global features of the difference sample video frame can be determined as the global feature distance between the reference sample video frame and the difference sample video frame; wherein, the feature distance can be Euclidean distance, cosine distance, etc., which is not limited in the embodiment of the present application.
[0097] Furthermore, if the global features of the reference sample video frame are Represented by, the global features of similar sample video frames are expressed as Represented by, the global features of the difference sample video frame are expressed as Represented as , the ternary loss function (TripletLoss) can be used as the optimization constraint, that is, the ternary loss function can be used as the loss function during the training of the global feature extraction model to determine the model loss value of the global feature extraction model; based on the global feature distance between the reference sample video frame and the similar sample video frame, and the global feature distance between the reference sample video frame and the difference sample video frame, the model loss value of the global feature extraction model can be determined as shown in the following formula 3:
[0098]
[0099] Among them, the model loss value of the global feature extraction model is composed of L t express, represents the global feature distance between the reference sample video frame and the similar sample video frame, It represents the global feature distance between the reference sample video frame and the difference sample video frame, and margin represents the preset threshold, which can be set according to the training requirements. The ternary loss function can ensure that the global feature distance between the reference sample video frame and the similar sample video frame is smaller than the global feature distance between the reference sample video frame and the difference sample video frame.
[0100] Furthermore, the electronic device can optimize and adjust the model parameters of the global feature extraction model in the direction of reducing the model loss value of the global feature extraction model to obtain a trained global feature extraction model. Since the ternary loss function is used as the optimization constraint, the model parameters of the global feature extraction model are optimized and adjusted in the direction of reducing the model loss value of the global feature extraction model, that is, the model parameters of the global feature extraction model are optimized and adjusted in the direction of reducing the global feature distance between the reference sample video frame and the similar sample video frame, and increasing the global feature distance between the reference sample video frame and the difference sample video frame to obtain a trained global feature extraction model. The model parameters of the trained global feature extraction model include: K target cluster centers obtained by optimizing and adjusting the K reference cluster centers in the model parameters of the global feature extraction model. Further optionally, during the training process of the global feature extraction model, the gradients of the positive samples and the negative samples can be separated (that is, the gradients of the positive samples and the negative samples can be separated). and gradients) to make the training process more stable.
[0101] In one embodiment, if a training sample set includes multiple training sample subsets, the training sample subsets in the training sample set can be input into the global feature extraction model in batches to optimize the global feature extraction model in batches, that is, the training sample subsets in the same batch share the same model parameters of the global feature extraction model, and the same model parameters of the global feature extraction model are optimized and adjusted based on the training sample subsets in the same batch. Optionally, when optimizing the global feature extraction model in batches based on the training sample subsets in the training sample set, the difference sample video frames included in any training sample subset in the same batch can be selected from the reference video frame pairs corresponding to each training sample subset in the same batch based on the selection rules of the difference sample video frames. Furthermore, in order to improve the accuracy of the global features extracted by the trained global feature extraction model of the video frames, each training sample subset in the same batch can be constructed based on the method of hard-to-distinguish sample mining to increase the training difficulty of the global feature extraction model, so that the global features of the video frames extracted by the trained global feature extraction model are more discriminative. In a feasible embodiment, the training sample subsets in each training sample subset of the same batch, in which the similarity between the reference features of the difference sample video frame and the reference features of the reference sample video frame is greater than a first similarity threshold and less than a preset similarity threshold, can be determined as difficult-to-distinguish training sample subsets, and the global feature extraction model can be trained based on the difficult-to-distinguish training sample subsets of each batch determined from each training sample subset of each batch. In another feasible embodiment, the first M training sample subsets in each training sample subset of the same batch, when arranged from large to small according to the similarity between the reference features of the difference sample video frame and the reference features of the reference sample video frame, can be determined as difficult-to-distinguish training sample subsets, where M is a positive integer. For example, the first 5 training sample subsets in each training sample subset of the same batch, when arranged from large to small according to the similarity between the reference features of the difference sample video frame and the reference features of the reference sample video frame, can be determined as difficult-to-distinguish training sample subsets; and the global feature extraction model can be trained based on the difficult-to-distinguish training sample subsets of each batch determined from each training sample subset of each batch.
[0102] In an embodiment of the present application, reference video frames can be selected from one or more reference videos as reference sample video frames, similar sample video frames, and difference sample video frames, respectively, and training sample subsets can be constructed based on the reference sample video frames, similar sample video frames, and difference sample video frames, thereby constructing a training sample set containing one or more training sample subsets; after obtaining the training sample set, the global features of the reference sample video frames, similar sample video frames, and difference sample video frames in the training sample subsets in the training sample set can be extracted respectively through a global feature extraction model, and the model loss value of the global feature extraction model can be determined based on the global feature distance between the reference sample video frame and the similar sample video frame, and the global feature distance between the reference sample video frame and the difference sample video frame; the model parameters of the global feature extraction model are optimized and adjusted in the direction of reducing the model loss value of the global feature extraction model to obtain a trained global feature extraction model. Since the ternary loss function is used as the optimization constraint of the global feature extraction model, the accuracy of the global features of the video frames extracted by the trained global feature extraction model can be improved, and the amount of data during model training can be reduced; and for the initialization of the model parameters of the global feature extraction model, at least one sample video frame can be selected from the multiple sample video frames included in the training sample set as the target sample video frame, and multiple texture features of each target sample video frame can be extracted as the local features of each target sample video frame; then, based on the feature distance between each local feature of each target sample video frame and each initial cluster center in the K initial cluster centers, each initial cluster center can be iteratively updated to obtain K iterative cluster centers; the K iterative cluster centers are initialized as K reference cluster centers in the model parameters of the global feature extraction model; using K iterative cluster centers instead of the random initialization of the K reference cluster centers in the model parameters of the global feature extraction model can reduce the optimization distance of the K target cluster centers obtained by optimizing and adjusting the K reference cluster centers, thereby reducing the training time of the global feature extraction model and increasing the training efficiency.
[0103] Based on the above-mentioned embodiments related to the model optimization method, the present application provides a model optimization device. Figure 7 , is a structural diagram of a model optimization device provided in an embodiment of the present application. The model optimization device may include a processing unit 701, an evaluation unit 702 and an optimization unit 703. Figure 7 The model optimization device shown can run the following units:
[0104] The processing unit 701 is configured to cluster, for any optimized video frame in the optimized sample set, multiple local features of the optimized video frame toward K target cluster centers using a trained global feature extraction model to obtain K primary global features of the optimized video frame; each of the K target cluster centers has the same influence weight; K is a positive integer;
[0105] The processing unit 701 is further configured to traverse the K target cluster centers, determine the influence weight of the t-th target cluster center currently traversed among the K target cluster centers to be zero, obtain the updated influence weights of the K target cluster centers, and perform conversion processing on the K primary global features of any optimized video frame based on the updated influence weights of the K target cluster centers to obtain the global features of any optimized video frame; until the global features of each optimized video frame in the optimized sample set are obtained; t is a positive integer less than or equal to K;
[0106] The processing unit 701 is further configured to recall similar video frames that are similar to the respective optimized video frames from a plurality of candidate video frames in a database based on the global features of the respective optimized video frames, and add the respective similar video frames to a t-th similar video frame group;
[0107] An evaluation unit 702 is configured to evaluate the t-th similar video frame group using a preset evaluation rule to obtain an evaluation value for the t-th similar video frame group, until evaluation values for K similar video frame groups are obtained; the evaluation value for the t-th similar video frame group is used to represent the performance of the trained global feature extraction model when the influence weight of the t-th target cluster center is zero, and the K similar video frame groups correspond one-to-one to the K target cluster centers;
[0108] The evaluation unit 702 is further configured to compare the K evaluation values, determine Z evaluation values from the K evaluation values, and determine target cluster centers corresponding to the Z evaluation values; Z is a natural number smaller than K;
[0109] The optimization unit 703 is used to optimize the influence weights of the Z target cluster centers corresponding to the trained global feature extraction model to zero, so as to obtain an optimized global feature extraction model.
[0110] In one embodiment, the processing unit 701 clusters the multiple local features of any optimized video frame toward K target cluster centers using the trained global feature extraction model to obtain K primary global features of any optimized video frame, and specifically performs the following operations:
[0111] Performing texture feature extraction processing on any of the optimized video frames using the trained global feature extraction model to obtain multiple texture features of the any of the optimized video frames, and using the multiple texture features of the any of the optimized video frames as local features of the any of the optimized video frames;
[0112] Based on the various local features of any optimized video frame, the probability of being assigned to each target cluster center in the K target cluster centers, and the residuals between the various local features of any optimized video frame and the various target cluster centers, the various local features of any optimized video frame are clustered to obtain K primary global features of any optimized video frame.
[0113] In one embodiment, the processing unit 701 clusters the local features of any optimized video frame based on the probability of each local feature of the optimized video frame being assigned to each of the K target cluster centers, and the residuals between each local feature of the optimized video frame and each target cluster center, to obtain K primary global features of the optimized video frame, and specifically performs the following operations:
[0114] Performing convolution processing on each local feature of any one of the optimized video frames to obtain a convolution result corresponding to each local feature of any one of the optimized video frames;
[0115] Predicting the probability of each local feature of any optimized video frame being assigned to each target cluster center based on the convolution result corresponding to each local feature of any optimized video frame;
[0116] For any target cluster center among the target cluster centers, based on the allocation probability of each local feature of the any optimized video frame to the any target cluster center, a weighted sum operation is performed on the residuals between each local feature of the any optimized video frame and the any target cluster center to obtain the primary global features corresponding to the any target cluster center; until the primary global features corresponding to the each target cluster center are obtained, which are used as the K primary global features of the any optimized video frame.
[0117] In one embodiment, the processing unit 701 converts the K primary global features of any optimized video frame based on the updated influence weights of the K target cluster centers to obtain the global features of any optimized video frame by specifically performing the following operations:
[0118] Performing weighted processing on the K primary global features of any one of the optimized video frames based on the updated influence weights of the K target cluster centers to obtain K weighted global features of any one of the optimized video frames;
[0119] Performing dimensionality reduction processing on the K weighted global features of any one of the optimized video frames to obtain a dimensionality reduction processing result corresponding to the any one of the optimized video frames;
[0120] Normalizing the dimensionality reduction processing result corresponding to any one of the optimized video frames to obtain the global features of any one of the optimized video frames.
[0121] In one embodiment, the processing unit 701 specifically performs the following operations when recalling similar video frames similar to each optimized video frame from multiple candidate video frames in the database based on the global features of each optimized video frame:
[0122] When the influence weight of the t-th target cluster center among the K target cluster centers is zero, extracting global features of each candidate video frame in the database by using the trained global feature extraction model;
[0123] For any optimized video frame among the optimized video frames, a candidate video frame whose global features have the greatest similarity with the global features of any optimized video frame is determined as a similar video frame similar to the any optimized video frame; until a similar video frame similar to each optimized video frame is obtained.
[0124] In one embodiment, the evaluation values include but are not limited to one or more of the following: accuracy, recall rate, F1 value; the Z evaluation values are greater than or equal to the remaining evaluation values of the K evaluation values except the Z evaluation values.
[0125] In one embodiment, the trained global feature extraction model is obtained based on training the global feature extraction model; the processing unit 701 is further configured to:
[0126] Acquire a training sample set; the training sample set includes a plurality of sample video frames, and the sample video frames included in the training sample set are different from the optimized video frames included in the optimized sample set;
[0127] Initializing model parameters of the global feature extraction model; the model parameters of the global feature extraction model include K reference cluster centers and the influence weight of each reference cluster center in the K reference cluster centers, and the influence weight of each reference cluster center is the same as the influence weight of each target cluster center;
[0128] The global features of each sample video frame in the training sample set are extracted by the global feature extraction model, and the global feature extraction model is optimized and trained based on the global features of each sample video frame to optimize and adjust the model parameters of the global feature extraction model to obtain the trained global feature extraction model; the global features of any sample video frame in the training sample set are obtained based on the influence weights of the respective reference cluster centers, the respective local features of any sample video frame, the probability of being assigned to each reference cluster center in the K reference cluster centers, and the residuals between the respective local features of any sample video frame and the respective reference cluster centers.
[0129] In one embodiment, the training sample set includes one or more training sample subsets, and any training sample subset in the training sample set includes three sample video frames, and the three sample video frames included in any training sample subset are respectively a reference sample video frame, a similar sample video frame, and a difference sample video frame;
[0130] The processing unit 701 extracts global features of each sample video frame in the training sample set through the global feature extraction model, and optimizes and trains the global feature extraction model based on the global features of each sample video frame to optimize and adjust model parameters of the global feature extraction model. When the trained global feature extraction model is obtained, the processing unit 701 specifically performs the following operations:
[0131] Extracting global features of reference sample video frames, similar sample video frames, and difference sample video frames in the training sample subset in the training sample set respectively through the global feature extraction model;
[0132] Based on the global features of the reference sample video frame, the global features of the similar sample video frame and the global features of the difference sample video frame, the global feature extraction model is optimized and trained to optimize and adjust the model parameters of the global feature extraction model to obtain the trained global feature extraction model.
[0133] In one embodiment, the processing unit 701 optimizes and trains the global feature extraction model based on the global features of the reference sample video frame, the global features of the similar sample video frame, and the global features of the difference sample video frame to optimize and adjust the model parameters of the global feature extraction model. When the trained global feature extraction model is obtained, the processing unit 701 specifically performs the following operations:
[0134] Determining a global feature distance between the reference sample video frame and the similar sample video frame, and a global feature distance between the reference sample video frame and the difference sample video frame based on the global features of the reference sample video frame, the global features of the similar sample video frame, and the global features of the difference sample video frame;
[0135] Determining a model loss value of the global feature extraction model based on a global feature distance between the reference sample video frame and the similar sample video frame, and a global feature distance between the reference sample video frame and the difference sample video frame;
[0136] In the direction of reducing the model loss value of the global feature extraction model, the model parameters of the global feature extraction model are optimized and adjusted to obtain the trained global feature extraction model.
[0137] In one embodiment, the influence weight of each reference cluster center is 1; when the processing unit 701 extracts the global features of the reference sample video frames in the training sample subset in the training sample set through the global feature extraction model, the processing unit 701 specifically performs the following operations:
[0138] Performing texture feature extraction processing on the reference sample video frame using the global feature extraction model to obtain multiple texture features of the reference sample video frame, and using the multiple texture features of the reference sample video frame as local features of the reference sample video frame;
[0139] Based on the allocation probabilities of the local features of the reference sample video frame to the reference cluster centers, and the residuals between the local features of the reference sample video frame and the reference cluster centers, the local features of the reference sample video frame are clustered to obtain K primary global features of the reference sample video frame;
[0140] The K primary global features of the reference sample video frame are converted to obtain the global features of the reference sample video frame.
[0141] In one embodiment, the global feature extraction model includes a global feature generation module, which includes multiple dimensionality reduction submodules and a normalization submodule; the multiple dimensionality reduction submodules are connected in sequence, and among the multiple dimensionality reduction submodules, except for the penultimate dimensionality reduction submodule, the input of the remaining two correspondingly connected dimensionality reduction submodules is the output of the previous dimensionality reduction submodule; the input of the penultimate dimensionality reduction submodule is the weighted processing result between the output of the penultimate dimensionality reduction submodule and the output of the first dimensionality reduction submodule;
[0142] The processing unit 701 performs the following operations when converting the K primary global features of the reference sample video frame to obtain the global features of the reference sample video frame:
[0143] Calling the multiple dimensionality reduction submodules to perform dimensionality reduction processing on the K primary global features of the reference sample video frame to obtain a dimensionality reduction processing result corresponding to the reference sample video frame;
[0144] The normalization submodule is called to perform normalization processing on the dimensionality reduction processing result corresponding to the reference sample video frame to obtain the global features of the reference sample video frame.
[0145] In one embodiment, when the processing unit 701 obtains the training sample set, it specifically performs the following operations:
[0146] Obtaining one or more reference videos, and segmenting each of the one or more reference videos based on a video scene contained in each reference video to obtain one or more reference video segments corresponding to each video scene;
[0147] For any reference video segment corresponding to any video scene in each of the video scenes, selecting two reference video frames from each reference video frame of the reference video segment to construct a reference video frame pair, and adding the reference video frame pair to a set of video frame pairs corresponding to the any video scene; until a set of video frame pairs corresponding to each of the video scenes is obtained;
[0148] The one or more training sample subsets are constructed based on the video frame pair sets corresponding to the respective video scenes; the reference sample video frames and the similar sample video frames in any training sample subset of the one or more training sample subsets are respectively two reference video frames in a reference video frame pair, and the difference sample video frames in any training sample subset are selected from other video frame pair sets that are different from the video frame pair set to which the reference sample video frames in any training sample subset belong;
[0149] A training sample set including the one or more training sample subsets is constructed.
[0150] In one embodiment, when the processing unit 701 selects two reference video frames from the reference video frames of any reference video segment to construct a reference video frame pair, the processing unit 701 specifically performs the following operations:
[0151] Performing scale-invariant feature transformation on each reference video frame of any reference video segment to obtain reference features of each reference video frame;
[0152] Constructing an initial reference video frame pair based on any two reference video frames whose corresponding reference features have a similarity greater than a preset similarity threshold in the reference video frames;
[0153] determining a distance between two reference video frames in each initial reference video frame pair in any one of the reference video segments;
[0154] From the various initial reference video frame pairs, the initial reference video frame pair corresponding to the maximum interval distance is used as the reference video frame pair.
[0155] In one embodiment, when the processing unit 701 initializes the model parameters of the global feature extraction model, it specifically performs the following operations:
[0156] selecting at least one sample video frame from the plurality of sample video frames as a target sample video frame, performing texture feature extraction processing on each target sample video frame in the at least one target sample video frame to obtain a plurality of texture features of each target sample video frame, and using the plurality of texture features of each target sample video frame as a local feature of each target sample video frame;
[0157] Based on the feature distances between each local feature of each target sample video frame and each initial cluster center in the K initial cluster centers, the initial cluster centers are iteratively updated to obtain K iterative cluster centers;
[0158] The K iterative cluster centers are initialized as K reference cluster centers in the model parameters of the global feature extraction model.
[0159] According to one embodiment of the present application, Figure 2 as well as Figure 6 The various steps involved in the model optimization method shown can be Figure 7 The model optimization device shown is executed by each unit. For example, Figure 2 Steps S201 to S203 shown in FIG. Figure 7 The processing unit 701 in the model optimization device shown is executed, Figure 2 Steps S204 to S205 shown in FIG. Figure 7 The evaluation unit 702 in the model optimization device shown is executed. Figure 2 Step S206 shown can be performed by Figure 7 The optimization unit 703 in the model optimization device shown in FIG. Figure 6 Steps S601 to S603 shown in FIG. Figure 7 The processing unit 701 in the model optimization device shown is executed,
[0160] According to another embodiment of the present application, Figure 7 The various units in the model optimization device shown can be separately or all merged into one or several other units to constitute, or a certain (some) unit therein can also be split into multiple smaller units in function to constitute, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In practical applications, the function of a unit can also be realized by multiple units, or the function of multiple units can be realized by one unit. In other embodiments of the present application, the model optimization device based on logical function division can also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.
[0161] According to another embodiment of the present application, the program can be executed by running on a general computing device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM) and other processing elements and storage elements. Figure 2 as well as Figure 6 A computer program (including program code) for each step of the corresponding method shown in FIG. Figure 7 The model optimization device shown in and the model optimization method of the embodiment of the present application are implemented. The computer program can be recorded on, for example, a computer-readable storage medium, and loaded into the above-mentioned computing device through the computer-readable storage medium and run therein.
[0162] In an embodiment of the present application, the trained global feature extraction model can cluster multiple local features of the video frame to K target clustering centers to obtain K primary global features of the video frame, and convert the K primary global features of the video frame based on the influence weights of the K target clustering centers to obtain the global features of the video frame, so as to realize the extraction of the global features of the video frame, wherein the influence weights of each target clustering center in the K target clustering centers are the same; the model optimization method proposed in the embodiment of the present application points out that the trained global feature extraction model can be used to extract the global features of each optimized video frame in the optimized sample set when the influence weight of the tth target clustering center in the K target clustering centers is zero, and based on the global features of each optimized video frame, similar video frames similar to each optimized video frame are recalled from multiple candidate video frames in the database, and each similar video frame is added to the tth similar In the video frame group; the preset evaluation rules are used to evaluate the t-th similar video frame group to obtain the evaluation value of the t-th similar video frame group, until the evaluation values of K similar video frame groups are obtained; and then based on the size relationship of the K evaluation values, Z target clustering centers that have a greater impact on the model performance of the trained global feature extraction model can be determined from the K target clustering centers; the influence weights of the Z target clustering centers corresponding to the trained global feature extraction model are optimized to zero to obtain the optimized global feature extraction model, so that when the optimized global feature extraction model extracts the global features of the video frame, the negative impact of the local features of the video frame that are clustered to the Z target clustering centers on the global features of the video frame can be reduced, thereby improving the model performance and improving the accuracy of the global features of the extracted video frame, and can well deal with attacks such as black-edged frosted glass implemented on the video.
[0163] Based on the above-mentioned related embodiments of the model optimization method and the model optimization device embodiment, the present application also provides an electronic device. Figure 8 , is a structural diagram of an electronic device provided in an embodiment of the present application. Figure 8 The electronic device shown may include at least a processor 801, an input interface 802, an output interface 803, and a computer storage medium 804. The processor 801, the input interface 802, the output interface 803, and the computer storage medium 804 may be connected via a bus or other means.
[0164] Computer storage medium 804 can be stored in the memory of the electronic device. Computer storage medium 804 is used to store computer programs, which include program instructions. Processor 801 is used to execute the program instructions stored in computer storage medium 804. Processor 801 (or CPU (Central Processing Unit)) is the computing core and control core of the electronic device. It is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the above-mentioned model optimization method process or corresponding functions.
[0165] The embodiment of the present application also provides a computer storage medium (Memory), which is a memory device in an electronic device for storing programs and data. It is understandable that the computer storage medium here can include both the built-in storage medium in the terminal and, of course, the extended storage medium supported by the terminal. The computer storage medium provides a storage space that stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor 801 are also stored in the storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer storage medium here can be a high-speed random access memory (RAM) memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage; optionally, it can also be at least one computer storage medium located away from the aforementioned processor.
[0166] In one embodiment, the processor 801 and the input interface 802 can load and execute one or more instructions stored in the computer storage medium to implement the above-mentioned Figure 2 as well as Figure 6 In the corresponding steps of the method in the embodiment of the model optimization method, in a specific implementation, one or more instructions in the computer storage medium are loaded by the processor 801 and the following steps are executed:
[0167] For any optimized video frame in the optimized sample set, multiple local features of the optimized video frame are clustered towards K target cluster centers using the trained global feature extraction model to obtain K primary global features of the optimized video frame; the influence weight of each target cluster center in the K target cluster centers is the same; K is a positive integer;
[0168] Traversing the K target cluster centers, determining the influence weight of the t-th target cluster center currently traversed among the K target cluster centers to be zero, obtaining the updated influence weights of the K target cluster centers, and converting the K primary global features of any optimized video frame based on the updated influence weights of the K target cluster centers to obtain the global features of any optimized video frame; until the global features of each optimized video frame in the optimized sample set are obtained; t is a positive integer less than or equal to K;
[0169] Based on the global features of each optimized video frame, recall similar video frames similar to each optimized video frame from multiple candidate video frames in the database, and add each similar video frame to the tth similar video frame group;
[0170] The t-th similar video frame group is evaluated using a preset evaluation rule to obtain an evaluation value of the t-th similar video frame group, until evaluation values of K similar video frame groups are obtained; the evaluation value of the t-th similar video frame group is used to characterize the performance of the trained global feature extraction model when the influence weight of the t-th target cluster center is zero, and the K similar video frame groups correspond to the K target cluster centers one-to-one;
[0171] Comparing the K evaluation values, determining Z evaluation values from the K evaluation values, and determining the target cluster centers corresponding to the Z evaluation values; Z is a natural number less than K;
[0172] The influence weights of the Z target cluster centers corresponding to the trained global feature extraction model are optimized to zero to obtain an optimized global feature extraction model.
[0173] In one embodiment, the processor 801 clusters the multiple local features of any optimized video frame toward K target cluster centers using the trained global feature extraction model to obtain K primary global features of any optimized video frame, and specifically performs the following operations:
[0174] Performing texture feature extraction processing on any of the optimized video frames using the trained global feature extraction model to obtain multiple texture features of the any of the optimized video frames, and using the multiple texture features of the any of the optimized video frames as local features of the any of the optimized video frames;
[0175] Based on the various local features of any optimized video frame, the probability of being assigned to each target cluster center in the K target cluster centers, and the residuals between the various local features of any optimized video frame and the various target cluster centers, the various local features of any optimized video frame are clustered to obtain K primary global features of any optimized video frame.
[0176] In one embodiment, the processor 801 clusters the local features of any optimized video frame based on the probability of each local feature of the optimized video frame being assigned to each of the K target cluster centers, and the residuals between each local feature of the optimized video frame and each target cluster center, to obtain K primary global features of the optimized video frame, and specifically performs the following operations:
[0177] Performing convolution processing on each local feature of any one of the optimized video frames to obtain a convolution result corresponding to each local feature of any one of the optimized video frames;
[0178] Predicting the probability of each local feature of any optimized video frame being assigned to each target cluster center based on the convolution result corresponding to each local feature of any optimized video frame;
[0179] For any target cluster center among the target cluster centers, based on the allocation probability of each local feature of the any optimized video frame to the any target cluster center, a weighted sum operation is performed on the residuals between each local feature of the any optimized video frame and the any target cluster center to obtain the primary global features corresponding to the any target cluster center; until the primary global features corresponding to the each target cluster center are obtained, which are used as the K primary global features of the any optimized video frame.
[0180] In one embodiment, the processor 801 converts the K primary global features of any optimized video frame based on the updated influence weights of the K target cluster centers to obtain the global features of any optimized video frame, specifically performing the following operations:
[0181] Performing weighted processing on the K primary global features of any one of the optimized video frames based on the updated influence weights of the K target cluster centers to obtain K weighted global features of any one of the optimized video frames;
[0182] Performing dimensionality reduction processing on the K weighted global features of any one of the optimized video frames to obtain a dimensionality reduction processing result corresponding to the any one of the optimized video frames;
[0183] Normalizing the dimensionality reduction processing result corresponding to any one of the optimized video frames to obtain the global features of any one of the optimized video frames.
[0184] In one embodiment, when the processor 801 recalls similar video frames similar to each optimized video frame from multiple candidate video frames in the database based on the global features of each optimized video frame, the processor 801 specifically performs the following operations:
[0185] When the influence weight of the t-th target cluster center among the K target cluster centers is zero, extracting global features of each candidate video frame in the database by using the trained global feature extraction model;
[0186] For any optimized video frame among the optimized video frames, a candidate video frame whose global features have the greatest similarity with the global features of any optimized video frame is determined as a similar video frame similar to the any optimized video frame; until a similar video frame similar to each optimized video frame is obtained.
[0187] In one embodiment, the evaluation values include but are not limited to one or more of the following: accuracy, recall rate, F1 value; the Z evaluation values are greater than or equal to the remaining evaluation values of the K evaluation values except the Z evaluation values.
[0188] In one embodiment, the trained global feature extraction model is obtained based on training the global feature extraction model; the processor 801 is further configured to:
[0189] Acquire a training sample set; the training sample set includes a plurality of sample video frames, and the sample video frames included in the training sample set are different from the optimized video frames included in the optimized sample set;
[0190] Initializing model parameters of the global feature extraction model; the model parameters of the global feature extraction model include K reference cluster centers and the influence weight of each reference cluster center in the K reference cluster centers, and the influence weight of each reference cluster center is the same as the influence weight of each target cluster center;
[0191] The global features of each sample video frame in the training sample set are extracted by the global feature extraction model, and the global feature extraction model is optimized and trained based on the global features of each sample video frame to optimize and adjust the model parameters of the global feature extraction model to obtain the trained global feature extraction model; the global features of any sample video frame in the training sample set are obtained based on the influence weights of the respective reference cluster centers, the respective local features of any sample video frame, the probability of being assigned to each reference cluster center in the K reference cluster centers, and the residuals between the respective local features of any sample video frame and the respective reference cluster centers.
[0192] In one embodiment, the training sample set includes one or more training sample subsets, and any training sample subset in the training sample set includes three sample video frames, and the three sample video frames included in any training sample subset are respectively a reference sample video frame, a similar sample video frame, and a difference sample video frame;
[0193] The processor 801 extracts global features of each sample video frame in the training sample set through the global feature extraction model, and optimizes and trains the global feature extraction model based on the global features of each sample video frame to optimize and adjust model parameters of the global feature extraction model. When the trained global feature extraction model is obtained, the processor 801 specifically performs the following operations:
[0194] Extracting global features of reference sample video frames, similar sample video frames, and difference sample video frames in the training sample subset in the training sample set respectively through the global feature extraction model;
[0195] Based on the global features of the reference sample video frame, the global features of the similar sample video frame and the global features of the difference sample video frame, the global feature extraction model is optimized and trained to optimize and adjust the model parameters of the global feature extraction model to obtain the trained global feature extraction model.
[0196] In one embodiment, the processor 801 optimizes and trains the global feature extraction model based on the global features of the reference sample video frame, the global features of the similar sample video frame, and the global features of the difference sample video frame to optimize and adjust the model parameters of the global feature extraction model. When the trained global feature extraction model is obtained, the processor 801 specifically performs the following operations:
[0197] Determining a global feature distance between the reference sample video frame and the similar sample video frame, and a global feature distance between the reference sample video frame and the difference sample video frame based on the global features of the reference sample video frame, the global features of the similar sample video frame, and the global features of the difference sample video frame;
[0198] Determining a model loss value of the global feature extraction model based on a global feature distance between the reference sample video frame and the similar sample video frame, and a global feature distance between the reference sample video frame and the difference sample video frame;
[0199] In the direction of reducing the model loss value of the global feature extraction model, the model parameters of the global feature extraction model are optimized and adjusted to obtain the trained global feature extraction model.
[0200] In one embodiment, the influence weight of each reference cluster center is 1; when the processor 801 extracts the global features of the reference sample video frames in the training sample subset in the training sample set using the global feature extraction model, the processor 801 specifically performs the following operations:
[0201] Performing texture feature extraction processing on the reference sample video frame using the global feature extraction model to obtain multiple texture features of the reference sample video frame, and using the multiple texture features of the reference sample video frame as local features of the reference sample video frame;
[0202] Based on the allocation probabilities of the local features of the reference sample video frame to the reference cluster centers, and the residuals between the local features of the reference sample video frame and the reference cluster centers, the local features of the reference sample video frame are clustered to obtain K primary global features of the reference sample video frame;
[0203] The K primary global features of the reference sample video frame are converted to obtain the global features of the reference sample video frame.
[0204] In one embodiment, the global feature extraction model includes a global feature generation module, which includes multiple dimensionality reduction submodules and a normalization submodule; the multiple dimensionality reduction submodules are connected in sequence, and among the multiple dimensionality reduction submodules, except for the penultimate dimensionality reduction submodule, the input of the remaining two correspondingly connected dimensionality reduction submodules is the output of the previous dimensionality reduction submodule; the input of the penultimate dimensionality reduction submodule is the weighted processing result between the output of the penultimate dimensionality reduction submodule and the output of the first dimensionality reduction submodule;
[0205] The processor 801 converts the K primary global features of the reference sample video frame to obtain the global features of the reference sample video frame, and specifically performs the following operations:
[0206] Calling the multiple dimensionality reduction submodules to perform dimensionality reduction processing on the K primary global features of the reference sample video frame to obtain a dimensionality reduction processing result corresponding to the reference sample video frame;
[0207] The normalization submodule is called to perform normalization processing on the dimensionality reduction processing result corresponding to the reference sample video frame to obtain the global features of the reference sample video frame.
[0208] In one embodiment, when the processor 801 obtains the training sample set, it specifically performs the following operations:
[0209] Obtaining one or more reference videos, and segmenting each of the one or more reference videos based on a video scene contained in each reference video to obtain one or more reference video segments corresponding to each video scene;
[0210] For any reference video segment corresponding to any video scene in each of the video scenes, selecting two reference video frames from each reference video frame of the reference video segment to construct a reference video frame pair, and adding the reference video frame pair to a set of video frame pairs corresponding to the any video scene; until a set of video frame pairs corresponding to each of the video scenes is obtained;
[0211] The one or more training sample subsets are constructed based on the video frame pair sets corresponding to the respective video scenes; the reference sample video frames and the similar sample video frames in any training sample subset of the one or more training sample subsets are respectively two reference video frames in a reference video frame pair, and the difference sample video frames in any training sample subset are selected from other video frame pair sets that are different from the video frame pair set to which the reference sample video frames in any training sample subset belong;
[0212] A training sample set including the one or more training sample subsets is constructed.
[0213] In one embodiment, when the processor 801 selects two reference video frames from the reference video frames of any reference video segment to construct a reference video frame pair, the processor 801 specifically performs the following operations:
[0214] Performing scale-invariant feature transformation on each reference video frame of any reference video segment to obtain reference features of each reference video frame;
[0215] Constructing an initial reference video frame pair based on any two reference video frames whose corresponding reference features have a similarity greater than a preset similarity threshold in the reference video frames;
[0216] determining a distance between two reference video frames in each initial reference video frame pair in any one of the reference video segments;
[0217] From the various initial reference video frame pairs, the initial reference video frame pair corresponding to the maximum interval distance is used as the reference video frame pair.
[0218] In one embodiment, when the processor 801 initializes the model parameters of the global feature extraction model, it specifically performs the following operations:
[0219] selecting at least one sample video frame from the plurality of sample video frames as a target sample video frame, performing texture feature extraction processing on each target sample video frame in the at least one target sample video frame to obtain a plurality of texture features of each target sample video frame, and using the plurality of texture features of each target sample video frame as a local feature of each target sample video frame;
[0220] Based on the feature distances between each local feature of each target sample video frame and each initial cluster center in the K initial cluster centers, the initial cluster centers are iteratively updated to obtain K iterative cluster centers;
[0221] The K iterative cluster centers are initialized as K reference cluster centers in the model parameters of the global feature extraction model.
[0222] The embodiment of the present application provides a computer program product, which includes a computer program stored in a computer storage medium; a processor of an electronic device reads the computer program from the computer storage medium, and the processor executes the computer program, so that the electronic device performs the above-mentioned Figure 2 as well as Figure 6 The computer-readable storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0223] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A model optimization method, characterized in that: include: For any optimized video frame in the optimized sample set, multiple local features of the optimized video frame are clustered towards K target cluster centers using the trained global feature extraction model to obtain K primary global features of the optimized video frame; the influence weight of each target cluster center in the K target cluster centers is the same; K is a positive integer; Traversing the K target cluster centers, determining the influence weight of the t-th target cluster center currently traversed among the K target cluster centers to be zero, obtaining the updated influence weights of the K target cluster centers, and converting the K primary global features of any optimized video frame based on the updated influence weights of the K target cluster centers to obtain the global features of any optimized video frame; until the global features of each optimized video frame in the optimized sample set are obtained; t is a positive integer less than or equal to K; Based on the global features of each optimized video frame, recall similar video frames similar to each optimized video frame from multiple candidate video frames in the database, and add each similar video frame to the tth similar video frame group; The t-th similar video frame group is evaluated using a preset evaluation rule to obtain an evaluation value of the t-th similar video frame group, until evaluation values of K similar video frame groups are obtained; the evaluation value of the t-th similar video frame group is used to characterize the performance of the trained global feature extraction model when the influence weight of the t-th target cluster center is zero, and the K similar video frame groups correspond to the K target cluster centers one-to-one; Comparing the K evaluation values, determining Z evaluation values from the K evaluation values, and determining target cluster centers corresponding to the Z evaluation values; Z is a natural number less than K; The influence weights of the Z target cluster centers corresponding to the trained global feature extraction model are optimized to zero to obtain an optimized global feature extraction model.
2. The method according to claim 1, wherein The trained global feature extraction model is used to cluster the multiple local features of any optimized video frame toward K target cluster centers to obtain K primary global features of any optimized video frame, including: Performing texture feature extraction processing on any of the optimized video frames using the trained global feature extraction model to obtain multiple texture features of the any of the optimized video frames, and using the multiple texture features of the any of the optimized video frames as local features of the any of the optimized video frames; Based on the various local features of any optimized video frame, the probability of being assigned to each target cluster center in the K target cluster centers, and the residuals between the various local features of any optimized video frame and the various target cluster centers, the various local features of any optimized video frame are clustered to obtain K primary global features of any optimized video frame.
3. The method according to claim 2, wherein The method comprises: clustering the local features of any optimized video frame based on the probability of each local feature of any optimized video frame being assigned to each of the K target cluster centers, and the residual between each local feature of any optimized video frame and each target cluster center to obtain K primary global features of any optimized video frame, including: Performing convolution processing on each local feature of any one of the optimized video frames to obtain a convolution result corresponding to each local feature of any one of the optimized video frames; Predicting the probability of each local feature of any optimized video frame being assigned to each target cluster center based on the convolution result corresponding to each local feature of any optimized video frame; For any target cluster center among the target cluster centers, based on the allocation probability of each local feature of the any optimized video frame to the any target cluster center, a weighted sum operation is performed on the residuals between each local feature of the any optimized video frame and the any target cluster center to obtain the primary global features corresponding to the any target cluster center; until the primary global features corresponding to the each target cluster center are obtained, which are used as the K primary global features of the any optimized video frame.
4. The method according to claim 1, wherein The converting process of the K primary global features of any one of the optimized video frames based on the updated influence weights of the K target cluster centers to obtain the global features of any one of the optimized video frames includes: Performing weighted processing on the K primary global features of any one of the optimized video frames based on the updated influence weights of the K target cluster centers to obtain K weighted global features of any one of the optimized video frames; Performing dimensionality reduction processing on the K weighted global features of any one of the optimized video frames to obtain a dimensionality reduction processing result corresponding to the any one of the optimized video frames; Normalizing the dimensionality reduction processing result corresponding to any one of the optimized video frames to obtain the global features of any one of the optimized video frames.
5. The method according to claim 1, wherein The recalling, based on the global features of the respective optimized video frames, similar video frames to the respective optimized video frames from a plurality of candidate video frames in a database comprises: When the influence weight of the t-th target cluster center among the K target cluster centers is zero, extracting global features of each candidate video frame in the database by using the trained global feature extraction model; For any optimized video frame among the optimized video frames, a candidate video frame whose global features have the greatest similarity with the global features of any optimized video frame is determined as a similar video frame similar to the any optimized video frame; until a similar video frame similar to each optimized video frame is obtained.
6. The method according to claim 1, wherein The evaluation values include but are not limited to one or more of the following: accuracy, recall rate, F1 value; the Z evaluation values are greater than or equal to the remaining evaluation values of the K evaluation values except the Z evaluation values.
7. The method according to claim 1, wherein The trained global feature extraction model is obtained based on training the global feature extraction model; the method further includes: Acquire a training sample set; the training sample set includes a plurality of sample video frames, and the sample video frames included in the training sample set are different from the optimized video frames included in the optimized sample set; Initializing model parameters of the global feature extraction model; the model parameters of the global feature extraction model include K reference cluster centers and the influence weight of each reference cluster center in the K reference cluster centers, and the influence weight of each reference cluster center is the same as the influence weight of each target cluster center; The global features of each sample video frame in the training sample set are extracted by the global feature extraction model, and the global feature extraction model is optimized and trained based on the global features of each sample video frame to optimize and adjust the model parameters of the global feature extraction model to obtain the trained global feature extraction model; the global features of any sample video frame in the training sample set are obtained based on the influence weights of the respective reference cluster centers, the respective local features of any sample video frame, the probability of being assigned to each reference cluster center in the K reference cluster centers, and the residuals between the respective local features of any sample video frame and the respective reference cluster centers.
8. The method according to claim 7, wherein The training sample set includes one or more training sample subsets, and any training sample subset in the training sample set includes three sample video frames, and the three sample video frames included in any training sample subset are respectively a reference sample video frame, a similar sample video frame, and a difference sample video frame; The global features of each sample video frame in the training sample set are extracted by the global feature extraction model, and the global feature extraction model is optimized and trained based on the global features of each sample video frame to optimize and adjust the model parameters of the global feature extraction model to obtain the trained global feature extraction model, including: Extracting global features of reference sample video frames, similar sample video frames, and difference sample video frames in the training sample subset in the training sample set respectively through the global feature extraction model; Based on the global features of the reference sample video frame, the global features of the similar sample video frame and the global features of the difference sample video frame, the global feature extraction model is optimized and trained to optimize and adjust the model parameters of the global feature extraction model to obtain the trained global feature extraction model.
9. The method according to claim 8, wherein The optimizing training of the global feature extraction model based on the global features of the reference sample video frame, the global features of the similar sample video frame, and the global features of the difference sample video frame to optimize and adjust the model parameters of the global feature extraction model to obtain the trained global feature extraction model includes: Determining a global feature distance between the reference sample video frame and the similar sample video frame, and a global feature distance between the reference sample video frame and the difference sample video frame based on the global features of the reference sample video frame, the global features of the similar sample video frame, and the global features of the difference sample video frame; Determining a model loss value of the global feature extraction model based on a global feature distance between the reference sample video frame and the similar sample video frame, and a global feature distance between the reference sample video frame and the difference sample video frame; In the direction of reducing the model loss value of the global feature extraction model, the model parameters of the global feature extraction model are optimized and adjusted to obtain the trained global feature extraction model.
10. The method according to claim 8, wherein The influence weight of each reference cluster center is 1; and the method of extracting the global features of the reference sample video frames in the training sample subset in the training sample set by using the global feature extraction model includes: Performing texture feature extraction processing on the reference sample video frame using the global feature extraction model to obtain multiple texture features of the reference sample video frame, and using the multiple texture features of the reference sample video frame as local features of the reference sample video frame; Based on the allocation probabilities of the local features of the reference sample video frame to the reference cluster centers, and the residuals between the local features of the reference sample video frame and the reference cluster centers, the local features of the reference sample video frame are clustered to obtain K primary global features of the reference sample video frame; The K primary global features of the reference sample video frame are converted to obtain the global features of the reference sample video frame.
11. The method according to claim 10, wherein The global feature extraction model includes a global feature generation module, which includes multiple dimensionality reduction submodules and a normalization submodule; the multiple dimensionality reduction submodules are connected in sequence, and among the multiple dimensionality reduction submodules, except for the penultimate dimensionality reduction submodule, the input of the remaining two corresponding dimensionality reduction submodules is the output of the previous dimensionality reduction submodule; The input of the penultimate dimensionality reduction submodule is a weighted processing result between the output of the penultimate dimensionality reduction submodule and the output of the first dimensionality reduction submodule; The converting process of the K primary global features of the reference sample video frame to obtain the global features of the reference sample video frame includes: Calling the multiple dimensionality reduction submodules to perform dimensionality reduction processing on the K primary global features of the reference sample video frame to obtain a dimensionality reduction processing result corresponding to the reference sample video frame; The normalization submodule is called to perform normalization processing on the dimensionality reduction processing result corresponding to the reference sample video frame to obtain the global features of the reference sample video frame.
12. The method according to claim 8, wherein The obtaining of the training sample set includes: Obtaining one or more reference videos, and segmenting each of the one or more reference videos based on a video scene contained in each reference video to obtain one or more reference video segments corresponding to each video scene; For any reference video segment corresponding to any video scene in each of the video scenes, selecting two reference video frames from each reference video frame of the reference video segment to construct a reference video frame pair, and adding the reference video frame pair to a set of video frame pairs corresponding to the any video scene; until a set of video frame pairs corresponding to each of the video scenes is obtained; The one or more training sample subsets are constructed based on the video frame pair sets corresponding to the respective video scenes; the reference sample video frames and the similar sample video frames in any training sample subset of the one or more training sample subsets are respectively two reference video frames in a reference video frame pair, and the difference sample video frames in any training sample subset are selected from other video frame pair sets that are different from the video frame pair set to which the reference sample video frames in any training sample subset belong; A training sample set including the one or more training sample subsets is constructed.
13. The method according to claim 12, wherein: The selecting two reference video frames from the reference video frames of any reference video segment to construct a reference video frame pair includes: Performing scale-invariant feature transformation on each reference video frame of any reference video segment to obtain reference features of each reference video frame; Constructing an initial reference video frame pair based on any two reference video frames whose corresponding reference features have a similarity greater than a preset similarity threshold in the reference video frames; determining a distance between two reference video frames in each initial reference video frame pair in any one of the reference video segments; From the various initial reference video frame pairs, the initial reference video frame pair corresponding to the maximum interval distance is used as the reference video frame pair.
14. The method according to claim 7, wherein Initializing the model parameters of the global feature extraction model includes: selecting at least one sample video frame from the plurality of sample video frames as a target sample video frame, performing texture feature extraction processing on each target sample video frame in the at least one target sample video frame to obtain a plurality of texture features of each target sample video frame, and using the plurality of texture features of each target sample video frame as a local feature of each target sample video frame; Based on the feature distances between each local feature of each target sample video frame and each initial cluster center in the K initial cluster centers, the initial cluster centers are iteratively updated to obtain K iterative cluster centers; The K iterative cluster centers are initialized as K reference cluster centers in the model parameters of the global feature extraction model.
15. A model optimization device, characterized in that: include: a processing unit configured to cluster, for any optimized video frame in the optimized sample set, a plurality of local features of the any optimized video frame toward K target clustering centers using a trained global feature extraction model to obtain K primary global features of the any optimized video frame; wherein the influence weight of each target clustering center in the K target clustering centers is the same; and K is a positive integer; The processing unit is further configured to traverse the K target cluster centers, determine the influence weight of the t-th target cluster center currently traversed among the K target cluster centers to be zero, obtain the updated influence weights of the K target cluster centers, and perform conversion processing on the K primary global features of any optimized video frame based on the updated influence weights of the K target cluster centers to obtain the global features of any optimized video frame; until the global features of each optimized video frame in the optimized sample set are obtained; t is a positive integer less than or equal to K; The processing unit is further configured to recall similar video frames that are similar to the respective optimized video frames from a plurality of candidate video frames in a database based on the global features of the respective optimized video frames, and add the respective similar video frames to a t-th similar video frame group; an evaluation unit, configured to evaluate the t-th similar video frame group using a preset evaluation rule to obtain an evaluation value for the t-th similar video frame group, until evaluation values for K similar video frame groups are obtained; the evaluation value for the t-th similar video frame group is used to characterize the performance of the trained global feature extraction model when the influence weight of the t-th target cluster center is zero, and the K similar video frame groups correspond one-to-one to the K target cluster centers; The evaluation unit is further configured to compare the K evaluation values, determine Z evaluation values from the K evaluation values, and determine target cluster centers corresponding to the Z evaluation values; Z is a natural number less than K; The optimization unit is used to optimize the influence weights of the Z target cluster centers corresponding to the trained global feature extraction model to zero, so as to obtain an optimized global feature extraction model.
Citation Information
Patent Citations
Video data processing method and device and storage medium
CN112118494A
Video classification method and device, electronic equipment and storage medium
CN112131978A