Album video recognition method, album video recognition model training method and device

By introducing feature extraction layer, feature mining layer and self-attention mechanism layer into the album video recognition model, combining color information and dense optical flow information, the problem of poor album video recognition effect in the prior art is solved, and more accurate recognition and higher user experience are achieved.

CN114519840BActive Publication Date: 2025-05-23CTRIP TRAVEL INFORMATION TECH (SHANGHAI) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210176321.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-05-23
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

The existing technology has poor results in photo album video recognition, which makes it difficult to guarantee the quality of videos in online travel platforms and consumes a lot of labor costs.

Method used

Deep learning technology is used to design an album video recognition model, which includes a feature extraction layer, a feature mining layer and a self-attention mechanism layer. By extracting the color information and dense optical flow information of the video image frame, and weighted processing is combined with the self-attention mechanism to achieve accurate identification of album videos.

Benefits of technology

It improves the recognition accuracy of album videos, saves operation and maintenance costs, ensures the accuracy of video display and recommendations, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114519840B_ABST
    Figure CN114519840B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying album videos, comprising: extracting color information and dense optical flow information of video image frames in a video to be identified; inputting the color information and the dense optical flow information into an album video identification model, so that the feature extraction layer of the album video identification model extracts features of the video data according to the color information and the dense optical flow information, the feature mining layer of the album video identification model mines the output result of the feature extraction layer, and the self-attention mechanism layer of the album video identification model performs weighted processing on the output result of the feature mining layer; wherein the album video identification model is obtained by training multiple video data samples; and judging whether the video is an album video according to the output result of the album video identification model. The present invention can accurately identify album videos, timely discover defects in video content, and improve user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of album video recognition, and in particular to an album video recognition method, an album video recognition model training method and device, electronic equipment, and a storage medium. Background Art

[0002] In the current Internet environment, video is an important information medium. There are many types of video production, including electronic albums generated by pictures, videos shot, and videos that combine shot videos with album videos. Experiments have shown that album videos are much less attractive to users than normal videos. In online travel platforms, videos uploaded by users are mixed with normal videos. The quality of recommended videos needs to be guaranteed on online travel platforms, so video recognition is particularly important. The quality of current videos mainly depends on manual review, and the amount of videos in the hotel industry is increasing, so maintenance requires a large amount of manpower costs. Common methods in the field of video classification include 3D convolutional networks and long-term recurrent convolutional networks based on convolutional long short-term memory. However, although 3D convolutional networks have speed advantages, they are not as accurate as long-term recurrent convolutional networks. Traditional long-term recurrent convolutional networks are mainly used in the field of video action classification, and they are not very effective in album video detection tasks. Summary of the invention

[0003] The technical problem to be solved by the present invention is to overcome the defect of poor album video recognition effect in the prior art, and to provide an album video recognition method, an album video recognition model and device, an electronic device, and a storage medium.

[0004] The present invention solves the above technical problems through the following technical solutions:

[0005] In a first aspect, a method for identifying an album video is provided, comprising:

[0006] Extracting color information and dense optical flow information of video image frames in the video to be identified;

[0007] The color information and the dense optical flow information are input into an album video recognition model, so that the feature extraction layer of the album video recognition model extracts the features of the video to be recognized according to the color information and the dense optical flow information, the feature mining layer of the album video recognition model mines the output result of the feature extraction layer, and the self-attention mechanism layer of the album video recognition model performs weighted processing on the output result of the feature mining layer; wherein the album video recognition model is trained by multiple video data samples;

[0008] According to the output result of the album video recognition model, it is determined whether the video to be recognized is an album video.

[0009] Optionally, the output result of the album video recognition model is represented by the weight of the video image frame;

[0010] Determining whether the video to be identified is an album video includes:

[0011] If the similarities of the weights output by the album video recognition model fall within a preset range, the judgment result is that the video to be recognized is an album video;

[0012] If the similarities of the weights do not fall within the preset range, the judgment result is that the video to be identified is a non-album video.

[0013] Optionally, a first fully connected layer is provided between the feature extraction layer and the feature mining layer, and the first fully connected layer is used to reduce the dimension of the output result of the feature extraction layer and then input it into the feature mining layer.

[0014] Optionally, a second fully connected layer is further provided at the output end of the self-attention mechanism layer, and the second fully connected layer is used to integrate and output the output results of the self-attention mechanism layer.

[0015] In a second aspect, a training method for an album video recognition model is provided, comprising:

[0016] Acquire multiple video data samples, each of which is annotated with annotation information, wherein the annotation information indicates whether the video data sample is an album video;

[0017] For each video data sample, extracting color information and dense optical flow information of a video image frame in the video data sample;

[0018] Inputting the color information and the dense optical flow information into a feature extraction layer, a feature mining layer and a self-attention mechanism layer, so that the feature extraction layer extracts features of the video data sample according to the color information and the dense optical flow information, the feature mining layer mines an output result of the feature extraction layer, and the self-attention mechanism layer performs weighted processing on the output result of the feature mining layer;

[0019] The loss error is calculated according to the output result of the feature mining layer and the annotation information, and the network parameters of the feature extraction layer, the feature mining layer and the self-attention mechanism layer are adjusted according to the loss error until the iteration stopping condition is reached.

[0020] Optionally, the album video recognition model also includes a first fully connected layer arranged between the feature extraction layer and the feature mining layer, and the first fully connected layer is used to reduce the dimension of the output result of the feature extraction layer and then input it into the feature mining layer.

[0021] Optionally, the album video recognition model also includes a second fully connected layer arranged at the output end of the self-attention mechanism layer, and the second fully connected layer is used to integrate and output the output results of the self-attention mechanism layer.

[0022] Optionally, the self-attention mechanism layer is characterized by the following formula:

[0023] α=softmax(w s2 tanh(W s1 H T )

[0024] Wherein, α represents the output result of the self-attention mechanism layer; w s2 , W s1 They represent custom learnable parameter vectors, w s2 The size of W is d×1, s1 The size of is d×u, d and u both represent hyperparameters; H represents the output result of the feature mining layer.

[0025] In a third aspect, a device for identifying album videos is provided, comprising:

[0026] A preprocessing module, used to extract color information and dense optical flow information of a video image frame in a video to be identified;

[0027] An input module, used for inputting the color information and the dense optical flow information into an album video recognition model, so that the feature extraction layer of the album video recognition model extracts the features of the video to be recognized according to the color information and the dense optical flow information, the feature mining layer of the album video recognition model mines the output result of the feature extraction layer, and the self-attention mechanism layer of the album video recognition model performs weighted processing on the output result of the feature mining layer; wherein the album video recognition model is trained by multiple video data samples;

[0028] The judgment module is used to judge whether the video to be identified is an album video according to the output result of the album video recognition model.

[0029] Optionally, the output result of the album video recognition model is represented by the weight of the video image frame to be recognized;

[0030] The judgment module is specifically used for:

[0031] If the similarities of the weights output by the album video recognition model fall within a preset range, the judgment result is that the video to be recognized is an album video;

[0032] If the similarities of the various weights fall within a preset range, the judgment result is that the video to be identified is a non-album video.

[0033] Optionally, a first fully connected layer is provided between the feature extraction layer and the feature mining layer, and the first fully connected layer is used to reduce the dimension of the output result of the feature extraction layer and then input it into the feature mining layer.

[0034] Optionally, a second fully connected layer is further provided at the output end of the self-attention mechanism layer, and the second fully connected layer is used to integrate and output the output results of the self-attention mechanism layer.

[0035] In a fourth aspect, a training device for an album video recognition model is provided, wherein the album video recognition model includes a feature extraction layer, a feature mining layer, and a self-attention mechanism layer cascaded in sequence; the training device includes:

[0036] An acquisition module, used to acquire multiple video data samples, each of which is annotated with annotation information, wherein the annotation information indicates whether the video data sample is an album video;

[0037] A preprocessing module, for extracting color information and dense optical flow information of a video image frame in each video data sample;

[0038] An input module, used for inputting the color information and dense optical flow information into a feature extraction layer, a feature mining layer and a self-attention mechanism layer, so that the feature extraction layer extracts features of the video data sample according to the color information and the dense optical flow information, the feature mining layer mines an output result of the feature extraction layer, and the self-attention mechanism layer performs weighted processing on the output result of the feature mining layer;

[0039] A calculation module, used for calculating the loss error according to the output result of the feature mining layer and the annotation information;

[0040] An adjustment module adjusts the network parameters of the feature extraction layer, the feature mining layer and the self-attention mechanism layer according to the loss error until an iteration stop condition is reached.

[0041] Optionally, the album video recognition model also includes a first fully connected layer arranged between the feature extraction layer and the feature mining layer, and the first fully connected layer is used to reduce the dimension of the output result of the feature extraction layer and then input it into the feature mining layer.

[0042] Optionally, the album video recognition model also includes a second fully connected layer arranged at the output end of the self-attention mechanism layer, and the second fully connected layer is used to integrate and output the output results of the self-attention mechanism layer.

[0043] Optionally, the self-attention mechanism layer is characterized by the following formula:

[0044] α=softmax(w s2 tanh(W s1 H T )

[0045] Wherein, α represents the output result of the self-attention mechanism layer; w s2 , W s1 They represent custom learnable parameter vectors, w s2 The size of W is d×1, s1 The size of is d×u, d and u both represent hyperparameters; H represents the output result of the feature mining layer.

[0046] In a fifth aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-described methods when executing the computer program.

[0047] In a sixth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, any of the methods described above is implemented.

[0048] The positive progressive effect of the present invention is that based on the massive videos in the online travel platform scenario, the album video recognition model combines the feature extraction layer, the feature mining layer and the self-attention mechanism layer by using deep learning, which can more accurately identify the album videos, timely discover the defects of the video content, control the quality of the uploaded videos, save the operation and maintenance costs, ensure the accuracy of the front-end video display videos and recommended videos, and effectively improve the user experience in the online travel platform scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 A flowchart of a method for identifying album videos provided by an exemplary embodiment of the present invention;

[0050] Figure 2 A flowchart of a method for training an album video recognition model provided by an exemplary embodiment of the present invention;

[0051] Figure 3 An architecture diagram of an album video recognition model provided by an exemplary embodiment of the present invention;

[0052] Figure 4 Another architecture diagram of an album video recognition model provided by an exemplary embodiment of the present invention;

[0053] Figure 5 A schematic diagram of a module of an album video recognition device provided by an exemplary embodiment of the present invention;

[0054] Figure 6 A module schematic diagram of a device for training an album video recognition model provided by an exemplary embodiment of the present invention;

[0055] Figure 7 A schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present invention; DETAILED DESCRIPTION

[0056] The present invention is further described below by way of examples, but the present invention is not limited to the scope of the examples.

[0057] Figure 1 A flowchart of a method for identifying an album video provided by an exemplary embodiment of the present invention, the method comprising the following steps:

[0058] Step 101: extract color information and dense optical flow information of a video image frame in a video to be identified.

[0059] First, the color information of the video image frame is converted into video RGB features. The video RGB features can be represented by, but are not limited to, a matrix.

[0060] Dense optical flow information is usually extracted using the farneback method of OpenCV (OpenCV is an open source computer vision and machine learning software library) (the farneback method is an algorithm for extracting dense optical flow in OpenCV), and the extracted dense optical flow information is represented by a grayscale image sequence.

[0061] Step 102: Input the color information and the dense optical flow information into an album video recognition model.

[0062] The RGB feature matrix and the grayscale image sequence of dense optical flow information extracted in the previous step are input into the feature extraction layer of the album video recognition model, so that the feature extraction layer extracts the features of the video to be recognized according to the color information and the dense optical flow information. After the feature mining layer of the album video recognition model mines the output result of the feature extraction layer, the self-attention mechanism layer of the album video recognition model performs weighted processing on the output result of the feature mining layer. Among them, the album video recognition model is trained by multiple video data samples, and the specific training process of the album video recognition model is described below.

[0063] In one embodiment, after the RGB feature matrix and the dense optical flow information are input into the feature extraction layer, each outputs a corresponding 1×2048-dimensional vector, and the two 1×2048-dimensional vectors obtained after the RGB feature matrix and the dense optical flow information are input into the feature extraction layer are fused. The fusion method may include but is not limited to at least one of the following methods: direct addition, splicing, etc., to obtain a fused feature vector as the output result of the feature extraction layer. It should be noted that 1×2048 dimensions are only an example, and the dimensions of the grayscale image sequence of the RGB feature matrix and the dense optical flow information can be set according to requirements in actual applications.

[0064] In one embodiment, a first fully connected layer is provided between the feature extraction layer and the feature mining layer, and the first fully connected layer is used to reduce the dimension of the output result of the feature extraction layer and then input it into the feature mining layer. The input and output data sizes of the first fully connected layer can be specified, and the main purpose of dimensionality reduction is to retain useful information, remove redundant data, reduce the workload of subsequent neural networks, and improve the speed of album video recognition.

[0065] In one embodiment, the self-attention mechanism layer performs weighted processing on the output results of the feature mining layer. Figure 3 The following is a schematic diagram of an album video recognition model provided by an exemplary embodiment of the present invention. Figure 4 Another architecture diagram of an album video recognition model provided by an exemplary embodiment of the present invention, wherein the feature mining layer is not limited to being implemented by the LSTM network shown in the figure, but can also be implemented by other neural networks. The example parameters of 0.5, 0.3, 0.1, and 0.1 marked in the figure represent the output results of the self-attention mechanism layer, which are calculated by the following formula:

[0066] α=softmax(w s2 tanh(W s1 H T )

[0067] Wherein, α represents the output result of the self-attention mechanism layer; w s2 , W s1 They represent custom learnable parameter vectors, w s2 The size of W is d×1, s1 The size of is d×u, d and u both represent hyperparameters; H represents the output result of the feature mining layer.

[0068] In one embodiment, a second fully connected layer is provided at the output end of the self-attention mechanism layer, and the second fully connected layer is used to integrate the high-dimensional results of the self-attention mechanism layer and output a 2×1 dimensional matrix. It should be noted that 2×1 dimension is only an example, and the dimension of the integrated matrix can be set according to requirements in actual applications.

[0069] Step 103: Determine whether the video to be identified is an album video based on the output result of the album video recognition model.

[0070] In one embodiment, the output of the video recognition model is represented by the weights of the video image frames:

[0071] If the phase velocity of the weight of each video image frame falls within the preset range, the judgment result is that the video to be identified is an album video, and if the similarity of the weight of each video image frame does not fall within the preset range, the judgment result is that the video to be identified is a non-album video. The similarity of the weight of each video image frame is calculated by comparing each frame and expressed as a percentage.

[0072] In one embodiment, if the similarity of the weight of each video image frame is less than 20%, the video is judged as a non-album video; if the similarity of the weight of each video image frame is higher than 80%, the video is judged as an album video; if the similarity of the weight of each video image frame is between 20% and 80%, it is necessary to manually judge whether the video is an album video. It should be noted that 20%, 80%, and 20% to 80% are just examples, and the weight similarity judgment standard can be set according to needs in actual applications.

[0073] Based on the massive videos in the online travel platform scenario, the album video recognition model uses deep learning and combines the feature extraction layer, feature mining layer and self-attention mechanism layer to more accurately identify album videos, timely discover defects in video content, control the quality of uploaded videos, save operation and maintenance costs, ensure the accuracy of front-end displayed videos and recommended videos, and effectively improve the user experience in the online travel platform scenario.

[0074] The similarity of the weight of each frame of album videos is very high, while the similarity of the weight of each frame of normally shot videos is very low. By utilizing this difference between album videos and normally shot videos and applying the self-attention mechanism layer to album video recognition, album videos can be identified more accurately and specifically, thereby discovering defects in video content in a timely manner and controlling the quality of uploaded videos.

[0075] Figure 2 A flowchart of a method for training an album video recognition model provided by an exemplary embodiment of the present invention, Figure 3 An architecture diagram of an album video recognition model provided for an exemplary embodiment of the present invention, wherein features 1, 2, ...T+1 represent color information of video image frames, and the extracted dense optical flow information is dense optical flow information between frames.

[0076] The album video recognition model includes a feature extraction layer, a feature mining layer and a self-attention mechanism layer cascaded in sequence, and the training method includes the following steps:

[0077] Step 201: Acquire multiple video data samples, each of which is annotated with annotation information, and the annotation information indicates whether the video data sample is an album video.

[0078] Step 202: extract color information and dense optical flow information of the video image frame of the video data sample.

[0079] The color information of the extracted video image frame is a matrix. The dense optical flow information is usually extracted using the farneback method of OpenCV, and the extracted dense optical flow information is represented by a grayscale image sequence.

[0080] When performing model training, the color information and dense optical flow information of the video image frames of the extracted video data samples are used as inputs for model training.

[0081] Step 203: input the color information and dense optical flow information into the feature extraction layer, feature mining layer and self-attention mechanism layer.

[0082] The color information feature matrix of the video image frame and the grayscale image sequence of the dense optical flow information extracted in the previous step are input into the feature extraction layer of the album video recognition model, so that the feature extraction layer extracts the features of the video data samples according to the color information and the dense optical flow information, the feature mining layer mines the output results of the feature extraction layer, and the self-attention mechanism layer performs weighted processing on the output results of the feature mining layer.

[0083] After the RGB feature matrix and the dense optical flow information are input into the feature extraction layer, each outputs a corresponding 1×2048-dimensional vector. The two 1×2048-dimensional vectors obtained after the RGB feature matrix and the dense optical flow information are input into the feature extraction layer are fused. The fusion method may include but is not limited to at least one of the following methods: direct addition, splicing, etc., to obtain a fused feature vector as the output result of the feature extraction layer. It should be noted that 1×2048 dimensions are only an example. In actual applications, the dimensions of the RGB feature matrix and the grayscale image sequence of the dense optical flow information can be set according to requirements.

[0084] In one embodiment, a first fully connected layer is provided between the feature extraction layer and the feature mining layer, and the first fully connected layer is used to reduce the dimension of the output result of the feature extraction layer and then input it into the feature mining layer. The input and output data sizes of the first fully connected layer can be specified, and the main purpose of dimensionality reduction is to retain useful information, remove redundant data, reduce the workload of subsequent neural networks, and improve the speed of album video recognition.

[0085] The feature extraction layer can use, but is not limited to, a convolutional neural network (CNN); the feature mining layer can use, but is not limited to, a recurrent neural network (RNN); and the self-attention mechanism layer can use, but is not limited to, a Self-Attention network.

[0086] Step 204: Calculate the loss error based on the output results of the feature mining layer and the annotation information.

[0087] The loss function for calculating the loss error may refer to the description of related technologies, and the embodiments of the present invention do not specifically limit this.

[0088] Step 205: adjust the network parameters of the feature extraction layer, feature mining layer, and self-attention mechanism layer according to the loss error until the iteration stop condition is reached.

[0089] The iteration stopping condition may include but is not limited to: the loss error of the album video recognition model converges; or the number of iterations reaches a number threshold. The number threshold can be set according to actual conditions.

[0090] After the training is completed, the album video recognition model can be obtained. The album video recognition model is used to realize the effective recognition of album videos. During the training process, the loss error is calculated according to the output results of the feature mining layer and the annotation information. The loss error is used to constrain the training of the album video recognition model to make the judgment results of the album video recognition model more accurate.

[0091] In one embodiment, the album recognition model can be trained using videos of different scenes to adapt to different usage scenarios. The videos of different scenes include at least one of the following scenes: hotel, travel photography. The album video recognition model generated for different scenes does not need to rebuild the recognition environment every time recognition is performed, thereby improving the recognition efficiency of album videos.

[0092] Corresponding to the aforementioned embodiments of the method for identifying album videos and the method for training an album video recognition model, the present invention also provides embodiments of an apparatus for identifying album videos and an apparatus for training an album video recognition model.

[0093] Figure 5 A schematic diagram of a module of a device for identifying album videos provided by an exemplary embodiment of the present invention, the device comprising:

[0094] A preprocessing module 51 is used to extract color information and dense optical flow information of a video image frame in a video to be identified;

[0095] An input module 52 is used to input the color information and the dense optical flow information into the album video recognition model, so that the feature extraction layer of the album video recognition model extracts the features of the video to be recognized according to the color information and the dense optical flow information, the feature mining layer of the album video recognition model mines the output result of the feature extraction layer, and the self-attention mechanism layer of the album video recognition model performs weighted processing on the output result of the feature mining layer; wherein the album video recognition model is trained by multiple video data samples;

[0096] The judgment module 53 is used to judge whether the video to be identified is an album video according to the output result of the album video recognition model.

[0097] Optionally, the output result of the album video recognition model is represented by the weight of the video image frame;

[0098] The judgment module is specifically used for:

[0099] If the similarities of the weights output by the album video recognition model fall within a preset range, the judgment result is that the video to be recognized is an album video;

[0100] If the similarities of the various weights fall within a preset range, the judgment result is that the video to be identified is a non-album video.

[0101] Optionally, a first fully connected layer is provided between the feature extraction layer and the feature mining layer, and the first fully connected layer is used to reduce the dimension of the output result of the feature extraction layer and then input it into the feature mining layer.

[0102] Optionally, a second fully connected layer is further provided at the output end of the self-attention mechanism layer, and the second fully connected layer is used to integrate and output the output results of the self-attention mechanism layer.

[0103] Figure 6 A schematic diagram of a module of a device for training an album video recognition model provided by an exemplary embodiment of the present invention, the device comprising:

[0104] An acquisition module 61 is used to acquire multiple video data samples, each of which is annotated with annotation information, and the annotation information indicates whether the video data sample is an album video;

[0105] A preprocessing module 62, configured to extract color information and dense optical flow information of a video image frame in each video data sample;

[0106] An input module 63, configured to input the color information and the dense optical flow information into a feature extraction layer, a feature mining layer and a self-attention mechanism layer, so that the feature extraction layer extracts features of the video data sample according to the color information and the dense optical flow information, the feature mining layer mines an output result of the feature extraction layer, and the self-attention mechanism layer performs weighted processing on the output result of the feature mining layer;

[0107] A calculation module 64, configured to calculate a loss error based on an output result of the feature mining layer and the annotation information;

[0108] The adjustment module 65 adjusts the network parameters of the feature extraction layer, the feature mining layer and the self-attention mechanism layer according to the loss error until the iteration stop condition is reached.

[0109] Optionally, the album video recognition model also includes a first fully connected layer arranged between the feature extraction layer and the feature mining layer, and the first fully connected layer is used to reduce the dimension of the output result of the feature extraction layer and then input it into the feature mining layer.

[0110] Optionally, the album video recognition model also includes a second fully connected layer arranged at the output end of the self-attention mechanism layer, and the second fully connected layer is used to integrate and output the output results of the self-attention mechanism layer.

[0111] Optionally, the self-attention mechanism layer is characterized by the following formula:

[0112] α=softmax(w s2 tanh(W s1 H T )

[0113] Wherein, α represents the output result of the self-attention mechanism layer; w s2 , W s1 They represent custom learnable parameter vectors, w s2 The size is d×1, W s1 The size of is d×u, d and u both represent hyperparameters; H represents the output result of the feature mining layer.

[0114] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.

[0115] Figure 7 The schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present invention shows a block diagram of an exemplary electronic device 70 suitable for implementing the embodiment of the present invention. Figure 7 The electronic device 70 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0116] like Figure 7 As shown, the electronic device 70 may be in the form of a general-purpose computing device, for example, it may be a server device. The components of the electronic device 70 may include, but are not limited to: at least one processor 71, at least one memory 72, and a bus 73 connecting different system components (including the memory 72 and the processor 71).

[0117] The bus 73 includes a data bus, an address bus, and a control bus.

[0118] The memory 72 may include a volatile memory, such as a random access memory (RAM) 721 and / or a cache memory 722 , and may further include a read-only memory (ROM) 723 .

[0119] The memory 72 may also include a program tool 725 (or utility) having a set (at least one) of program modules 724, such program modules 724 including but not limited to: an operating system, one or more application programs, other program modules and program data, each of which or some combination may include an implementation of a network environment.

[0120] The processor 71 executes various functional applications and data processing by running the computer program stored in the memory 72, such as the method provided in any of the above embodiments.

[0121] The electronic device 70 may also communicate with one or more external devices 74 (e.g., keyboards, pointing devices, etc.). Such communication may be performed via an input / output (I / O) interface 75. Furthermore, the model-generated electronic device 70 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 76. As shown, the network adapter 76 communicates with other modules of the model-generated electronic device 70 via a bus 73. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the model-generated electronic device 70, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (disk array) systems, tape drives, and data backup storage systems, etc.

[0122] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to an embodiment of the present invention, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided into multiple units / modules to be embodied.

[0123] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the method provided in any of the above embodiments are implemented.

[0124] The readable storage medium may include but is not limited to: a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device or any suitable combination of the above.

[0125] In a possible implementation manner, the present invention may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of the method provided in any of the above embodiments.

[0126] The program code for executing the present invention may be written in any combination of one or more programming languages, and may be executed entirely on a user device, partially on a user device, as an independent software package, partially on a user device and partially on a remote device, or entirely on a remote device.

[0127] Although the specific embodiments of the present invention are described above, it should be understood by those skilled in the art that this is only for illustration and the protection scope of the present invention is defined by the appended claims. Those skilled in the art may make various changes or modifications to these embodiments without departing from the principles and essence of the present invention, but these changes and modifications all fall within the protection scope of the present invention.

Claims

1. A method for identifying album videos, It is characterized in that include: Extracting color information and dense optical flow information of video image frames in the video to be identified; The color information and the dense optical flow information are input into an album video recognition model, so that the feature extraction layer of the album video recognition model extracts the features of the video to be recognized according to the color information and the dense optical flow information, the feature mining layer of the album video recognition model mines the output result of the feature extraction layer, and the self-attention mechanism layer of the album video recognition model performs weighted processing on the output result of the feature mining layer; wherein the album video recognition model is trained by multiple video data samples; According to the output result of the album video recognition model, determining whether the video to be recognized is an album video; The output result of the album video recognition model is represented by the weight of the video image frame; Determining whether the video to be identified is an album video includes: If the similarities of the weights output by the album video recognition model fall within a preset range, the judgment result is that the video to be recognized is an album video; If the similarities of the weights do not fall within the preset range, the judgment result is that the video to be identified is a non-album video.

2. The method for identifying album videos as claimed in claim 1, It is characterized in that A first fully connected layer is provided between the feature extraction layer and the feature mining layer. The first fully connected layer is used to reduce the dimension of the output result of the feature extraction layer and then input it into the feature mining layer.

3. The method for identifying album videos as claimed in claim 1, It is characterized in that A second fully connected layer is also provided at the output end of the self-attention mechanism layer, and the second fully connected layer is used to integrate the output results of the self-attention mechanism layer and then output them.

4. A training method for an album video recognition model, It is characterized in that The album video recognition model includes a feature extraction layer, a feature mining layer and a self-attention mechanism layer cascaded in sequence; the training method includes: Acquire multiple video data samples, each of which is annotated with annotation information, wherein the annotation information indicates whether the video data sample is an album video; For each video data sample, extracting color information and dense optical flow information of a video image frame in the video data sample; Inputting the color information and the dense optical flow information into a feature extraction layer, a feature mining layer and a self-attention mechanism layer, so that the feature extraction layer extracts features of the video data sample according to the color information and the dense optical flow information, the feature mining layer mines an output result of the feature extraction layer, and the self-attention mechanism layer performs weighted processing on the output result of the feature mining layer; Calculating a loss error according to an output result of the feature mining layer and the annotation information, and adjusting network parameters of the feature extraction layer, the feature mining layer, and the self-attention mechanism layer according to the loss error until an iteration stop condition is reached; The output result of the album video recognition model is represented by the weight of the video image frame; Determining whether the video data sample is an album video includes: If the similarity of each weight output by the album video recognition model falls within a preset range, the judgment result is that the video data sample is an album video; If the similarities of the weights do not fall within the preset range, the judgment result is that the video data sample is a non-album video.

5. The method for training the album video recognition model as claimed in claim 4, It is characterized in that The album video recognition model also includes a first fully connected layer arranged between the feature extraction layer and the feature mining layer, and the first fully connected layer is used to reduce the dimension of the output result of the feature extraction layer and then input it into the feature mining layer.

6. The method for training the album video recognition model as claimed in claim 4, It is characterized in that The album video recognition model also includes a second fully connected layer arranged at the output end of the self-attention mechanism layer, and the second fully connected layer is used to integrate and output the output results of the self-attention mechanism layer.

7. The method for training the album video recognition model as claimed in claim 6, It is characterized in that The self-attention mechanism layer is characterized by the following formula: in, Represents the output result of the self-attention mechanism layer; , They represent custom learnable parameter vectors, The size is , , d and u both represent hyperparameters; Represents the output result of the feature mining layer.

8. A device for identifying album videos, It is characterized in that include: A preprocessing module, used to extract color information and dense optical flow information of a video image frame in a video to be identified; An input module, used for inputting the color information and the dense optical flow information into an album video recognition model, so that the feature extraction layer of the album video recognition model extracts the features of the video to be recognized according to the color information and the dense optical flow information, the feature mining layer of the album video recognition model mines the output result of the feature extraction layer, and the self-attention mechanism layer of the album video recognition model performs weighted processing on the output result of the feature mining layer; wherein the album video recognition model is trained by multiple video data samples; A judgment module, used to judge whether the video to be identified is an album video according to the output result of the album video recognition model; The output result of the album video recognition model is represented by the weight of the video image frame; the judgment module is specifically used to: if the similarity of each weight output by the album video recognition model falls within a preset range, the judgment result is that the video to be identified is an album video; if the similarity of each weight does not fall within the preset range, the judgment result is that the video to be identified is a non-album video.

9. A training device for an album video recognition model, It is characterized in that The album video recognition model includes a feature extraction layer, a feature mining layer and a self-attention mechanism layer cascaded in sequence; the training device includes: An acquisition module, used to acquire multiple video data samples, each of which is annotated with annotation information, wherein the annotation information indicates whether the video data sample is an album video; A preprocessing module, for extracting color information and dense optical flow information of a video image frame in each video data sample; An input module, used for inputting the color information and dense optical flow information into a feature extraction layer, a feature mining layer and a self-attention mechanism layer, so that the feature extraction layer extracts features of the video data sample according to the color information and the dense optical flow information, the feature mining layer mines an output result of the feature extraction layer, and the self-attention mechanism layer performs weighted processing on the output result of the feature mining layer; A calculation module, used for calculating the loss error according to the output result of the feature mining layer and the annotation information; An adjustment module, which adjusts the network parameters of the feature extraction layer, the feature mining layer and the self-attention mechanism layer according to the loss error until an iteration stop condition is reached; The output result of the album video recognition model is represented by the weight of the video image frame; the judgment module is specifically used to: if the similarity of each weight output by the album video recognition model falls within a preset range, the judgment result is that the video data sample is an album video; if the similarity of each weight does not fall within the preset range, the judgment result is that the video data sample is a non-album video.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

11. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Photo album video recognition method, system and device and storage medium

    CN112749672A