A video description method and system based on multi-level prediction architecture

A video description and short video technology, applied in the field of video description methods and systems based on multi-level prediction architecture, can solve problems such as exposure deviation, failure, gradient disappearance, etc., and achieve results with high accuracy, high practicability, and fine description. Effect

Active Publication Date: 2022-06-28
SHANDONG INSPUR SCI RES INST CO LTD
View PDF7 Cites 0 Cited by
  • Summary
  • Abstract
  • Description
  • Claims
  • Application Information

AI Technical Summary

Problems solved by technology

[0010] The technical task of the present invention is to provide a video description method and system based on a multi-level prediction architecture to solve how to generate fine-grained language descriptions, avoid gradient disappearance caused by increased model complexity, and fundamentally solve the problem of exposure deviation. Avoid the accumulation of errors and the failure of the final result

Method used

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
View more

Image

Smart Image Click on the blue labels to locate them in the text.
Viewing Examples
Smart Image
  • A video description method and system based on multi-level prediction architecture
  • A video description method and system based on multi-level prediction architecture
  • A video description method and system based on multi-level prediction architecture

Examples

Experimental program
Comparison scheme
Effect test

Embodiment 1

[0056] as attached figure 1 As shown, the video description method based on the multi-level prediction architecture of the present invention, the specific steps of the method are as follows:

[0057] S1. Obtain original data: Cut the obtained original surveillance video into short videos. The short video is to extract frames at equal short time intervals for analysis, and manually mark each short video. At the same time, the short video is divided into training set and test set;

[0058] S2. Use nltk to screen and segment the description: screen and segment the manual annotations in each short video, and sieve the annotations into words;

[0059] S3. Make a word list: make a word list according to the annotations of the training set completed by screening, and form a word list according to the order of the number of words in the annotations from high to low;

[0060] S4. Pre-training YOLO: Use the trained training set model to extract k salient regions;

[0061] S5. The lan...

Embodiment 2

[0079] as attached image 3 As shown, the video description system based on the multi-level prediction architecture of the present invention includes,

[0080] The original data acquisition module is used to cut the acquired original surveillance video into short videos. The short video is to extract frames at equal short time intervals for analysis, and manually mark each short video. At the same time, the short video is divided into training set and test set;

[0081] Filter word segmentation module, used to use nltk to filter and segment the description, filter and segment the manual annotations in each short video, and filter the annotations into words;

[0082] The word list making module is used to make a word list according to the annotation of the training set completed by screening, and form a word list according to the number of words in the annotation from high to low;

[0083] The YOLO pre-training module is used to extract k salient regions using the trained tra...

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

PUM

No PUM Login to View More

Abstract

The invention discloses a video description method and system based on a multi-level prediction framework, which belongs to the field of computer vision and natural language processing in deep learning. The technical problem to be solved by the invention is how to generate a fine-grained language description to avoid model complexity The improvement causes the gradient to disappear, and at the same time fundamentally solves the problem of exposure deviation, avoiding the accumulation of errors and causing the failure of the final result. The technical solution adopted is as follows: The steps of the method are as follows: S1. Obtain the original data; S2. Use nltk to describe Screening word segmentation; S4, pre-training YOLO; S5, obtaining the language description through the multi-layer decoder LSTM and stacking attention mechanism; S6, calculating the cross-entropy of the obtained language description and the real label, and at the same time using the sum of the obtained language description as overall loss. The system includes raw data acquisition module, screening word segmentation module, vocabulary making module, YOLO pre-training module, language description acquisition module and gradient calculation module.

Description

technical field [0001] The invention relates to the field of computer vision and natural language processing in deep learning, and can be used in various video scenarios, such as surveillance video, social video, entertainment video, etc., in particular to a video description method and system based on a multi-level prediction architecture. Background technique [0002] In recent years, as my country has entered the Internet+ era, computers and related technologies have been increasingly integrated into our lives and production, and have become important productive forces. Thanks to the rapid increase in network penetration rate, the scale of online video users in my country also ranks first in the world. As of December 2018, the scale of online video users reached 612 million, and the scale is still growing rapidly. There are a large number of media files such as videos flooding the Internet, and the quality is uneven. It has become an impossible task to completely rely on ...

Claims

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

Application Information

Patent Timeline
no application Login to View More
Patent Type & AuthorityPatents(China)
IPC IPC(8): G06V20/40G06V10/774G06V10/82G06F40/289G06N3/04
CPCG06N3/049G06V20/41G06N3/045G06F18/214
Inventor尹晓雅李锐于治楼
OwnerSHANDONG INSPUR SCI RES INST CO LTD