Video file description text generation method and device, and storage medium
By extracting global static, dynamic, and local features from bank surveillance videos to generate descriptive text, the problem of time-consuming and labor-intensive content searching in bank surveillance videos is solved, achieving rapid location and efficient retrieval.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INDUSTRIAL AND COMMERCIAL BANK OF CHINA
- Filing Date
- 2023-03-10
- Publication Date
- 2026-04-28
AI Technical Summary
Searching for relevant content in bank surveillance videos requires enormous human resources and time, and it is impossible to quickly locate the required video file to obtain user behavior.
By extracting global static features, global dynamic features, and local features from video files, descriptive text is generated to help staff quickly locate the action and behavior features of target objects. The descriptive text is then stored in a database for easy retrieval.
It enables rapid location of video files, improves work efficiency, and supports rapid retrieval and incident investigation of bank surveillance videos.
Smart Images

Figure CN116229328B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data technology, and in particular to a method, apparatus and storage medium for generating video file description text. Background Technology
[0002] Currently, banks have installed numerous surveillance video systems. These systems allow for the analysis of customer behavior, real-time monitoring of customer activity within the bank's premises, and timely identification of unexpected issues, thus facilitating incident investigation.
[0003] Currently, if you want to find relevant content in a bank's surveillance videos, you need to review each video individually, which requires a huge amount of human resources and time.
[0004] Therefore, there is an urgent need for a method to generate video file description text, which can help staff quickly locate the required video files, thereby enabling them to quickly obtain user behavior information through video files and improve work efficiency. Summary of the Invention
[0005] This application provides a method, apparatus, and storage medium for generating video file description text, which can help staff quickly locate the required video files, thereby enabling them to quickly obtain user behavior through video files and improve work efficiency.
[0006] Firstly, this application provides a method for generating video file description text, including:
[0007] Obtain the video file to be analyzed;
[0008] Global static features, global dynamic features, and local features are extracted from the video file to be analyzed; wherein, the global static features are used to characterize the background features of the video file to be analyzed; the global dynamic features are used to characterize the movement features of each object in the video file to be analyzed; and the local features are used to characterize the features of a preset region of the video file to be analyzed.
[0009] Based on the global static features, the global dynamic features, and the local features, the action behavior features of the target object in the video file to be analyzed are determined;
[0010] Based on the global static features, the global dynamic features, the local features, and the action and behavior features of the target object, a descriptive text is generated; wherein, the descriptive text is used to characterize the content of the video file.
[0011] In one example, determining the action behavior features of the target object in the video file to be analyzed based on the global static features, the global dynamic features, and the local features includes:
[0012] Based on the global static features, the global dynamic features, and the local features, a first probability value, a second probability value, and a third probability value are determined in the video file to be analyzed; wherein, the first probability value represents the probability distribution of objects in the video file to be analyzed; the second probability value represents the probability distribution of the actions of each object in the video file to be analyzed; and the third probability value represents the probability distribution of the behavior of each object in the video file to be analyzed.
[0013] Based on the first probability value, the second probability value, and the third probability value, the action behavior characteristics of the target object in the video file to be analyzed are determined; wherein, the target object is one of the plurality of objects; wherein, the object of the first probability value, the second probability value, and the third probability value is the same.
[0014] In one example, determining a first probability value for multiple objects in the video file to be analyzed, a second probability value for each object's action, and a third probability value for each object's behavior based on the global static features, the global dynamic features, and the local features includes:
[0015] Based on the global static features and the local features, scene features are constructed;
[0016] Each object is determined based on the local features;
[0017] Based on the scene features, each object, and the global dynamic features, a first probability value, a second probability value, and a third probability value are determined in the video file to be analyzed.
[0018] In one example, generating descriptive text based on the global static features, the global dynamic features, the local features, and the action behavior features of the target object includes:
[0019] The average value of the features is determined based on the global static features, the global dynamic features, and the local features;
[0020] Descriptive text is generated based on the average value of the features and the action and behavior characteristics of the target object.
[0021] In one example, generating descriptive text based on the average value of the features and the action behavior features of the target object includes:
[0022] The average value of the features and the action behavior features of the target object are fused to obtain the fused features;
[0023] The fused features are input into the encoder to generate word probability distributions;
[0024] Descriptive text is generated based on the probability distribution of the words.
[0025] In one example, after generating the descriptive text, the following is also included:
[0026] The description text and the generation time of the video file to be analyzed are stored in a preset database.
[0027] In one example, the method further includes:
[0028] Responding to the user's index message;
[0029] Based on the index message, query the preset database for video files;
[0030] The video file will be sent back to the user.
[0031] Secondly, this application provides a video file description text generation apparatus, the apparatus comprising:
[0032] The acquisition unit is used to acquire the video file to be analyzed.
[0033] An extraction unit is used to extract global static features, global dynamic features, and local features from the video file to be analyzed; wherein, the global static features are used to characterize the background features of the video file to be analyzed; the global dynamic features are used to characterize the movement features of each object in the video file to be analyzed; and the local features are used to characterize the features of a preset region of the video file to be analyzed.
[0034] The determining unit is configured to determine the action behavior features of the target object in the video file to be analyzed based on the global static features, the global dynamic features, and the local features.
[0035] The generation unit is configured to generate descriptive text based on the global static features, the global dynamic features, the local features, and the action behavior features of the target object; wherein the descriptive text is used to characterize the content of the video file.
[0036] In one example, the identified unit includes:
[0037] The first determining module is used to determine a first probability value, a second probability value, and a third probability value in the video file to be analyzed based on the global static features, the global dynamic features, and the local features; wherein, the first probability value represents the probability distribution of objects in the video file to be analyzed; the second probability value represents the probability distribution of the actions of each object in the video file to be analyzed; and the third probability value represents the probability distribution of the behavior of each object in the video file to be analyzed.
[0038] The second determining module is used to determine the action behavior characteristics of a target object in the video file to be analyzed based on the first probability value, the second probability value, and the third probability value; wherein the target object is one of the plurality of objects; wherein the object of the first probability value, the second probability value, and the third probability value is the same.
[0039] In one example, the first determined module includes:
[0040] A submodule is constructed to build scene features based on the global static features and the local features;
[0041] The first determining submodule is used to determine each object based on the local features;
[0042] The second determining submodule is used to determine a first probability value, a second probability value, and a third probability value in the video file to be analyzed based on the scene features, each object, and the global dynamic features.
[0043] In one example, the generating unit includes:
[0044] The third determining module is used to determine the average value of the features based on the global static features, the global dynamic features, and the local features;
[0045] The generation module is used to generate descriptive text based on the average value of the features and the action and behavior features of the target object.
[0046] In one example, the generated module includes:
[0047] The fusion submodule is used to fuse the average value of the features and the action behavior features of the target object to obtain fused features;
[0048] The first generation submodule is used to input the fused features into the encoder to generate word probability distributions;
[0049] The second generation submodule is used to generate descriptive text based on the word probability distribution.
[0050] In one example, the device includes:
[0051] The storage unit is used to store the description text and the generation time of the video file to be analyzed into a preset database.
[0052] In one example, the device includes:
[0053] A response unit is used to respond to the user's index message;
[0054] The query unit is used to query video files in the preset database based on the index message;
[0055] The feedback unit is used to send the video file back to the user.
[0056] Thirdly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0057] The memory stores computer-executed instructions;
[0058] The processor executes computer execution instructions stored in the memory to implement the method as described in the first aspect.
[0059] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method described in the first aspect.
[0060] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0061] This application provides a method, apparatus, and storage medium for generating descriptive text for video files. The method involves acquiring a video file to be analyzed and extracting global static features, global dynamic features, and local features from the video file. The global static features characterize the background features of the video file; the global dynamic features characterize the movement features of each object in the video file; and the local features characterize the features of a preset region of the video file. Based on the global static features, global dynamic features, and local features, the action behavior features of a target object in the video file are determined. Descriptive text is generated based on the global static features, global dynamic features, local features, and the action behavior features of the target object. The descriptive text characterizes the content of the video file. This technical solution can assist staff in quickly locating the required video file, thereby enabling rapid acquisition of user behavior through video files and improving work efficiency. Attached Figure Description
[0062] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0063] Figure 1 This is a flowchart illustrating a method for generating video file description text according to Embodiment 1 of this application;
[0064] Figure 2 This is a flowchart illustrating a method for generating video file description text according to Embodiment 2 of this application;
[0065] Figure 3 This is a schematic diagram of a video file description text generation device according to Embodiment 3 of this application;
[0066] Figure 4 This is a schematic diagram of a video file description text generation device according to Embodiment 4 of this application;
[0067] Figure 5 This is a block diagram illustrating an electronic device according to an exemplary embodiment.
[0068] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0069] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0070] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0071] Figure 1 This is a flowchart illustrating a method for generating video file description text according to Embodiment 1 of this application. Embodiment 1 includes the following steps:
[0072] S101. Obtain the video file to be analyzed.
[0073] In one example, the video file to be analyzed is retrieved from the bank's pre-set database. The video file is then input into the bank's monitoring system for analysis.
[0074] S102. Extract global static features, global dynamic features, and local features from the video file to be analyzed; wherein, global static features are used to characterize the background features of the video file to be analyzed; global dynamic features are used to characterize the movement features of each object in the video file to be analyzed; and local features are used to characterize the features of a preset region in the video file to be analyzed.
[0075] In this embodiment, the video file to be analyzed includes multiple objects, which may be a table or different people. Global static features are the background features of the video file to be analyzed; for example, the background features may be the scene features inside a bank. Global dynamic features are the movement features of each object in the video file to be analyzed; for example, the trajectory information of person A moving from position A to window B is the global dynamic feature. It is worth noting that the global dynamic feature is the movement feature of multiple objects, not just the movement feature of a single specified object. Local features refer to the features of a preset area, where the preset area may be a partial area; for example, a local feature may be the hand features of a specified person in the bank.
[0076] S103. Based on global static features, global dynamic features, and local features, determine the action and behavior features of the target object in the video file to be analyzed.
[0077] In this embodiment, the action behavior characteristics of the target object refer to the actions performed by the target object. Further, the action behavior characteristics of the target object include the target object, the verb, and the specific action performed. In this embodiment, the action behavior characteristics of the target object are jointly determined by global static features, global dynamic features, and local features.
[0078] S104. Generate descriptive text based on global static features, global dynamic features, local features, and the action and behavior features of the target object; wherein, the descriptive text is used to characterize the content of the video file.
[0079] In this embodiment, the description file includes keywords and text content, through which the content of the video file can be obtained. Global static features, global dynamic features, local features, and local features are input into a multi-head attention encoding system for processing, and then the action and behavior features of the target object in the video file to be analyzed are output.
[0080] This application provides a method for generating descriptive text for video files. The method involves acquiring the video file to be analyzed, extracting global static features, global dynamic features, and local features from the video file, and then determining the action and behavior features of the target object in the video file based on these features. Descriptive text is then generated based on the global static features, global dynamic features, local features, and the action and behavior features of the target object. This descriptive text characterizes the content of the video file. Using this technical solution, staff can quickly locate the required video file, thereby enabling rapid acquisition of user behavior through the video file and improving work efficiency.
[0081] Figure 2 This is a flowchart illustrating a method for generating video file description text according to Embodiment 2 of this application. Embodiment 2 includes the following steps:
[0082] S201. Obtain the video file to be analyzed.
[0083] For example, this step can refer to step S101 above, and will not be repeated here.
[0084] S202. Extract global static features, global dynamic features, and local features from the video file to be analyzed; wherein, global static features are used to characterize the background features of the video file to be analyzed; global dynamic features are used to characterize the movement features of each object in the video file to be analyzed; and local features are used to characterize the features of a preset region in the video file to be analyzed.
[0085] In this embodiment, the global static features in the video file to be analyzed are extracted using the two-dimensional convolutional neural network InceptionResnetV2. Here, the global static features can be denoted as V. r Global dynamic features in the video file to be analyzed are extracted using a 3D convolutional neural network (C3D). Here, the global dynamic features can be denoted as V. c Local features are extracted from the video file to be analyzed using the object detector Faster R-CNN. Here, the local features can be denoted as V. o .
[0086] S203. Based on global static features, global dynamic features, and local features, determine the first probability value, the second probability value, and the third probability value in the video file to be analyzed.
[0087] In this embodiment, the first probability value represents the probability value of each object. For example, the first probability value of object A is P1(S1), and the first probability value of object B is P1(S2). The second probability value represents the probability value of each action of each object. For example, if object A performs a service at a window, and object B performs a service at a self-service machine, then the second probability value of object A is P2(a1), and the second probability value of object B is P2(a2). The behavior of each object represents the object it executes, for example, it could be data. Therefore, the third probability value of object A is P3(o1), and the second probability value of object B is P3(o2).
[0088] In one example, based on global static features, global dynamic features, and local features, the first probability value, the second probability value, and the third probability value in the video file to be analyzed are determined, including:
[0089] Construct scene features based on global static features and local features;
[0090] Each object is identified based on its local features;
[0091] Based on scene features, each object, and global dynamic features, determine the first probability value, the second probability value, and the third probability value in the video file to be analyzed.
[0092] In this embodiment, scene features are used to characterize the current scene information. Specifically, they can be constructed by concatenating a global static feature matrix and a local feature matrix. In this embodiment, local features V... o And the Faster R-CNN object detector can identify multiple objects. After identifying multiple objects, the relative position information of the objects is input into the scene features to determine the position information of each object. Specifically, this can be represented by the following formula:
[0093] L i =Re lu(W L [X,Y,W,H]);
[0094] Where Li represents the object's position information, ReLU is the activation function, X represents the object's x-coordinate, Y represents the object's y-coordinate, W represents the object's width, and H represents the object's height. Further, the identified multiple objects are combined with global dynamic features to obtain a second probability value for each object's action. Specifically, this can be achieved using the following formula:
[0095] P2(a) = soft max(W) a Re lu[Emb(s n ):V c ]);
[0096] Among them, s n It is an object, V c It represents global dynamic features, Emb is the word embedding in the vocabulary, [:] represents the concatenation of two matrices, and ReLU is the activation function.
[0097] S204. Based on the first probability value, the second probability value, and the third probability value, determine the action and behavior characteristics of the target object in the video file to be analyzed; wherein, the target object is one of multiple objects.
[0098] In this embodiment, the action and behavior features of the target object in the video file to be analyzed are determined through a self-attention mechanism and according to the following formula:
[0099] P1(s n =SelfAttention(Emb(s) n V o V o );
[0100] P2(a)=SelfAttention(Emb(a),V c V c );
[0101] P3(o)=SelfAttention(Emb(o),V o V o );
[0102] F action =P(α)·Emb(β)
[0103] Where s n 'a' and 'o' represent the object, the action performed, and the object that the object executes, respectively. o and V c Representing local features and global dynamic features respectively; where α∈{s n ,a,o},β∈{s n ,a,o},[·] represents the dot product operation,F action It represents the action and behavior characteristics of the target object, and Emb is the word embedding in the vocabulary.
[0104] S205. Determine the average value of the features based on the global static features, global dynamic features, and local features.
[0105] In this embodiment, the feature average is the mean of the global static features, global dynamic features, and local features. Specifically, the feature average can be calculated from F. visual express.
[0106] S206. Generate descriptive text based on the average feature value and the action and behavior characteristics of the target object.
[0107] In one example, descriptive text is generated based on the average feature value and the target object's action behavior characteristics, including:
[0108] The average feature value and the action behavior feature of the target object are fused to obtain the fused feature;
[0109] The fused features are input into the encoder to generate word probability distributions;
[0110] Descriptive text is generated based on the probability distribution of words.
[0111] In this embodiment, the average feature value and the action behavior features of the target object are fused to obtain the fused feature, which can be achieved by the following formula:
[0112]
[0113]
[0114] Among them, F action F represents the action and behavior characteristics of the target object. visual Denotes the characteristic average, where, and The fused features are input into two fully connected layers and then into the encoder to obtain the word probability distribution, which can be specifically obtained using the following formula:
[0115]
[0116]
[0117] P(w t ) = soft max(F n );
[0118] Among them, P(w t ) represents the probability distribution of words.
[0119] In this embodiment, after determining the probability distribution of words, the word with the highest probability value is used as the descriptive text.
[0120] S207. Store the generation time of the description text and the video file to be analyzed into a preset database.
[0121] In this embodiment, the preset database includes multiple descriptive texts and the generation time of each video file. The preset database is used for subsequent user queries, enabling quick location of the required video file.
[0122] In one example, in response to a user's index message.
[0123] In this embodiment, the user's index message consists of keywords, key fields, and time range entered by the user.
[0124] In one example, a video file is queried in a preset database based on an index message.
[0125] In this embodiment, the keywords entered by the user are compared with the description files or generation times in the preset database. If they match, the relevant statements and times are filtered out and displayed in the information bar on the right. At the same time, the time stamp is displayed in the time bar of the bank monitoring system, thereby determining the video file.
[0126] In one example, a video file is sent back to the user.
[0127] In this embodiment, the video file is displayed to the user through the monitoring system's interface.
[0128] This application provides a method for generating video file description text. Based on global static features, global dynamic features, and local features, it determines the first probability value of multiple objects in the video file to be analyzed, the second probability value of each object's action, and the third probability value of each object's behavior. Based on the first, second, and third probability values, it determines the action and behavior characteristics of the target object in the video file to be analyzed. Based on the global static features, global dynamic features, and local features, it determines the average value of these features. Based on the average value of the features and the action and behavior characteristics of the target object, it generates description text. The description text and the generation time of the video file to be analyzed are stored in a preset database. Based on subsequent user queries, the video file is returned to the user. This technical solution enables the investigation of some emergencies and accidents. Through keyword search and location of relevant content, it enables rapid retrieval of bank surveillance data even with massive amounts of video data. Furthermore, for situations where video files occupy large storage space and cannot be stored long-term, the video file description text can be stored to assist in later verification.
[0129] Figure 3 This is a schematic diagram of a video file description text generation device according to Embodiment 3 of this application. Specifically, the device 30 in Embodiment 3 includes:
[0130] Acquisition unit 301 is used to acquire the video file to be analyzed.
[0131] Extraction unit 302 is used to extract global static features, global dynamic features and local features from the video file to be analyzed; wherein, global static features are used to characterize the background features of the video file to be analyzed; global dynamic features are used to characterize the movement features of each object in the video file to be analyzed; and local features are used to characterize the features of a preset region in the video file to be analyzed.
[0132] The determination unit 303 is used to determine the action behavior features of the target object in the video file to be analyzed based on global static features, global dynamic features, and local features.
[0133] The generation unit 304 is used to generate descriptive text based on global static features, global dynamic features, local features, and the action and behavior features of the target object; wherein, the descriptive text is used to characterize the content of the video file.
[0134] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the above-described device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0135] Figure 4 This is a schematic diagram of a video file description text generation device according to Embodiment 4 of this application. Specifically, the device 40 in Embodiment 4 includes:
[0136] Acquisition unit 401 is used to acquire the video file to be analyzed.
[0137] Extraction unit 402 is used to extract global static features, global dynamic features and local features from the video file to be analyzed; wherein, global static features are used to characterize the background features of the video file to be analyzed; global dynamic features are used to characterize the movement features of each object in the video file to be analyzed; and local features are used to characterize the features of a preset region in the video file to be analyzed.
[0138] The determination unit 403 is used to determine the action behavior features of the target object in the video file to be analyzed based on global static features, global dynamic features, and local features.
[0139] The generation unit 404 is used to generate descriptive text based on global static features, global dynamic features, local features and the action and behavior features of the target object; wherein, the descriptive text is used to characterize the content of the video file.
[0140] In one example, unit 403 is defined as including:
[0141] The first determining module 4031 is used to determine a first probability value, a second probability value, and a third probability value in the video file to be analyzed based on global static features, global dynamic features, and local features; wherein, the first probability value represents the probability distribution of objects in the video file to be analyzed; the second probability value represents the probability distribution of the actions of each object in the video file to be analyzed; and the third probability value represents the probability distribution of the behavior of each object in the video file to be analyzed.
[0142] The second determining module 4032 is used to determine the action behavior characteristics of a target object in the video file to be analyzed based on a first probability value, a second probability value, and a third probability value; wherein the target object is one of multiple objects; wherein the object of the first probability value, the second probability value, and the third probability value is the same object.
[0143] In one example, the first determining module 4031 includes:
[0144] Submodule 40311 is constructed to build scene features based on global static features and local features.
[0145] The first determination submodule 40312 is used to determine each object based on local features.
[0146] The second determination submodule 40313 is used to determine the first probability value, the second probability value, and the third probability value in the video file to be analyzed based on scene features, each object, and global dynamic features.
[0147] In one example, generation unit 404 includes:
[0148] The third determining module 4041 is used to determine the average value of features based on global static features, global dynamic features, and local features.
[0149] The generation module 4042 is used to generate descriptive text based on the average value of features and the action and behavior features of the target object.
[0150] In one example, module 4042 is generated, including:
[0151] The fusion submodule 40421 is used to fuse the average feature value and the action behavior feature of the target object to obtain the fused feature.
[0152] The first generation submodule 40422 is used to input the fused features into the encoder to generate word probability distributions.
[0153] The second generation submodule 40423 is used to generate descriptive text based on the word probability distribution.
[0154] In one example, device 40 includes:
[0155] Storage unit 405 is used to store the description text and the generation time of the video file to be analyzed into a preset database.
[0156] In one example, device 40 includes:
[0157] Response unit 406 is used to respond to the user's index message.
[0158] The query unit 407 is used to query video files in a preset database based on the index message.
[0159] Feedback unit 408 is used to send video files back to the user.
[0160] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the above-described device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0161] Figure 5 This is a block diagram illustrating an electronic device according to an exemplary embodiment. The device may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness device, personal digital assistant, etc.
[0162] The device 500 may include one or more of the following components: processing component 502, memory 504, power supply component 506, multimedia component 508, audio component 510, input / output (I / O) interface 512, sensor component 514, and communication component 516.
[0163] Processing component 502 typically controls the overall operation of device 500, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 502 may include one or more modules to facilitate interaction between processing component 502 and other components. For example, processing component 502 may include a multimedia module to facilitate interaction between multimedia component 508 and processing component 502.
[0164] Memory 504 is configured to store various types of data to support the operation of device 500. Examples of such data include instructions for any application or method operating on device 500, contact data, phonebook data, messages, pictures, videos, etc. Memory 504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0165] Power supply component 506 provides power to various components of device 500. Power supply component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 500.
[0166] Multimedia component 508 includes a screen that provides an output interface between device 500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 508 includes a front-facing camera and / or a rear-facing camera. When device 500 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0167] Audio component 510 is configured to output and / or input audio signals. For example, audio component 510 includes a microphone (MIC) configured to receive external audio signals when device 500 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 504 or transmitted via communication component 516. In some embodiments, audio component 510 also includes a speaker for outputting audio signals.
[0168] I / O interface 512 provides an interface between processing component 502 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0169] Sensor assembly 514 includes one or more sensors for providing state assessments of various aspects of device 500. For example, sensor assembly 514 may detect the on / off state of device 500, the relative positioning of components such as the display and keypad of device 500, changes in the position of device 500 or a component of device 500, the presence or absence of user contact with device 500, the orientation or acceleration / deceleration of device 500, and temperature changes of device 500. Sensor assembly 514 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 514 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 514 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0170] Communication component 516 is configured to facilitate wired or wireless communication between device 500 and other devices. Device 500 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 516 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 516 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0171] In an exemplary embodiment, the apparatus 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0172] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 504 including instructions, which can be executed by a processor 520 of the device 500 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0173] A non-transitory computer-readable storage medium, wherein instructions in the storage medium, when executed by a processor of an electronic device, enable the electronic device to perform a video file description text generation method of the electronic device.
[0174] This application also discloses a computer program product, including a computer program that, when executed by a processor, implements the method described in this embodiment.
[0175] Various embodiments of the systems and technologies described above in this application can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0176] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or electronic device.
[0177] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0178] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0179] The systems and technologies described herein can be implemented in computing systems that include back-end components (e.g., as data electronic devices), or computing systems that include middleware components (e.g., application electronic devices), or computing systems that include front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0180] Computer systems can include client and electronic devices. Clients and electronic devices are generally geographically separated and typically interact via communication networks. The client-electronic device relationship is created by computer programs running on the respective computers and having a client-electronic device relationship with each other. The electronic device can be a cloud electronic device, also known as a cloud computing electronic device or cloud host, a host product within the cloud computing service system, addressing the shortcomings of traditional physical hosts and VPS services ("Virtual Private Server," or simply "VPS") in terms of management difficulty and weak business scalability. The electronic device can also be an electronic device in a distributed system or an electronic device incorporating blockchain technology. It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application is achieved, and this is not limited herein.
[0181] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0182] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A method for generating descriptive text for video files, characterized in that, The method includes: Obtain the video file to be analyzed; Global static features, global dynamic features, and local features are extracted from the video file to be analyzed; wherein, the global static features are used to characterize the background features of the video file to be analyzed; the global dynamic features are used to characterize the movement features of each object in the video file to be analyzed; and the local features are used to characterize the features of a preset region of the video file to be analyzed. Based on the global static features, the global dynamic features, and the local features, the action behavior features of the target object in the video file to be analyzed are determined; Based on the global static features, the global dynamic features, the local features, and the action and behavior features of the target object, a descriptive text is generated; wherein, the descriptive text is used to characterize the content of the video file; The step of determining the action behavior features of the target object in the video file to be analyzed based on the global static features, the global dynamic features, and the local features includes: Scene features are constructed based on the global static features and the local features; each object is determined based on the local features; a first probability value, a second probability value, and a third probability value are determined in the video file to be analyzed based on the scene features, each object, and the global dynamic features; wherein, the first probability value represents the probability distribution of objects in the video file to be analyzed; the second probability value represents the probability distribution of actions of each object in the video file to be analyzed; and the third probability value represents the probability distribution of behavior of each object in the video file to be analyzed. Based on the first probability value, the second probability value, and the third probability value, the action behavior characteristics of the target object in the video file to be analyzed are determined; wherein, the target object is one of multiple objects; wherein, the object of the first probability value, the second probability value, and the third probability value is the same object.
2. The method according to claim 1, characterized in that, The step of generating descriptive text based on the global static features, the global dynamic features, the local features, and the action and behavior features of the target object includes: The average value of the features is determined based on the global static features, the global dynamic features, and the local features; Descriptive text is generated based on the average value of the features and the action and behavior characteristics of the target object.
3. The method according to claim 2, characterized in that, The step of generating descriptive text based on the average value of the features and the action behavior features of the target object includes: The average value of the features and the action behavior features of the target object are fused to obtain the fused features; The fused features are input into the encoder to generate word probability distributions; Descriptive text is generated based on the probability distribution of the words.
4. The method according to claim 1, characterized in that, After generating the description text, the following is also included: The description text and the generation time of the video file to be analyzed are stored in a preset database.
5. The method according to claim 4, characterized in that, The method further includes: Responding to the user's index message; Based on the index message, query the preset database for video files; The video file will be sent back to the user.
6. A device for generating descriptive text for video files, characterized in that, The device includes: The acquisition unit is used to acquire the video file to be analyzed. An extraction unit is used to extract global static features, global dynamic features, and local features from the video file to be analyzed; wherein, the global static features are used to characterize the background features of the video file to be analyzed; the global dynamic features are used to characterize the movement features of each object in the video file to be analyzed; and the local features are used to characterize the features of a preset region of the video file to be analyzed. The determining unit is configured to determine the action behavior features of the target object in the video file to be analyzed based on the global static features, the global dynamic features, and the local features. The generation unit is configured to generate descriptive text based on the global static features, the global dynamic features, the local features, and the action behavior features of the target object; wherein the descriptive text is used to characterize the content of the video file; The generation unit is specifically configured to: construct scene features based on the global static features and the local features; determine each object based on the local features; determine a first probability value, a second probability value, and a third probability value in the video file to be analyzed based on the scene features, each object, and the global dynamic features; wherein the first probability value represents the probability distribution of objects in the video file to be analyzed; the second probability value represents the probability distribution of actions of each object in the video file to be analyzed; the third probability value represents the probability distribution of behavior of each object in the video file to be analyzed; and determine the action behavior features of the target object in the video file to be analyzed based on the first probability value, the second probability value, and the third probability value; wherein the target object is one of multiple objects; and wherein the object for the first probability value, the second probability value, and the third probability value is the same object.
7. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-5.
9. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Video description method based on action guidance
CN115376039A