Time segment positioning in media streams
By extracting two-dimensional timing feature maps and combining vector representations of natural language queries, the time-series neighbor network is used to determine the correlation between time periods and behaviors in the media stream, solving the problem of difficult to deal with the correlation between time periods and behaviors in the prior art, and achieving efficient time-series positioning effect.
Patent Information
- Application Number
- CN201911059082.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-01
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2039-11-01
AI Technical Summary
The prior art is difficult to effectively handle the correlation between multiple periods in media streams and behavior, especially in the case of natural language queries.
By extracting the two-dimensional timing feature map, representing the start and end times of multiple periods, and combining vector representations of natural language query, the correlation between each period and behavior is determined using a timing adjacent network.
It realizes efficient evaluation of the correlation between multiple periods and media streaming behavior, and can more accurately locate the behavioral periods in the video that match natural language queries.
Smart Images

Figure CN112765377B_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to computer-implemented methods, electronic devices, and computer program products for media stream processing. Background Art
[0002] Currently, video understanding is an important research direction in the field of computer vision. For example, using natural language to locate time periods is an important research direction in video understanding. For example, a natural language query that specifies an action can be used to determine the location of the time period of an action in a video that matches the query. Summary of the invention
[0003] According to some implementations of the present disclosure, a neural network-based scheme for processing a media stream is provided. A two-dimensional temporal feature map representing multiple time periods in the media stream is extracted from the media stream. The two-dimensional temporal feature map includes a first dimension representing the start of a time period in the multiple time periods and a second dimension representing the end of a time period in the multiple time periods. Then, based on the two-dimensional temporal feature map, correlations between the multiple time periods and behaviors in the media stream are determined.
[0004] This Summary is provided to introduce a selection of concepts in a simplified form that are further described in the Detailed Description below. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Figure 1 A block diagram showing a computing device capable of implementing various implementations of the present disclosure is shown;
[0006] Figure 2 A schematic diagram showing the architecture of a two-dimensional temporal neighbor network according to some implementations of the present disclosure;
[0007] Figure 3A and Figure 3B A schematic diagram showing a two-dimensional temporal feature map according to some implementations of the present disclosure is shown;
[0008] Figure 4 A schematic diagram of a two-dimensional temporal feature map extractor according to some implementations of the present disclosure is shown;
[0009] Figure 5 A schematic diagram showing a sparse sampling method according to some implementations of the present disclosure is shown;
[0010] Figure 6 A schematic diagram of a speech encoder according to some implementations of the present disclosure is shown;
[0011] Figure 7A schematic diagram showing a timing neighbor network according to some implementations of the present disclosure is shown;
[0012] Figure 8 A schematic diagram showing scoring according to some implementations of the present disclosure; and
[0013] Fig. 9 A flow chart illustrating a method according to some implementations of the present disclosure is shown.
[0014] In these drawings, the same or similar reference symbols are used to designate the same or similar elements. DETAILED DESCRIPTION
[0015] The present disclosure will now be discussed with reference to several example implementations. It should be understood that these implementations are discussed only to enable those skilled in the art to better understand and thus implement the present disclosure, and do not imply any limitation on the scope of the present subject matter.
[0016] As used herein, the term "including" and variations thereof are to be interpreted as open-ended terms meaning "including but not limited to". The term "based on" is to be interpreted as "based at least in part on". The terms "an implementation" and "an implementation" are to be interpreted as "at least one implementation". The term "another implementation" is to be interpreted as "at least one other implementation". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0017] The basic principles and several example implementations of the present disclosure are explained below with reference to the accompanying drawings. Figure 1 1 is a block diagram of a computing device 100 capable of implementing multiple implementations of the present disclosure. It should be understood that Figure 1 The computing device 100 shown is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described in the present disclosure. Figure 1 As shown, computing device 100 comprises a computing device in the form of a general purpose computing device 100. Components of computing device 100 may include, but are not limited to, one or more processors or processing units 110, memory 120, storage device 130, one or more communication units 140, one or more input devices 150, and one or more output devices 160.
[0018] In some implementations, the computing device 100 can be implemented as various user terminals or service terminals with computing capabilities. The service terminal can be a server, a large computing device, etc. provided by various service providers. The user terminal is such as any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a site, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is also foreseeable that the computing device 100 can support any type of interface for the user (such as a "wearable" circuit, etc.).
[0019] Processing unit 110 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 120. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to increase the parallel processing capabilities of computing device 100. Processing unit 110 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0020] The computing device 100 typically includes a plurality of computer storage media. Such media may be any available media accessible to the computing device 100, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 120 may be a volatile memory (e.g., registers, caches, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The memory 120 may include a neural network (NN) module 122, which are program modules configured to perform the functions of the various implementations described herein. The neural network (NN) module 122 may be accessed and run by the processing unit 110 to implement the corresponding functions.
[0021] Storage device 130 may be a removable or non-removable medium and may include machine-readable media that can be used to store information and / or data and can be accessed within computing device 100. Computing device 100 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 1As shown in , a disk drive for reading or writing from a removable, nonvolatile disk and an optical drive for reading or writing from a removable, nonvolatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces.
[0022] The communication unit 140 enables communication with another computing device via a communication medium. Additionally, the functions of the components of the computing device 100 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Therefore, the computing device 100 can operate in a networked environment using a logical connection with one or more other servers, a personal computer (PC), or another general network node.
[0023] Input device 150 may be one or more various input devices, such as a mouse, keyboard, tracking ball, voice input device, etc. Output device 160 may be one or more output devices, such as a display, a speaker, a printer, etc. Computing device 100 may also communicate with one or more external devices (not shown) through communication unit 140 as needed, such as storage devices, display devices, etc., communicate with one or more devices that allow a user to interact with computing device 100, or communicate with any device that allows computing device 100 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0024] In some implementations, in addition to being integrated on a single device, some or all of the various components of the computing device 100 can also be set in the form of a cloud computing architecture. In a cloud computing architecture, these components can be remotely arranged and can work together to implement the functions described in the present disclosure. In some implementations, cloud computing provides computing, software, data access and storage services, which do not require end users to know the physical location or configuration of the system or hardware that provides these services. In various implementations, cloud computing uses appropriate protocols to provide services through a wide area network (such as the Internet). For example, a cloud computing provider provides applications through a wide area network, and they can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be merged at a remote data center location or they can be dispersed. Cloud computing infrastructure can provide services through a shared data center, even if they appear as a single access point for users. Therefore, the components and functions described herein can be provided from a service provider at a remote location using a cloud computing architecture. Alternatively, they can be provided from a conventional server, or they can be installed directly or otherwise on a client device.
[0025] The computing device 100 can be used to implement the scheme of artificial neural network according to multiple implementations of the present disclosure. The computing device 100 can receive input data, such as media streams such as video and audio, through the input device 150. In addition, the computing device 100 can also receive queries for specific actions in videos, etc. through the input device 150. The query can be in text form or in the form of audio, etc. The neural network (NN) module 122 can process the input data to obtain corresponding output data, such as the location of the action specified in the query in the video. The output data can be provided to the output device 160 to be provided to the user, etc. as output 180.
[0026] As mentioned above, moment positioning is an important branch application of video understanding, and its purpose is to determine the time period of the behavior matching the query in the video according to the natural language query, for example, the start and end time of the time period. Behavior does not necessarily mean strong body movements, but can also mean continuous states. For example, in a saxophone teaching video, the teacher may be playing the saxophone in some time periods, explaining how to play the saxophone in some time periods, and playing the saxophone again after the explanation. If the user wants to query which time period or time periods the teacher is playing the saxophone, the user can provide a corresponding natural language query, for example, "someone is playing the saxophone". If the user wants to query which time period or time periods the teacher is explaining how to play the saxophone, the user can provide a corresponding natural language query, for example, "someone is explaining playing the saxophone". If the user wants to query which time period or time periods the teacher is playing the saxophone again, the user can provide a corresponding natural language query, for example, "someone is playing the saxophone again".
[0027] It should be understood that although the application scenarios of the implementations of the present disclosure are introduced above in conjunction with the time period positioning of the video, multiple implementations of the present disclosure may be applied to any other suitable scenarios. For example, multiple implementations of the present disclosure may not be limited to video understanding, and may also be applied to other media streams, such as audio or voice. For another example, multiple implementations of the present disclosure may also be applied to other directions of video understanding, such as temporal behavior nomination or temporal behavior detection. For example, temporal behavior nomination is to determine the time period that may contain behavior from the video. Therefore, in this application, a query may not be required as input. Temporal behavior detection may classify nominations based on temporal behavior nominations, or may not require a query as input.
[0028] Figure 2 A schematic diagram of a neural network 200 for performing time segment location through natural language according to some implementations of the present disclosure is shown. A time segment, also called a moment or a time, represents a segment of a media stream such as a video.
[0029] like Figure 2 As shown, the neural network 200 receives a media stream 202 as an input, wherein the media stream 202 may be an untrimmed media stream, such as a video and / or an audio. The media stream 202 is typically a piece of media content having a certain duration, wherein the duration is defined by the start time and the end time of the media stream 202. The media stream may be divided into frames in the time dimension as a basic unit, and thus the media stream may also be referred to as a frame sequence. For ease of description, the neural network 200 is described below using a video V as an example. For example, the video V may be an untrimmed video, which may contain multiple behaviors.
[0030] In the time segment location scenario using natural language, the neural network 200 is expected to extract the video segment matching the query 206, that is, the time segment M. Specifically, the query statement can be expressed as where s i Indicates a word or phrase in a sentence and l S is the total number of words. A video V can be represented as a sequence of frames, i.e., where x i represents a frame in the video V, and l V represents the total number of frames. The neural network 200 is expected to make the extracted i To frame x j The period conveys the same meaning as the input sentence S.
[0031] like Figure 2 As shown, feature map extractor 204 extracts a two-dimensional temporal feature map from media stream 202 . Figure 3A and Figure 3B An example of a two-dimensional time series feature map according to some implementations of the present disclosure is shown. The two-dimensional time series feature map includes three dimensions, two of which are time dimensions X and Y, and one dimension is a feature dimension Z. For example, the X dimension represents the start time of a time period, and the Y dimension represents the end time of a time period. Figure 3B Shows Figure 3A The two time dimensions of the two-dimensional time series feature map are shown for ease of illustration.
[0032] For example, the video V can be divided into N video clips of the same length, where the length of each video clip is τ. Figure 3B As shown, the unit with coordinates (0, 1) represents a time period starting at 0 and ending at 2τ, including the first and second segments. The unit with coordinates (3, 3) represents a time period starting at 3τ and ending at 4τ, i.e., the fourth segment. Figure 3BIn the example, the portion where the X coordinate value is greater than the Y coordinate value is invalid, the portion where the X coordinate value is equal to the Y coordinate value represents a period of time of the length of a video segment, and the portion where the X coordinate value is greater than the Y coordinate value represents a period of time greater than the length of a video segment.
[0033] Alternatively, the duration of the time period can be used as the Y dimension. In this case, Figure 3A and Figure 3B A simple transformation of the coordinates is sufficient, so it will not be described in detail.
[0034] In addition, the neural network 200 may also receive a query 206 for the media stream 202. For example, the query 206 may be a sentence, or a segment of speech, etc. For convenience, the following description is made by taking the sentence S as the query 206 as an example.
[0035] like Figure 2 As shown, the language encoder 208 can encode the query 206 to generate a vector representation of the query 206. The query 206 can be encoded using various suitable methods currently known or developed in the future. The vector representation of the query 206 is provided to the temporal neighbor network 210 together with the two-dimensional temporal feature map. The temporal neighbor network 210 can combine the vector representation of the query 206 with the two-dimensional temporal feature map to determine the score associated with the behavior specified by the query 206 for each time period. For example, the temporal neighbor network 210 can perceive its context when determining the score of each time period to more conveniently extract the discriminative features of the time period.
[0036] As described above, the implementation of the present disclosure may also be applied to any other suitable scenarios, such as temporal behavior nomination or temporal behavior detection, etc. In these scenarios, query 206 and language encoder 208 may not be required, and only the two-dimensional temporal feature map needs to be provided to the temporal neighbor network for discrimination.
[0037] By using a two-dimensional temporal feature map, multiple time periods can be evaluated simultaneously, allowing for consideration of video context information between different time periods. Compared to evaluating a time period alone, the implementation of the present disclosure can extract discriminative features to distinguish different time periods. In this way, it is easier to extract the time period that best matches the query.
[0038] Figure 4 FIG. 4 is a schematic diagram of a feature map extractor 204 according to some implementations of the present disclosure. The feature map extractor 204 includes a sampler 401, which can divide a video V into small video segments, wherein each video segment v iK frames may be included. Then, sampling may be performed at fixed time intervals on these video segments, for example, at least a portion of the frames may be selected from the K frames for sampling. In this way, the sampler 401 may obtain N video segments, represented as
[0039] The feature extractor 402 can extract the features of each video segment and output a representation of the video segment. where d V is the feature dimension. For example, the feature extractor 402 may be implemented by a convolutional neural network (CNN). In some implementations, a convolutional neural network having d may be provided after the convolutional neural network. V The fully connected layer with output channels has an output feature dimension of d V The feature vector of .
[0040] The N video clips can be used as basic elements for constructing the candidate time period. Therefore, the feature map of the candidate time period can be constructed by using the features of the video clips obtained by the feature extractor 402. For example, the pooling layer 404 can extract the features of each time period from the features of the N video clips to obtain a two-dimensional temporal feature map 406, such as Figure 3A and Figure 3B As shown. For example, the pooling layer 404 may be a maximum pooling layer. As a non-parameter method, maximum pooling is helpful to reduce the amount of calculation of the system. For example, for each time period, the features of all the segments corresponding to the time period may be subjected to maximum pooling, and their features may be calculated. Where a and b represent the index of the start and end segments, and 0≤a≤b≤N-1. For example, for Figure 3B The (0, 1) coordinates are shown for the time period, and the maximum pooling is performed on the 1st and 2nd segments. Therefore, the longer time period is pooled on more adjacent segments, while the shorter time period is pooled on fewer segments.
[0041] It should be understood that other suitable methods may be used to replace the pooling layer 404. For example, a convolutional neural network may be used to replace the pooling layer 404. As an example, a stacked multi-layer convolutional neural network may be used to replace the pooling layer 404.
[0042] The feature map extractor 204 can construct all time periods into a two-dimensional temporal feature map 406, represented as As mentioned above, the two-dimensional temporal feature map F M It consists of three dimensions, two of which N represent the start and end segment indexes, and the third dimension d V Represents the dimension of the feature. a to v b The period of time is characterized by FM [a, b, :], where
[0043] The start and end segment indexes a and b of the time period should satisfy a≤b. Therefore, in the two-dimensional temporal feature map, all time periods located in the region a>b (i.e., the lower triangular part of the feature map) are invalid. For example, the features in this region can be filled with zeros.
[0044] In some implementations, all possible adjacent video segments can be used as candidate time periods. However, for subsequent processing, this processing method may bring about a large computational overhead. In some implementations, a sparse sampling scheme can be used to reduce the computational overhead. In the sparse sampling scheme, redundant time periods with more overlaps can be deleted from the candidate time periods. Specifically, dense sampling can be performed for time periods of shorter duration, and the sampling interval can be increased as the time period length increases.
[0045] Figure 5 An example of such a sparse sampling scheme is shown. Figure 5 Shown with Figure 3B A similar two-dimensional time series feature map, where each small square represents a time period. The start time and the index of the start time are shown on the X-axis, and the end time and the index of the end time are shown on the Y-axis. Figure 5 The lower triangular part is the invalid part, and the upper triangular part is the valid part. Figure 5 The upper triangular part of includes two patterns, black-filled and white-filled squares, where the black-filled squares are sampled as candidate time periods. Figure 5 It can be seen that the overlap between time periods with a small difference between the X and Y indexes is low, so all of them can be used as candidate time periods. As the difference between the X and Y indexes increases, the overlap between time periods gradually increases, so the selection of candidate time periods becomes increasingly sparse.
[0046] This coefficient sampling scheme can be expressed in mathematical language. For example, when the sampling segments are small (for example, N≤16), all possible time periods can be listed as candidate time periods, such as Figure 3B When N increases (for example, N>16), the following conditions can be selected from the fragment v a to v b The time periods are selected as candidate time periods:
[0047]
[0048] where a and b represent the indices of the fragment, mod represents modulo, & represents logical AND, and s and s' are defined as follows:
[0049]
[0050] in, and Indicates rounding up. If G(a, b) = 1, the period is selected as a candidate period, otherwise the period is not selected as a candidate period. This sampling strategy can greatly reduce the number of candidate periods, thereby reducing computational overhead.
[0051] Figure 6 2 shows a schematic diagram of a speech encoder 208 according to some implementations of the present disclosure. Figure 6 As shown, the language encoder 208 includes an embedding layer 602 and a long short-term memory unit (LSTM) network 604. For example, for each word S in the query 206 (e.g., input sentence S), i , the embedding layer 602 can generate an embedding vector representing the word where d S Indicates the vector length. The embedding layer 602 can embed the word The output is sent to the LSTM network 604, which may include a multi-layer bidirectional LSTM, for example, a three-layer bi-layer LSTM. The LSTM network 604 may use its final hidden state as the feature representation of the input sentence S, i.e., the feature vector 606, represented as The feature vector 606 encodes the language structure of the query 206, so as to reflect the information of the time period that the query 206 is interested in. It should be understood that Figure 6 The illustrated encoder 208 is provided by way of example only, and any other suitable encoder architecture currently known or developed in the future may also be used.
[0052] Figure 7 FIG. 2 shows a schematic diagram of a timing adjacent network 210 according to some implementations of the present disclosure. Figure 7 As shown, the fusion unit 702 can map the two-dimensional time series feature 406 or F M With the eigenvector 606 or f S In one implementation, the fuser 702 may include a fully connected layer that projects the two features into a unified subspace. The fuser 702 may also include a Hadamard product operation and a 2 Regularization, which fuses the projected features. The fused feature map can also have N×N×d V The fusion unit 702 can be expressed mathematically as:
[0053]
[0054] Among them, w S and W M represents the learning parameters of the fully connected layer, represents the transpose of a vector of all 1s, ⊙ represents the Hadamard product, and ||·|| 2 represents 12 regularization.
[0055] It should be understood that other neural network architectures can be used to implement the fuser 702. For example, the fuser 702 can be implemented by an attention network.
[0056] The context sensor 704 can build a temporal neighbor network on the fused two-dimensional temporal feature map F. For example, the context sensor 704 can include a convolutional neural network, for example, L convolutional layers, and the size of the convolution kernel is k. The output of the context sensor 704 can maintain the same shape as the input fused feature map, that is, N×N. The context sensor 704 allows the model to gradually perceive more context of adjacent candidate time periods while learning the difference between candidate time periods. In addition, the receptive field of the network is large, so the content of the entire video and sentence can be observed to learn temporal dependencies.
[0057] In some implementations, the convolutional neural network may include a dilated convolution. Sparse sampling as described above may be achieved by adaptively modifying the stride parameter of the dilated convolution. For example, for a 3×3 convolution kernel, if the stride is set to 1, it is equivalent to using a 5×5 convolution kernel, which includes 16 holes, that is, only 9 of the points are sampled.
[0058] Use a larger moving step size at a low sampling rate and a smaller moving step size at a high sampling rate. As mentioned above, the sampling rate is related to the length of the time period. For a longer time period, a lower sampling rate can be used; conversely, for a shorter time period, a higher sampling rate can be used. Therefore, for a longer time period, a larger moving step size can be used; conversely, for a shorter time period, a smaller moving step size can be used.
[0059] In the two-dimensional fused feature map, there are areas filled with zeros. In some implementations, when performing convolution on these areas, only the values on the valid areas can be calculated. In other words, the features filled with zeros are not considered in the calculation, thereby reducing the computational overhead.
[0060] The score predictor 706 can predict the score of the candidate time segment matching the natural language query based on the feature map output by the context sensor. For example, the score predictor 706 can include a fully connected layer and an activation function layer. For example, a sigmoid function can be used as an activation function.
[0061] Figure 8 A schematic diagram of a two-dimensional timing mapping according to some implementations of the present disclosure is shown. Figure 3A and Figure 3B The two-dimensional temporal feature map shown corresponds to the two-dimensional temporal feature map. The two-dimensional temporal map is also called a two-dimensional score map, in which each square represents the score of the corresponding candidate time period. For example, the coordinate (1, 4) shows the highest score, and the time period from time τ to time 5τ is the time period that best matches the query.
[0062] For example, all valid scores on the two-dimensional temporal map can be collected according to the candidate indication G(a, b) in equation (1), expressed as Where C is the total number of candidate time slots. Each value p on the two-dimensional time series map i Represents the matching score between the candidate time period and the query statement, where the maximum value represents the best matching time period.
[0063] During the training process, the scaled overlap area (IoU) value can be used as a supervisory signal to reduce the computational overhead during training. Specifically, for each candidate period, the IoU score with the true period can be calculated. i Then, we can use two thresholds t min and t max To scale the IoU score i :
[0064]
[0065] And use y i As the supervision label. In some implementations, a binary cross-coupling loss function as shown in equation (5) can be used for training.
[0066]
[0067] where p i is the output score for the time period.
[0068] Fig. 9 A flowchart of a method 900 according to some implementations of the present disclosure is shown. The method 900 may be combined with Figure 1-Figure 8 For example, method 900 can be implemented in Figure 1 The NN module 122 shown in FIG. Figure 2 The neural network 200 shown is implemented.
[0069] At block 902, a two-dimensional temporal feature map representing a plurality of time periods in the media stream is extracted from the media stream, wherein the two-dimensional temporal feature map comprises a first dimension representing the beginning of a time period in the plurality of time periods and a second dimension representing the end of a time period in the plurality of time periods. For example, the two-dimensional temporal feature map may be as follows: Figure 3A-3B and Figure 5The two-dimensional temporal feature map shown. For example, the first dimension may be the beginning of the candidate period, and the second dimension may be the end of the candidate period. For example, the media stream may include unintercepted media streams, such as video and / or audio.
[0070] In some implementations, extracting the two-dimensional timing feature map includes: segmenting the media stream into multiple segments; extracting features of each of the multiple segments to obtain a feature map of the media stream; and extracting features of a segment corresponding to a time period in the multiple time periods from the feature map of the media stream as part of the two-dimensional timing feature map.
[0071] In some implementations, extracting the features of the specific candidate time period includes: extracting the features of the specific candidate time period from features of a segment corresponding to the specific candidate time period by pooling. For example, the pooling may be maximum pooling.
[0072] In block 904, based on the two-dimensional temporal feature map, the correlation between the plurality of candidate time periods and the behavior in the media stream is determined. The correlation may be embodied in the form of a score, a probability, etc. In a time period positioning application using natural language, the behavior may be a behavior specified by a query. For temporal behavior nomination and temporal behavior detection, the behavior may be any possible behavior in the media stream.
[0073] In some implementations, determining the correlation includes: sampling the multiple time periods at a sampling rate to determine multiple candidate time periods, wherein the sampling rate is adaptively adjusted based on the duration of each of the multiple time periods; and determining the correlation between the multiple candidate time periods and the behavior in the media stream.
[0074] In some implementations, the sampling rate is set such that as the length of a time period increases, the sampling rate decreases.
[0075] In some implementations, determining the correlation includes: applying a convolutional layer to the two-dimensional temporal feature map to obtain another feature map having the same dimension as the two-dimensional temporal feature map; and determining, based on the other feature map, scores associated with the multiple time periods and behaviors in the media stream.
[0076] In some implementations, the convolution layer includes a dilated convolution, and a moving step of the dilated convolution is configured such that the moving step increases as the length of the time period increases.
[0077] In some implementations, determining the correlation includes: in response to receiving a query for a specific behavior in the media stream, extracting a feature vector of the query; and determining the correlation based on the feature vector of the query and the two-dimensional temporal feature map.
[0078] In some implementations, determining the correlation includes: fusing the feature vector of the query and the two-dimensional time series feature map to generate another two-dimensional time series feature map having the same dimension as the two-dimensional time series feature map; and determining the correlation between the multiple time periods and the specific behavior based on the other two-dimensional time series feature map.
[0079] In some implementations, fusing the query feature vector and the two-dimensional time series feature map includes: generating the another two-dimensional time series feature map by applying a Hadamard product to the query feature vector and the two-dimensional time series feature map.
[0080] Some example implementations of the present disclosure are listed below.
[0081] In a first aspect, a computer-implemented method is provided. The method includes: extracting a two-dimensional temporal feature map representing a plurality of time periods in a media stream, wherein the two-dimensional temporal feature map includes a first dimension representing a start of a time period in the plurality of time periods and a second dimension representing an end of a time period in the plurality of time periods; and determining, based on the two-dimensional temporal feature map, a correlation of the plurality of time periods with behaviors in the media stream.
[0082] In some implementations, extracting the two-dimensional timing feature map includes: segmenting the media stream into multiple segments; extracting features of each of the multiple segments to obtain a feature map of the media stream; and extracting features of a segment corresponding to a time period in the multiple time periods from the feature map of the media stream as part of the two-dimensional timing feature map.
[0083] In some implementations, determining the correlation includes: sampling the multiple time periods at a sampling rate to determine multiple candidate time periods, wherein the sampling rate is adaptively adjusted based on the duration of each of the multiple time periods; and determining the correlation between the multiple candidate time periods and the behavior in the media stream.
[0084] In some implementations, the sampling rate is set such that as the length of a time period increases, the sampling rate decreases.
[0085] In some implementations, determining the correlation includes: applying a convolutional layer to the two-dimensional temporal feature map to obtain another feature map having the same dimension as the two-dimensional temporal feature map; and determining, based on the other feature map, scores associated with the multiple time periods and behaviors in the media stream.
[0086] In some implementations, the convolution layer includes a dilated convolution, and a moving step of the dilated convolution is configured such that the moving step increases as the length of the time period increases.
[0087] In some implementations, determining the correlation includes: in response to receiving a query for a specific behavior in the media stream, extracting a feature vector of the query; and determining the correlation based on the feature vector of the query and the two-dimensional temporal feature map.
[0088] In some implementations, determining the correlation includes: fusing the feature vector of the query and the two-dimensional time series feature map to generate another two-dimensional time series feature map having the same dimension as the two-dimensional time series feature map; and determining the correlation between the multiple time periods and the specific behavior based on the other two-dimensional time series feature map.
[0089] In some implementations, fusing the query feature vector and the two-dimensional time series feature map includes: generating the another two-dimensional time series feature map by applying a Hadamard product to the query feature vector and the two-dimensional time series feature map.
[0090] In some implementations, the first dimension is a start of a time period in the plurality of time periods and the second dimension is an end of a time period in the plurality of time periods.
[0091] In some implementations, the media stream includes an unintercepted media stream.
[0092] In some implementations, extracting the features of the specific candidate time period includes: extracting the features of the specific candidate time period from features of a segment corresponding to the specific candidate time period by pooling.
[0093] In some implementations, the pooling includes max pooling.
[0094] In some implementations, the media stream includes at least one of a video stream and an audio stream.
[0095] In a second aspect, a device is provided, comprising: a processing unit; and a memory coupled to the processing unit and containing instructions stored thereon, wherein the instructions, when executed by the processing unit, cause the device to perform the following actions: extract a two-dimensional timing feature map representing multiple time periods within the media stream from a media stream, wherein the two-dimensional timing feature map comprises a first dimension representing the start of a time period in the multiple time periods and a second dimension representing the end of a time period in the multiple time periods; and determine, based on the two-dimensional timing feature map, a correlation between the multiple time periods and behaviors in the media stream.
[0096] In some implementations, extracting the two-dimensional timing feature map includes: segmenting the media stream into multiple segments; extracting features of each of the multiple segments to obtain a feature map of the media stream; and extracting features of a segment corresponding to a time period in the multiple time periods from the feature map of the media stream as part of the two-dimensional timing feature map.
[0097] In some implementations, determining the correlation includes: sampling the multiple time periods at a sampling rate to determine multiple candidate time periods, wherein the sampling rate is adaptively adjusted based on the duration of each of the multiple time periods; and determining the correlation between the multiple candidate time periods and the behavior in the media stream.
[0098] In some implementations, the sampling rate is set such that as the length of a time period increases, the sampling rate decreases.
[0099] In some implementations, determining the correlation includes: applying a convolutional layer to the two-dimensional temporal feature map to obtain another feature map having the same dimension as the two-dimensional temporal feature map; and determining, based on the other feature map, scores associated with the multiple time periods and behaviors in the media stream.
[0100] In some implementations, the convolution layer includes a dilated convolution, and a moving step of the dilated convolution is configured such that the moving step increases as the length of the time period increases.
[0101] In some implementations, determining the correlation includes: in response to receiving a query for a specific behavior in the media stream, extracting a feature vector of the query; and determining the correlation based on the feature vector of the query and the two-dimensional temporal feature map.
[0102] In some implementations, determining the correlation includes: fusing the feature vector of the query and the two-dimensional time series feature map to generate another two-dimensional time series feature map having the same dimension as the two-dimensional time series feature map; and determining the correlation between the multiple time periods and the specific behavior based on the other two-dimensional time series feature map.
[0103] In some implementations, fusing the query feature vector and the two-dimensional time series feature map includes: generating the another two-dimensional time series feature map by applying a Hadamard product to the query feature vector and the two-dimensional time series feature map.
[0104] In some implementations, the first dimension is a start of a time period in the plurality of time periods and the second dimension is an end of a time period in the plurality of time periods.
[0105] In some implementations, the media stream includes an unintercepted media stream.
[0106] In some implementations, extracting the features of the specific candidate time period includes: extracting the features of the specific candidate time period from features of a segment corresponding to the specific candidate time period by pooling.
[0107] In some implementations, the pooling includes max pooling.
[0108] In some implementations, the media stream includes at least one of a video stream and an audio stream.
[0109] In a third aspect, the present disclosure provides a computer program product, which is tangibly stored in a non-transitory computer storage medium and includes computer executable instructions, which when executed by a device cause the device to perform the method in the first aspect of the present disclosure.
[0110] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by a device, cause the device to perform the method in the first aspect of the present disclosure.
[0111] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0112] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0113] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0114] In addition, although each operation is described in a specific order, this should be understood as requiring such operation to be performed in the specific order shown or in a sequential order, or requiring that all illustrated operations should be performed to obtain the desired result. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of a separate implementation can also be implemented in a single implementation in combination. On the contrary, the various features described in the context of a single implementation can also be implemented in multiple implementations individually or in any suitable sub-combination.
[0115] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.
Claims
1. A computer-implemented method, include: Extracting a two-dimensional temporal feature map representing a plurality of time periods within the media stream from a media stream, wherein the two-dimensional temporal feature map comprises a first dimension representing a start of a time period in the plurality of time periods and a second dimension representing an end of a time period in the plurality of time periods; Encode sentence features extracted from the input; fusing the encoded sentence features and the two-dimensional temporal feature map into a unified subspace as a fused two-dimensional temporal map; generating a temporal neighbor network using the fused two-dimensional temporal map; Determining, based on the fused two-dimensional temporal map, correlations of the plurality of time periods with behaviors in the media stream using the temporal neighbor network; as well as Based on the correlation, a set of candidate time periods that match the input are identified.
2. The method according to claim 1, wherein the two-dimensional temporal feature map is extracted include: dividing the media stream into a plurality of segments; Extracting features of each of the multiple segments to obtain a feature map of the media stream; as well as The feature of the time period is extracted from the feature of the segment corresponding to the time period in the plurality of time periods in the feature map of the media stream to serve as a part of the two-dimensional time series feature map.
3. The method according to claim 1, wherein determining the correlation include: Sampling the multiple time periods at a sampling rate to determine a plurality of candidate time periods, wherein the sampling rate is adaptively adjusted based on a duration of each of the multiple time periods; as well as Determine correlations between the plurality of candidate time periods and behaviors in the media stream. The method according to claim 3 , wherein the sampling rate is set such that the sampling rate decreases as the duration of the time period increases.
5. The method of claim 1, wherein determining the correlation include: Applying a convolutional layer to the fused two-dimensional temporal feature map to obtain another feature map having the same dimension as the two-dimensional temporal feature map; as well as Based on the another feature map, scores are determined for the plurality of time periods to be associated with behaviors in the media stream.
6. The method according to claim 5, wherein the convolution layer comprises a dilated convolution, and the moving step of the dilated convolution is configured such that the moving step increases as the length of the time period increases.
7. The method of claim 1, wherein determining the correlation include: In response to receiving a query for a specific behavior in the media stream, extracting a feature vector of the query; as well as The correlation is determined based on the feature vector of the query and the two-dimensional temporal feature map.
8. The method of claim 7, wherein determining the correlation include: fusing the query feature vector and the two-dimensional time series feature map to generate another two-dimensional time series feature map having the same dimension as the two-dimensional time series feature map; as well as The correlation between the multiple time periods and the specific behavior is determined based on the another two-dimensional time series feature map.
9. The method according to claim 8, wherein the query feature vector and the two-dimensional temporal feature map are fused include: The another two-dimensional time series feature map is generated by applying a Hadamard product to the query feature vector and the two-dimensional time series feature map.
10. The method of any one of claims 7-9, wherein the query comprises a natural language query.
11. The method of claim 1, wherein the media stream comprises an unintercepted media stream.
12. An electronic device, include: Processing unit; as well as A memory coupled to the processing unit and containing instructions stored thereon, which, when executed by the processing unit, cause the device to perform the following actions: Extracting a two-dimensional temporal feature map representing a plurality of time periods within the media stream from a media stream, wherein the two-dimensional temporal feature map comprises a first dimension representing a start of a time period in the plurality of time periods and a second dimension representing an end of a time period in the plurality of time periods; Encode sentence features extracted from the input; fusing the encoded sentence features and the two-dimensional temporal feature map into a unified subspace as a fused two-dimensional temporal map; generating a temporal neighbor network using the fused two-dimensional temporal map; Determining, based on the fused two-dimensional temporal map, correlations of the plurality of time periods with behaviors in the media stream using the temporal neighbor network; as well as Based on the correlation, a matching set of candidate time periods for the input is identified.
13. The apparatus according to claim 12, wherein the two-dimensional temporal feature map is extracted include: dividing the media stream into a plurality of segments; Extracting features of each of the multiple segments to obtain a feature map of the media stream; as well as The feature of the time period is extracted from the feature of the segment corresponding to the time period in the plurality of time periods in the feature map of the media stream to serve as a part of the two-dimensional time series feature map.
14. The apparatus of claim 12, wherein determining the correlation include: Sampling the multiple time periods at a sampling rate to determine a plurality of candidate time periods, wherein the sampling rate is adaptively adjusted based on a duration of each of the multiple time periods; as well as Determine correlations between the plurality of candidate time periods and behaviors in the media stream.
15. The apparatus of claim 12, wherein determining the correlation include: Applying a convolutional layer to the two-dimensional temporal feature map to obtain another feature map having the same dimension as the two-dimensional temporal feature map; as well as Based on the another feature map, scores are determined for the plurality of time periods to be associated with behaviors in the media stream.
16. The apparatus according to claim 15, wherein the convolution layer comprises a dilated convolution, and a moving step of the dilated convolution is configured such that the moving step increases as the length of a time period increases.
17. The apparatus of claim 15, wherein determining the correlation include: In response to receiving a query for a specific behavior in the media stream, extracting a feature vector of the query; as well as The correlation is determined based on the feature vector of the query and the two-dimensional temporal feature map.
18. The apparatus of claim 17, wherein determining the correlation include: Fuse the feature vector of the query and the two-dimensional temporal feature map to generate another two-dimensional temporal feature map having the same dimension as the two-dimensional temporal feature map; And Based on the another two-dimensional temporal feature map, determine the correlation between the multiple time periods and the specific behavior.
19. The apparatus according to claim 18, wherein fusing the feature vector of the query and the two-dimensional temporal feature map Comprises: Generate the another two-dimensional temporal feature map by applying a Hadamard product to the feature vector of the query and the two-dimensional temporal feature map.
20. A computer program product, the computer program product being stored in a computer storage medium and comprising computer-executable instructions, the computer-executable instructions causing the apparatus to perform actions when executed by the apparatus, the actions Comprises: Extract a two-dimensional temporal feature map representing multiple time periods within the media stream from the media stream, wherein the two-dimensional temporal feature map includes a first dimension representing the start of a time period among the multiple time periods and a second dimension representing the end of a time period among the multiple time periods; Encode the statement features extracted from the input; Fuse the encoded statement features and the two-dimensional temporal feature map into a unified subspace as the fused two-dimensional temporal map; Generate a temporal adjacent network using the fused two-dimensional temporal map; Based on the fused two-dimensional temporal map, use the temporal adjacent network to determine the correlation between the multiple time periods and the behavior in the media stream; And Based on the correlation, identify a set of candidate time periods matching the input.
Citation Information
Patent Citations
Video segment playlist generation in video management system
CN109564576A
Video attention moment retrieval method and device based on attention mechanism
CN110019849A