Media data processing method and apparatus, electronic device, and computer storage medium
By extracting and fusing global and local text features into video features, and utilizing a neural network model, the problem of inaccurate matching between video clips and text was solved, achieving video clip identification with higher semantic consistency.
Patent Information
- Application Number
- CN202110413772.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-16
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-07-09
AI Technical Summary
In existing technologies, video segments determined based on global text features and video features do not match the text descriptions, resulting in inaccurate matching.
Global and local text features of the text to be processed are extracted and fused into video features. The target segment that matches the text is then determined through a neural network model.
By fusing global and local text features, video clips and text are matched more accurately, improving semantic consistency.
Smart Images

Figure CN113761280B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence and cloud technology, in particular, the present application relates to a media data processing method and device, electronic equipment and computer storage medium. BACKGROUND
[0002] In the prior art, in order to obtain a video segment matching a text from a video, the global text feature of the text and the video feature of the video are usually used to determine a video segment matching the global text feature from the video feature.
[0003] In the prior art, since the global text feature cannot comprehensively represent all information of the text, the video segment determined based on the global text feature of the text and the video feature is not accurate, i.e., the video segment does not match the text description. SUMMARY
[0004] The present application aims to at least solve one of the above technical defects, and the following technical scheme is proposed to solve the problem that the video segment determined from the video does not match the text.
[0005] According to one aspect of the present application, a media data processing method is provided, which comprises:
[0006] obtaining a to-be-processed text and a to-be-processed video;
[0007] extracting a global text feature and a local text feature corresponding to the to-be-processed text, and a first video feature of the to-be-processed video, the global text feature comprising phrase features corresponding to each phrase contained in the to-be-processed text, and the local text feature comprising features corresponding to each unit text contained in the to-be-processed text;
[0008] fusing the global text feature into the first video feature to obtain a second video feature;
[0009] determining a target segment matching the to-be-processed text from the to-be-processed video according to the local text feature and the second video feature.
[0010] According to another aspect of the present application, a media data processing device is provided, which comprises:
[0011] a data acquisition module configured to acquire a to-be-processed text and a to-be-processed video;
[0012] a feature extraction module configured to extract a global text feature and a local text feature corresponding to the to-be-processed text, and a first video feature of the to-be-processed video, the global text feature comprising phrase features corresponding to each phrase contained in the to-be-processed text, and the local text feature comprising features corresponding to each unit text contained in the to-be-processed text;
[0013] The feature fusion module is configured to fuse the global text feature into the first video feature to obtain a second video feature.
[0014] The target segment determination module is configured to determine a target segment matching the to-be-processed text from the to-be-processed video according to the local text feature and the second video feature.
[0015] According to still another aspect of the present application, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the media data processing method of the present application when executing the computer program.
[0016] According to still another aspect of the present application, a computer readable storage medium is provided, which stores a computer program executable by a processor to implement the media data processing method of the present application.
[0017] The present application also provides a computer program product or a computer program, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the method provided in any of the optional implementation manners of the media data processing method.
[0018] The technical scheme provided by the present application has the beneficial effects that:
[0019] The media data processing method, device, electronic device, and computer readable storage medium provided by the present application can first perform preliminary processing on the first video feature based on the global text feature of the to-be-processed text and the first video feature of the to-be-processed video to obtain a second video feature, and then determine a target segment matching the to-be-processed text from each video segment in the to-be-processed video based on the local text feature of the to-be-processed text and the second video feature. Since the global text feature and the local text feature can describe all information of the to-be-processed text from different granularities, the target segment determined by the text features of different granularities (the global text feature and the local text feature) in the present application is more matched with the to-be-processed text and has a closer semantic relationship.
[0020] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the description of the embodiments of the present application will be briefly introduced.
[0022] Figure 1 A flowchart of a media data processing method provided by an embodiment of the present application is shown in FIG. 2.
[0023] Figure 2 A schematic diagram of a feature adjustment process of a text to be processed and a video to be processed provided by an embodiment of the present application is shown in FIG. 3.
[0024] Figure 3 A schematic diagram of a network structure of a neural network model provided by an embodiment of the present application is shown in FIG. 4.
[0025] Figure 4 A schematic diagram of a data processing flow in an encoder and a decoder provided by an embodiment of the present application is shown in FIG. 5.
[0026] Figure 5 A schematic diagram of an implementation environment of a media data processing method provided by an embodiment of the present application is shown in FIG. 6.
[0027] Figure 6 A schematic diagram of an implementation environment of another media data processing method provided by an embodiment of the present application is shown in FIG. 7.
[0028] Figure 7 A structural schematic diagram of a media data processing apparatus provided by an embodiment of the present application is shown in FIG. 8.
[0029] Figure 8 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 9. DETAILED DESCRIPTION
[0030] The embodiments of the present application will be described in detail below, examples of which are shown in the drawings, wherein the same or similar notations represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary only, and are used to explain the present application, but cannot be interpreted as a limitation of the present application.
[0031] Those skilled in the art of the technology can understand that the singular forms "a", "an" and "the" used herein also include the plural forms unless specifically stated otherwise. It should be further understood that the use of the term "include" in the specification of the application means that the features, integers, steps, operations, elements, and / or components described therein are present, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or there can be an intermediate element. In addition, "connected" or "coupled" used herein can include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any single unit and all combinations of the associated listed items.
[0032] The embodiments of the present application are to provide a media data processing method for accurately obtaining a video segment consistent with a text semantic expression in a video. The method can be applied to any scene where it is necessary to determine a video segment corresponding to a text to be processed in a video to be processed. The phrase features of each phrase in the text to be processed and the multi-modal features corresponding to the video to be processed and the text to be processed involved in the method can be realized by artificial intelligence technology, specifically involving machine learning and deep learning technology in the field of artificial intelligence technology. The data processing involved in the method can be realized by cloud technology.
[0033] In the schemes provided in the optional embodiments of the present application, the data processing (including but not limited to data calculation, etc.) involved in each optional embodiment can be realized by cloud computing. Cloud technology refers to a kind of hosting technology that unifies a series of resources such as hardware, software, network, etc. in a wide area network or local area network to realize data calculation, storage, processing and sharing. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on cloud computing business model application, which can form a resource pool, be used on demand, and be flexible and convenient. Cloud computing technology will become an important support. The background service of the technical network system needs a large amount of computing and storage resources, such as video websites, picture websites and more portals. With the high development and application of the Internet industry, every item in the future may have its own identification mark and needs to be transmitted to the background system for logical processing. Different levels of data will be processed separately, and data of various industries will need strong system support, which can only be realized by cloud computing.
[0034] Cloud computing is a computing model that distributes computing tasks on a large number of computing resources, so that various application systems can obtain computing power, storage space and information services according to needs. The network providing resources is called "cloud". The resources in the "cloud" are infinitely expandable to users and can be obtained at any time, used on demand, expanded at any time, and paid according to use.
[0035] As a basic capability provider of cloud computing, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform) is established, and a plurality of types of virtual resources are deployed in the resource pool for external customers to select and use. The cloud computing resource pool mainly includes: a computing device (a virtualized machine containing an operating system), a storage device, and a network device. According to logical functions, a PaaS (Platform as a Service) layer can be deployed on an IaaS (Infrastructure as a Service) layer, and a SaaS (Software as a Service) layer is deployed above the PaaS layer, or the SaaS is directly deployed on the IaaS. The PaaS is a platform for software running, such as a database and a web container. The SaaS is various business software, such as a web portal website and a short message massager. Generally, the SaaS and the PaaS are upper layers relative to the IaaS.
[0036] Cloud computing refers to a delivery and use model of IT infrastructure, that is, obtaining required resources in a scalable manner on demand through a network. Broadly, cloud computing refers to a delivery and use model of services, that is, obtaining required services in a scalable manner on demand through a network. Such services can be IT and software, Internet related, or other services. Cloud computing is a product of the development of grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, load balancing and other traditional computer and network technologies.
[0037] With the development of the Internet, real-time data flow, and the diversification of connected devices, and the demand for search services, social networks, mobile commerce, and open collaboration, cloud computing has rapidly developed. Unlike previous parallel distributed computing, the emergence of cloud computing will revolutionize the entire Internet model and enterprise management model.
[0038] The media data processing method provided in the application can also be implemented through an artificial intelligence cloud service, which is also commonly referred to as AIaaS (AI as a Service). This is a mainstream service mode of an artificial intelligence platform, specifically, the AIaaS platform will split several common AI services and provide independent or packaged services in the cloud.
[0039] This service mode is similar to opening an AI theme mall: all developers can access one or more artificial intelligence services provided by the platform through API interfaces, and some experienced developers can also use the AI framework and AI infrastructure provided by the platform to deploy and maintain their own cloud artificial intelligence services. In the present application, the AI framework and AI infrastructure provided by the platform can be used to implement the media data processing method provided in the present application.
[0040] Artificial intelligence (AI) is the use of digital computers or digital computer-controlled machines to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which aims to understand the essence of intelligence and produce a new intelligent machine that can reflect in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning, and decision-making.
[0041] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.
[0042] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithmic complexity theory, and other disciplines. It is a specialized study of how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure, and continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent. It is applied in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and adversarial learning.
[0043] With the research and progress of artificial intelligence technology, artificial intelligence technology is researched and applied in many fields, such as common smart home, smart wearable device, virtual assistant, smart speaker, smart marketing, unmanned vehicle, autonomous vehicle, unmanned aerial vehicle, robot, smart medical treatment, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0044] The scheme provided by the embodiments of the present application can be executed by any electronic device, which can be a user terminal device or a server. The server can be a physical server, a server cluster composed of multiple physical servers, or a distributed system, or a cloud server providing cloud computing services. The terminal device can include at least one of a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a smart television, and a smart vehicle device.
[0045] The technical scheme of the present application and how the technical scheme of the present application solves the above technical problems will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described in detail in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0046] The embodiments of the present application provide a possible implementation manner, as shown in Figure 1 A flowchart of a media data processing method is provided. The scheme can be executed by any electronic device, for example, the scheme of the embodiments of the present application can be executed on a terminal device or a server, or interactively executed by a terminal device and a server. For the convenience of description, the method provided by the embodiments of the present application will be described below with the server as an execution subject. As shown in the flowchart of Figure 1 The method can include the following steps:
[0047] In step S110, a to-be-processed text and a to-be-processed video are acquired.
[0048] In the present application, the data source of the to-be-processed video and the to-be-processed text is not limited. Optionally, the to-be-processed video can be an uncut video, at least one of the to-be-processed text or the to-be-processed video can be data sent by a user through a user terminal and received by a server corresponding to a multimedia data publishing platform, or can be data obtained from a preset storage space by the server corresponding to the multimedia publishing platform.
[0049] The to-be-processed text can be a text containing one or more languages, such as Chinese, English, etc. The language type of the to-be-processed text is not limited in the present application.
[0050] In step S120, the global text feature and the local text feature corresponding to the to-be-processed text and the first video feature of the to-be-processed video are extracted. The global text feature includes the phrase feature corresponding to each phrase contained in the to-be-processed text. The local text feature includes the feature corresponding to each unit text contained in the to-be-processed text.
[0051] The global text feature is a feature representing the overall information (overall semantic information) of the to-be-processed text, that is, the semantic information expressed by the to-be-processed text is represented by the global text feature. For example, if each phrase feature is spliced together as the global text feature of the to-be-processed text, the semantic information expressed by the to-be-processed text can be known based on the global text feature. The local text feature is a feature representing the local information (partial semantic information) of the to-be-processed text. For example, the local text feature can be a feature vector of each unit text contained in the to-be-processed text. The semantic information represented by the feature vector of each unit text represents the local feature of the to-be-processed text. The unit text can be at least one of a word or a word segment.
[0052] Optionally, if the unit text includes a word, the feature of each word contained in the to-be-processed text can be included in the local text feature. If the unit text includes a word segment, the word segment feature of each word segment contained in the to-be-processed text can be included in the local text feature. Since a word can be composed of at least one word, the word segment feature of each word segment can be obtained based on the feature of each word. A word can more accurately express the semantic information of the text than a word, and therefore, the word segment feature of each word segment contained in the to-be-processed text can also be included in the local text feature.
[0053] A video segment includes at least two adjacent video frame images. In the present application, the division manner of the to-be-processed video is not limited. For example, a set number of adjacent frame video images can be divided into a video segment, or a set time length of adjacent frame video images can be divided into a video segment.
[0054] The method of extracting the global text feature and the local text feature of the to-be-processed text and the first video feature will be described below, and will not be described here again.
[0055] In step S130, the global text feature is fused into the first video feature to obtain a second video feature.
[0056] The purpose of fusing the global text feature and the first video feature is to determine the second video feature preliminarily matched (which can be understood as coarse-grained positioning) with the to-be-processed text from the first video feature based on the global text feature.
[0057] In step S140, the target segment matched with the to-be-processed text is determined from the to-be-processed video according to the local text feature and the second video feature.
[0058] The local text feature can reflect the local information of the to-be-processed text from the local (which can be understood as fine-grained), and based on the local text feature and the second video feature, the association between the local text feature and the second video feature can be captured from the local, that is, based on the detailed features of the to-be-processed text provided by the local text feature, the determined target segment is closer to the semantics expressed by the to-be-processed text.
[0059] In an optional solution, the target segment can be determined based on the matching degree of the local text feature and the second video feature. The matching degree can be represented by the feature similarity. The more similar, the more matched, and the more similar the semantics.
[0060] The target segment refers to a segment of video in the to-be-processed video.
[0061] In the scheme of the present application, when the target segment matched with the to-be-processed text needs to be determined from the to-be-processed video, the first video feature can be preliminarily processed based on the global text feature of the to-be-processed text and the first video feature of the to-be-processed video to obtain the second video feature, and then the target segment matched with the to-be-processed text is determined from each video segment in the to-be-processed video based on the local text feature of the to-be-processed text and the second video feature. Since the global text feature and the local text feature can describe the overall information of the to-be-processed text from different granularities, the target segment determined by the text features of different granularities (global text feature and local text feature) in the scheme of the present application is more matched with the to-be-processed text, and the semantics is closer.
[0062] In an embodiment of the present application, the global text feature and the local text feature of the to-be-processed text are extracted, including:
[0063] Obtaining each unit text in the to-be-processed text and the positional relationship between each unit text;
[0064] determine features of each unit text based on the position relationship between each unit text and each unit text, and the local text features include the features of each unit text;
[0065] determine phrase features of each phrase contained in the to-be-processed text based on the features corresponding to each unit text;
[0066] fuse the phrase features of each phrase to obtain global text features.
[0067] The phrase features of each phrase can represent global information of the to-be-processed text from the perspective of a phrase, and the semantics of a phrase can more accurately reflect the global information of the to-be-processed text than a word. Therefore, in this scheme, the global text features are determined based on the phrase features of each phrase. Since a phrase can be composed of at least one unit text, the phrase features of each phrase can be determined based on the features of each unit text.
[0068] In this scheme, when extracting the features of each unit text, the context relationship (position relationship) between each unit text is also considered, that is, the relationship between a unit text and the unit text before and after the unit text. Therefore, based on the features of each unit text and the position relationship between each unit text, the features of each unit text are extracted, which can more accurately represent the semantic features of the to-be-processed text. The features of each unit text can represent the local information of the to-be-processed text from the perspective of a unit text, and therefore the features of each unit text can be used as local text features.
[0069] In an embodiment of the present application, the first video features include features of a plurality of video segments in the to-be-processed video; the global text features are fused into the first video features to obtain second video features, including:
[0070] fuse the global text features and the first video features to obtain fused video features;
[0071] obtain position information features corresponding to each video segment in the to-be-processed video;
[0072] superimpose the features of each video segment in the fused video features and the position information features corresponding to each video segment respectively to obtain superimposed video features;
[0073] obtain second video features based on the superimposed video features.
[0074] The purpose of fusing the global text features and the first video features is to determine video features in the first video features that preliminarily match the to-be-processed text based on the global text features.
[0075] The position information of each video segment in the second video feature is obtained by superimposing the feature of each video segment in the fused video feature and the position information feature corresponding to each video segment, that is, the superimposed video feature is the video feature containing the position information of each video segment. The position information feature is a feature corresponding to the position information.
[0076] In an embodiment of the present application, the first video feature includes the features of the plurality of video segments in the video to be processed; the second video feature is obtained based on the superimposed video feature, including:
[0077] The weight corresponding to each video segment is determined based on the association relationship of each video segment in the superimposed video feature;
[0078] For each video segment, the enhanced feature corresponding to the video segment is obtained based on the weight corresponding to the video segment and the feature corresponding to the video segment in the superimposed video feature;
[0079] The second video feature is extracted based on the enhanced feature corresponding to each video segment.
[0080] For a video segment, the association relationship of each video segment in the superimposed video feature includes the relationship between the video segment and itself, and the relationship between the video segment and other video segments except the video segment. In the superimposed video feature, each video segment has different importance in the video to be processed, which is represented by the weight corresponding to each video segment. The greater the weight, the more important it is.
[0081] The enhanced feature refers to adjusting the importance of the superimposed video feature corresponding to each video segment based on the weight of each video segment.
[0082] The enhanced feature corresponding to each video segment is extracted to obtain a deeper feature expression, that is, to make the second video feature contain more detailed video features.
[0083] Optionally, the weight corresponding to each video segment can be determined by a self-attention mechanism.
[0084] In an embodiment of the present application, the target segment matching the text to be processed is determined from the video to be processed according to the local text feature and the second video feature, including:
[0085] The association feature of the local text feature and the second video feature is determined;
[0086] According to the association feature and the second video feature, the local text feature is adjusted to obtain an adjusted local feature;
[0087] According to the association feature and the local text feature, the second video feature is adjusted to obtain an adjusted video feature;
[0088] Based on the adjusted local feature and the adjusted video feature, a target segment matching the to-be-processed text is determined from the to-be-processed video.
[0089] Among them, for the to-be-processed text and the to-be-processed video, the to-be-processed text can guide the to-be-processed video to focus on important information in the video, and the to-be-processed video can guide the to-be-processed text to focus on important information (such as keywords) in the text. Therefore, based on the two-way feature adjustment, the target segment obtained finally can be more matched with the to-be-processed text.
[0090] Among them, the association feature represents the association relationship between the second video feature and the local text feature. From the text point of view, the association feature can include features related to the second video feature in the local text feature. From the video point of view, the association feature can include features related to the local text feature in the second video feature.
[0091] Based on the association feature and the second video feature, the local text feature is adjusted, which means that based on the to-be-processed video, the to-be-processed text is guided to focus on important information in the text. Based on the association feature and the second video feature, it can be known which information in the local text feature is more important (fused with detailed information of the text). Based on the association feature and the local text feature, the second video feature is adjusted, which means that based on the to-be-processed text, the to-be-processed video is guided to focus on important information in the video. Based on the association feature and the local text feature, it can be known which information in the second video feature is more important.
[0092] Optionally, according to the association feature and the second video feature, the local text feature is adjusted to obtain an adjusted local feature, which can specifically include:
[0093] The first weight of each unit text in the local text feature is obtained; according to the association feature and the second video feature, the first weight of each unit text in the local text feature is adjusted, and based on the feature of each unit text and the adjusted weight corresponding to each unit text, the adjusted local feature is obtained.
[0094] The adjustment of the local text feature can be understood as adjusting the first weight corresponding to each unit text in the local text feature. The greater the adjusted weight, the more important the corresponding unit text is relative to the text to be processed. Since the second video feature is obtained based on the fusion of the global text feature and the first video feature, adjusting the local text feature can capture the details (local) information ignored in the second video feature through the adjusted local feature.
[0095] Similarly, the first video feature includes the features of multiple video segments in the video to be processed, and a video segment can be composed of at least one frame of video frame image. Adjusting the second video feature according to the association feature and the local text feature can mean adjusting the second weight corresponding to each video segment in the video to be processed according to the association feature and the local text feature. The greater the adjusted weight, the more important the corresponding video segment is relative to the video to be processed.
[0096] As an example, referring to the video to be processed and the text to be processed shown in Figure 2 In this example, the text to be processed is: a person is eating while standing and he is watching TV. Each word in the text to be processed is: "a, person, is, standing, eating, while, he is, watching, TV". In the video to be processed, there is a person standing and watching TV. In the first few frames of video frame images, the person is standing and watching TV while also eating. In the last few frames of video images, the person is only standing and watching TV and is not eating.
[0097] In this example, the first weight and the second weight are determined by the self-attention mechanism, so Figure 2 The video-text attention value shown in Figure 2 The text-video attention value shown in is the second weight. The grounding is the effective range of the attention value. The effective range of the text-video attention value is 0-1, and the effective range of the video-text attention value is 0-1.
[0098] As can be seen from Figure 2 , the attention values (adjusted first weights) corresponding to the words "eat" and "watch" in the text to be processed are relatively large (the group line marks on the words, the darker the color, the greater the attention value), so the words "eat" and "watch" in the text to be processed are relatively important information in the text to be processed. The attention values (adjusted second weights) corresponding to the A and B frames in the video to be processed are relatively large (the group line marks below A and B, the darker the color, the greater the attention value), so the A and B frames in the video to be processed are relatively important information in the video to be processed. Among them, the person in the A and B frames is standing and watching TV while eating.
[0099] Thus, based on the words "eat" and "see" in the text to be processed and the image A and image B in the video to be processed, the target segment matching the text to be processed can be accurately determined from the video to be processed (e.g. Figure 2 The video segment corresponding to the 0s-8.3s in the figure).
[0100] In an embodiment of the present application, the first video features include features of a plurality of video segments in the video to be processed; and determining the target segment matching the text to be processed from the video to be processed based on the adjusted local features and the adjusted video features includes:
[0101] determining the guidance information of the text to be processed to the video to be processed according to the adjusted local features;
[0102] adjusting the features of each video segment in the adjusted video features according to the guidance information to obtain third video features;
[0103] determining the target segment matching the text to be processed from the video to be processed according to the third video features.
[0104] The adjusted local text features serve as a supplement to the global text features and provide some detailed information in the text to be processed, so the guidance information determined based on the adjusted local features can fully convey the semantic information of the text to be processed. Adjusting the features of each video segment according to the guidance information means filtering out the features irrelevant to the semantic information of the text from the video segment, and obtaining the third video features, so that the target segment determined based on the third video features matches the text to be processed more.
[0105] Optionally, adjusting the features of each video segment in the adjusted video features according to the guidance information can be adjusting the weights corresponding to the features of each video segment, and the importance of the video features is represented by the size of the adjusted weights. The video features to be filtered out have relatively small adjusted weights, and the video features to be retained have relatively large adjusted weights.
[0106] In an embodiment of the present application, determining the target segment matching the text to be processed from the video to be processed according to the third video features includes:
[0107] determining the weights corresponding to each video segment according to the correlation between the features of each video segment included in the third video features;
[0108] weighting the third video features of each video segment based on the weights corresponding to each video segment to obtain fourth video features;
[0109] determine position information of the target segment in the to-be-processed video based on the fourth video feature.
[0110] determine the target segment in the to-be-processed video based on the position information.
[0111] The weight corresponding to the third video feature of a video segment represents the importance of the video segment in the to-be-processed video, and the fourth video feature determined based on the weight corresponding to each video segment can more accurately reflect the video feature in the to-be-processed video that matches the to-be-processed text.
[0112] The position information includes the start position and the end position of the video segment, and the position information can be represented by time information. The start time and the end time of a video segment represent the position information of the video segment.
[0113] Optionally, based on the fourth video feature, the position information corresponding to the target segment in the to-be-processed video can be determined by using a pre-trained video segment determination network. Based on the position information, it can be determined which video segment the target segment is.
[0114] The input of the video segment determination network is the video feature of a video, and the output is the position information corresponding to each video segment in the video. The video segment determination network can be trained based on the following method.
[0115] Obtain training data, which includes a plurality of sample videos carrying position labels. For a sample video, the position label represents the position information corresponding to each video segment in the sample video.
[0116] For a sample video, extract the video feature corresponding to the sample video, which includes the segment feature corresponding to each video segment.
[0117] For a sample video, input the video feature of the sample video into an initial neural network model to obtain the predicted position information corresponding to each video segment in the sample video.
[0118] Based on the predicted position information corresponding to each sample video and the position information corresponding to each position label, determine a training loss. For a sample video, the value of the training loss represents the difference between the predicted position information corresponding to the sample video and the position information corresponding to the position label of the sample video.
[0119] If the training loss satisfies a training end condition, the model corresponding to the end is used as the video segment determination network. If not, the model parameters of the initial neural network model are adjusted, and the initial neural network model is trained based on the training data.
[0120] In one embodiment of the present application, the global text features of the to-be-processed text are extracted, the global text features are fused into the first video features to obtain second video features, and the target segment matched with the to-be-processed text is determined from the to-be-processed video based on the local text features and the second video features based on a neural network model.
[0121] The neural network model comprises a phrase feature extraction network, a multi-modal feature extraction network, and a video segment determination network, and the neural network model is obtained by training in the following manner, which can specifically comprise the following steps:
[0122] Step 1: Obtain training data, the training data comprising a plurality of samples, each sample comprising a sample video and a sample text, and each sample carrying a position label. For a sample, the position label represents the position information of the corresponding target video segment of the sample text in the sample video.
[0123] The position label can be text, characters, etc., and the specific form of the position label is not limited in the present application.
[0124] Step 2: For each sample in the training data, extract the global text features and the local text features of the sample text in the sample, and the video features of the sample video.
[0125] The extraction method of the global text features and the local text features of the sample text is the same as that of the global text features and the local text features of the to-be-processed text, which is not described again here. The local text features of the sample text comprise the features of each unit text contained in the sample text, and the global text features of the sample text comprise the phrase features of each phrase contained in the sample text. The features of each unit text contained in the sample text can be extracted based on other methods, such as being extracted based on a pre-trained text feature extraction network. The video features of the sample video can also be extracted based on a trained network, thereby accelerating the model training speed.
[0126] Step 3: For the sample, input the features of each unit text in the sample text into the phrase feature extraction network to obtain the predicted phrase features of each phrase in the sample text.
[0127] Step 4: Determine a first loss value based on the matching degrees between the predicted phrase features of each phrase corresponding to each sample. For a sample, the first loss value represents the semantic difference between the phrases in the sample.
[0128] The matching degrees between the predicted phrase features of each phrase can be represented based on feature similarity. The more matched two phrases are, the closer the semantics between the two phrases are, and the smaller the semantic difference is.
[0129] Step 5, for the sample, input the global text feature and the local text feature of the sample text and the video feature of the sample video into the multi-modal feature extraction network to obtain the multi-modal video feature corresponding to the sample video.
[0130] The process of determining the multi-modal video feature based on the global text feature, the local text feature of the sample text, and the video feature of the sample video is consistent with the process of determining the third video feature based on the global text feature, the local text feature of the to-be-processed text, and the first video feature of the to-be-processed video, which will not be described here.
[0131] Specifically, at the decoder end, the video feature of the sample video can be adjusted based on the association feature between the local text feature of the sample text and the video feature of the sample video, and the local text feature of the sample text is adjusted based on the association feature, so that the model considers the mutual influence between the text feature and the video feature during training, and the determined multi-modal video feature is more accurate.
[0132] Optionally, at the decoder end, the association feature between the local text feature of the sample text and the video feature of the sample video can be determined based on the cooperative self-attention mechanism, and the local text feature and the video feature are alternately adjusted based on the association feature. This part will be described in detail below, and will not be described here.
[0133] Step 6, for the sample, input the multi-modal video feature into the video segment determination network to obtain the weight corresponding to each sample video segment in the multi-modal video feature, and based on the multi-modal video feature and the weight corresponding to each sample video segment, obtain the predicted position information of the corresponding predicted video segment of the sample text in the sample video.
[0134] Among the video features of the sample video, there are features of multiple sample video segments.
[0135] Step 7, based on the predicted position information corresponding to each sample and the position label of each sample, determine a second loss value, which represents the difference between the predicted position information corresponding to each sample and the position label of each sample.
[0136] For a sample, the second loss value corresponding to the sample represents the difference between the predicted position information corresponding to the sample and the position label corresponding to the sample, i.e. the difference between the position information corresponding to the predicted position information and the position label.
[0137] Step 8, based on each sample video segment in the multi-modal video feature corresponding to each sample and the position label corresponding to each sample, determine a third loss value, for a sample, the third loss value represents the possibility that each sample video segment in the sample is a target video segment.
[0138] wherein, for a sample, the greater the weight, the greater the likelihood that the corresponding sample video clip is the target video clip.
[0139] Step 9, based on the first loss value, the second loss value and the third loss value, determine the value of the training loss function corresponding to the neural network model; if the training loss function converges, the model corresponding to the convergence is taken as the final neural network model, if not, adjust the model parameters of the neural network model, and train the neural network model based on the training data.
[0140] Optionally, the second loss value can be represented by an L1 average absolute error loss function, and the second loss value is the average of the sum of the absolute differences between the position information corresponding to the position label and the predicted position information.
[0141] Optionally, the second loss value can also be represented by an L2 least square error loss function, that is, the second loss value is represented by the least square error, and the second loss value is the average of the sum of the squares of the differences between the position information corresponding to the position label and the predicted position information.
[0142] The inputs of the second loss function are normalized in the interval 0~1, when the second loss value is close to 0, that is, the difference between the predicted position information corresponding to the sample and the position information corresponding to the position label of the sample is small, the gradient of the L2 loss is smaller than that of the L1 loss, therefore, the second loss function is determined by L2, and the training stability is better. When the second loss value is larger, and considering that the input value is less than 1, therefore, the penalty effect of L1 loss on deviation is better than that of L2 loss, at this time, the second loss value is determined by L1, and the accuracy is higher, therefore, in the present scheme, the second loss value is determined based on L1 and L2, and a threshold parameter is introduced in L1 and L2 , so as to better balance the robustness and accuracy of the model.
[0143] The specific scheme is: in an embodiment of the present application, for a sample, based on the predicted position information corresponding to each sample and the position label, determine the second loss value, comprising:
[0144] Based on the predicted position information corresponding to the sample and the position label, determine the position deviation value;
[0145] If the absolute value of the position deviation value is less than the threshold parameter, determine the second loss value based on the least square error loss function corresponding to the position deviation value;
[0146] If the absolute value of the position deviation value is not less than the threshold parameter, determine the second loss value based on the loss function corresponding to the position deviation value, and the loss function includes the average absolute error loss function and the threshold parameter.
[0147] The second loss value corresponding to a sample is determined according to the following formula:
[0148]
[0149]
[0150] wherein, is a position deviation value corresponding to the i-th sample in the training data, that is, a difference between the predicted position information corresponding to the sample and the position information corresponding to the position label, is a threshold parameter, represents an absolute value of the position deviation value, is a loss function corresponding to the position deviation value, is a least square error loss function (L2) corresponding to the position deviation value, is a second loss value corresponding to any sample in the training data, is the second loss value corresponding to the i-th sample, represents a number of samples in the training data, wherein, .
[0151] When the absolute value of the position deviation value is less than the threshold parameter, the second loss value is determined according to the L2 loss function corresponding to the position deviation value.
[0152] When the absolute value of the position deviation value is not less than the threshold parameter, the second loss value is determined according to the loss function corresponding to the position deviation value.
[0153] Optionally, the predicted position information includes a predicted start position and a predicted end position, and the position information corresponding to the position label includes a labeled start position and a labeled end position.
[0154] Then, for a sample, the second loss value corresponding to the sample includes a start loss and an end loss, the start loss representing a difference between the start position and the predicted start position, and the end loss representing a difference between the start position and the predicted end position.
[0155] Then, the formula corresponding to the second loss value can be as follows:
[0156]
[0157] wherein, represents a difference (position deviation value) between the start position corresponding to the i-th sample and the predicted start position, is the start loss corresponding to the i-th sample, the end loss corresponding to the ith sample, the second loss value corresponding to the ith sample, the second loss value corresponding to the ith sample, denotes the number of samples in the training data.
[0158] The training of the neural network model in the present application will be described in detail below in combination with the neural network structure diagram shown in FIG. 1. Figure 3 The training of the neural network model in the present application will be described in detail below in combination with the neural network structure diagram shown in FIG. 1.
[0159] The neural network model includes a cascaded input encoding module, a multi-modal fusion module (multi-modal feature extraction network) and a time sequence positioning module. The input encoding module includes a phrase feature extraction network (SPE shown in the figure).
[0160] The training data includes a plurality of samples, each sample including a sample video and a sample text, and each sample carries a position label. For a sample, the position label represents the position information of the target video segment corresponding to the sample text in the sample video.
[0161] The processing flow of each module involved in the present application will be described below taking a sample as an example:
[0162] First, the sample is input into the input encoding module, which includes a pre-trained text feature extraction network, a phrase feature extraction network, and a video feature extraction network. In the present example, the text feature extraction network can be composed of a bidirectional LSTM (Long Short-Term Memory, Long Short-Term Memory Network) (Bi-LSTM shown in the figure).
[0163] For the sample text, the present example takes a word as an example, the features (feature vectors) of each word in the sample text can be extracted by GloVe. Specifically, the initial features of each word in the sample text can be obtained by GloVe, and optionally, a 300-dimensional embedding can be extracted. Then, based on the initial features of each word, the features of each word containing context relationship can be obtained based on bidirectional LSTM, which can be specifically represented as: wherein L is the number of words in the sample text, denotes the feature of each word.
[0164] wherein the features (which can be referred to as word vectors) of each word can be obtained by the following formula:
[0165]
[0166] wherein, denotes the number of words in the sample text, denotes a feature obtained by using the forward LSTM and containing historical information (words before a word), denotes a feature obtained by using the backward LSTM and containing future information (words after a word), denotes a feature corresponding to the ith word containing context information, i is greater than or equal to 1 and less than or equal to L.
[0167] The local text features of the sample text include features corresponding to the words.
[0168] A sample text can include multiple phrases, and a phrase can be a word or can be composed of at least two words. For a sample text, the overall information expressed by the text cannot be fully and accurately summarized by a text feature corresponding to the sample text. Therefore, in this embodiment, the phrases in the sample text can be determined based on the words and the positional relationship of the words, and then the phrase features of the phrases can be extracted by a phrase feature extraction network, and the global text features of the sample text can be obtained by fusing the phrase features of the phrases.
[0169] The global text features can be denoted as: wherein k is the number of phrases, denotes the phrase feature of the first phrase.
[0170] Optionally, for a sample text, the number of phrases in the sample text is 3, and the performance of the model is best.
[0171] In this example, the phrase features of the phrases can be extracted based on a phrase feature extraction network, and the phrase feature extraction network can be trained based on the following training method:
[0172] For each sample text in the training data, the features of the words in each sample text are extracted.
[0173] For each sample text, the features of the words in the sample text are input into the phrase feature extraction network, so that the phrase feature extraction network determines the predicted phrase features of the phrases in the sample text based on the features of the words.
[0174] Based on the matching degree between the predicted phrase features of the phrases corresponding to each sample, a first loss value is determined, and for a sample text, the first loss value represents the semantic difference between the phrases in the sample text.
[0175] Optionally, for a sample text, the phrase features (predicted phrase features of the phrases) of the phrases contained in the sample text can be determined based on the features of the words in the sample text and the weights of the words.
[0176] The weights of the phrases can be denoted as:
[0177]
[0178] wherein is a loss function, and is a network parameter, is a feature of each word, is a matrix of K L, L is the number of words in the sample text, K is the number of phrases in the sample text, the elements in each row represent the weight of each word in the sample text in the phrase (each word in the corresponding phrase), and the weight of a word represents the importance of the word in the current phrase (the phrase where the word is located) in the sample text, and K is less than or equal to L.
[0179] Multiplying and can include global text features of K different semantic phrases , i.e., the global text features corresponding to the sample text.
[0180] For a sample, the first loss value corresponding to the sample can be represented as:
[0181]
[0182] wherein, is the first loss value corresponding to the sample, is the transpose of , each element in may represent the similarity between any two phrases in the phrase, the elements on the diagonal of represent the similarity of each phrase to itself (1), each element on the non-diagonal line represents the similarity between a phrase and other phrases, is the Frobenius norm, is an identity matrix, which makes the elements on the diagonal of zero, the smaller the value of the element on the non-diagonal line, the better the difference between the extracted phrases.
[0183] Based on the features of each word in the sample text, the global text features of the sample text can be obtained by the trained phrase feature extraction network, which can be specifically represented as:
[0184]
[0185] wherein, is the transpose of , and G denotes global text features including K different semantic phrases for each word.
[0186] For a sample video, video features can be extracted from the sample video based on a video feature extraction network, which can include a 3D convolutional neural network (denoted as ) and at least one fully connected layer. First video features are extracted from the sample video by the 3D convolutional neural network, and the first video features pass through the fully connected layer to obtain an embedding expression of the first video features of the sample video:
[0187]
[0188] wherein, is the first video feature, the first video feature including features of a plurality of video segments in the sample video, is a feature of a first video segment.
[0189] wherein, the video feature extraction network can be denoted as:
[0190] wherein, is the sample video, is a parameter of the video feature extraction network, is a nonlinear activation function of the video feature extraction network.
[0191] For a sample, the local text features and the global text features of the sample text in the sample, and the video features of the sample video are input into a multi-modal feature extraction network (a multi-modal fusion module shown in the figure) to obtain a multi-modal video feature corresponding to the sample video, wherein the multi-modal video feature includes multi-modal video features corresponding to each sample video segment.
[0192] Inside the multi-modal feature extraction network, as shown in Figure 3 , the multi-modal fusion module includes an encoder and a decoder. At the encoder end, considering that the attention mechanism cannot express the time sequence relationship between each sample video segment in the sample video, the global text features (text shown in Figure 4 ) and the first video features (video features of the sample video, video shown in Figure 4 ) are fused to obtain a fused video feature. Then, corresponding position information features (PE shown in Figure 4 ) of the plurality of sample video segments in the sample video are obtained. The corresponding position information features are added to the fused video feature (denoted as The process involves combining the features of each video segment in the fused video features with the location information features corresponding to each video segment to obtain the superimposed video features, which are video features containing location information.
[0193] The encoder includes a feature extraction layer (in this example, the feature extraction layer can be a multilayer perceptron (MLP)) and a self-attention module. In this example, it includes an M-layer perceptron. After obtaining the superimposed video features, the features corresponding to each sample video segment in the superimposed video features can be input into the self-attention module. The self-attention module determines the weight corresponding to each sample video segment based on the correlation between each sample video segment in the superimposed video features. For each sample video segment, the enhanced features corresponding to the video segment are obtained based on the weight corresponding to the sample video segment and the features corresponding to the video segment in the superimposed video features.
[0194] In this example, the self-attention module consists of multi-head self-attention and a forward propagation network (FFN), as follows: Figure 3 The three shown And FFN. For a sample video segment, it is processed through three weight-sharing self-attention modules. Each time, the superimposed video features corresponding to the sample video segment are input into the multi-head self-attention module to obtain the superimposed video features. Specifically, the query (Q) vector, key (K) vector, and value vector (V) are all the same input features, and the input features are fused using the standard attention calculation formula. Then, addition and normalization operations are performed to alleviate the gradient vanishing problem that is prone to occur in deep networks during training, resulting in the enhanced features corresponding to the input features (superimposed video features). The enhanced features corresponding to each sample video segment are then input into the feedforward neural network (FFN) to further extract the deep fused features of each enhanced feature. The above process is repeated (one... (And the corresponding FFN processing steps) M times, which can obtain 3 enhanced video features. These 3 enhanced video features are then fused using a Multi-Layer Perceptron (MLP) to obtain the second video feature. The purpose of using multi-head self-attention is to reduce the feature dimensionality and increase the model's non-linearity.
[0195] Since the second video feature is determined based on two dimensions, it can be called the first multimodal feature. The first multimodal feature can represent the common features between the global text feature and the first video feature.
[0196] The working process of the above-mentioned encoder can be seen from the following formula:
[0197]
[0198] wherein, represents the first video feature, the first video feature including the features of a plurality of sample video segments in the sample video, is a global text feature, is a fusion function of fusing the first video feature and the global text feature, and N is the number of sample video segments, which can be 128 in this example; is the position information feature of the plurality of sample video segments in the sample video, is the superimposed video feature corresponding to the i-th encoder branch, represents the deep fusion feature (enhanced feature) of the superimposed video feature corresponding to the i-th encoder branch after further extraction by the multi-head self-attention module, represents the second video feature, and in this example, the encoder has three branches, i.e., each and FFN correspond to a branch, so i = 1, 2, 3.
[0199] In this example, the encoder has three branches, and the global text feature includes the phrase features of three phrases, so each branch in the encoder can fuse one phrase feature in the global text feature and the first video feature to obtain a preliminary fusion feature corresponding to each phrase feature, and then further extract and fuse the preliminary fusion features through the three branches to obtain the second video feature.
[0200] In this example, the global text feature is fused into the first video feature to obtain the second video feature, which can specifically include: aligning the phrase features in the global text feature with the features of the video segments in the dimension by copying, and then fusing the global text feature into the first video feature by using the Hadamard Product algorithm (element-level multiplication) to obtain the second video feature.
[0201] For details, see the processing flow diagram of the global text feature and the first video feature in the encoder shown in Figure 4 , which is consistent with the processing process in Figure 3 , and will not be described here.
[0202] At the decoder end, the local text feature (local text feature) can be based on Figure 4The second video features (corresponding to the text at the decoder end and the second video features output from the encoder end) are used to determine the target video segment that matches the sample text from the sample video. That is, the second video features are further processed based on the local text features to obtain video features that accurately represent the predicted video segment in the sample video that corresponds to the sample text.
[0203] Specifically, the decoder consists of two branches: a video branch and a text branch, each consisting of a self-attention module (as shown in the diagram). ) and a collaborative attention module (shown in the figure) It consists of a self-attention module, which can be a multi-head self-attention module, and a collaborative attention module, which can be a multi-head collaborative attention module.
[0204] Similarly, considering that attention mechanisms cannot represent the temporal relationships between characters, positional information features corresponding to each character (i.e., positional features corresponding to the positional information of each character in the sample text) can be added to the local text features. Figure 3 The PE location code shown, and Figure 4 The PE corresponding to the decoder.
[0205] First, the self-attention module of the video branch is used to extract deep features from the second video features to obtain deep video features. Then, the self-attention module of the text branch is used to extract deep features from the local text features to obtain deep text features.
[0206] In the video branch, the input to the branch is the second video feature output by the encoder, and the local text features of the sample text (feature vectors of each character); the function of this branch is to guide the local text features to pay attention to important information in the local text features through the second video features.
[0207] Specifically, deep video features and deep text features are input into the collaborative attention module, which adjusts the deep text features. Specifically, the correlation features between deep video features and deep text features are first determined, and the deep text features are adjusted based on the correlation features and deep video features to obtain the adjusted local features.
[0208] In the text branch, the input consists of the second video feature output by the encoder and the global text feature of the sample text. The purpose of this branch is to guide the second video feature to focus on important information within the second video feature using local text features. Specifically, the deep text feature and deep video feature are input into the collaborative attention module to adjust the deep video feature. This involves first determining the correlation features between the deep video feature and the deep text feature, and then adjusting the deep video feature based on these correlation features and the deep text feature to obtain the adjusted video feature.
[0209] In this example, the self-attention module in the decoder can be a multi-head self-attention module. For example, a multi-head self-attention module can be used to extract depth features from the second video features to obtain depth video features. The specific implementation process can be found in [link to relevant documentation]. Figure 4 The decoder processing shown is detailed in the same way as the encoder process described above, where the first fused feature is processed using a multi-head self-attention module. Therefore, it will not be repeated here. It is understood that extracting deep text features from local text features using a self-attention module to obtain deep text features is also consistent with the process described above, where the first fused feature is processed using a multi-head self-attention module. Therefore, it will not be repeated here.
[0210] The specific implementation process of the decoder can be found in the following formula:
[0211]
[0212] in, Encode the positional information of each character in the sample text. For local text features; For self-attention modules, The deep text features are the local text features processed by the self-attention module. As a second video feature, The deep video features are the result of processing the second video features using a self-attention module. For collaborative attention modules, For the adjusted local features, These are the adjusted video features.
[0213] After obtaining the outputs of the two branches, namely the adjusted local features and the adjusted video features, the information gate can be used ( Figure 3 As shown in the IG), based on the guidance information determined by the adjusted local features, features that are semantically irrelevant to the adjusted local features are filtered out from the adjusted video features, further ensuring the accuracy of the multimodal video features corresponding to the sample text in the determined sample video.
[0214] Specifically, one can first adjust the local features. By integrating (aggregating) information, guidance information can be obtained. Then, based on the guidance information... The features of each video segment in the adjusted video features are then adjusted to obtain the third video features.
[0215] Among them, the adjusted local features ( The information is then integrated to obtain guidance. It is possible accomplish.
[0216] For details, please refer to the following formula:
[0217]
[0218] in, For the adjusted local features, For guidance purposes, and For parameters, For the adjusted video features, This is the third video feature (the multimodal video feature corresponding to the sample video). This is the Sigmoid activation function.
[0219] After obtaining the multimodal video features, for each sample, the multimodal video features are input into the video segment determination network (…). Figure 3 The temporal localization module shown in the figure obtains the weights corresponding to each sample video segment in the multimodal video features, and based on the multimodal video features and the weights corresponding to each sample video segment, obtains the predicted position information of the predicted video segment corresponding to the sample text in the sample video, that is, the start time and end time of the predicted video segment in the sample video.
[0220] Specifically, based on the correlation between the features of each sample video segment contained in the third video features, the weight 'a' corresponding to each video segment is determined. Based on the weights corresponding to each sample video segment, the third video features of each sample video segment are weighted to obtain the fourth video features. Based on the fourth video feature, the predicted location information of the predicted video segment in the sample video is determined. The predicted location information includes the prediction start time and prediction end time.
[0221] The predicted location information for a given sample can be determined using the following formula:
[0222]
[0223] in, Let N be the weight corresponding to the i-th video segment in each sample, where N is the number of video segments in the sample, and i is greater than or equal to 1 and less than or equal to N. The multimodal video feature (third video feature) corresponding to the i-th video segment in each sample. As the fourth video feature, To determine the network for the pre-trained video clips, To predict location information, To predict the start time, To predict the end time.
[0224] wherein the third video features of each sample video segment are weighted to obtain fourth video features , and specifically comprising: weighted summing the third video features of each sample video segment in the feature dimension to obtain the fourth video features . As an example, for example, a video includes 128 video segments, and the dimension of each video segment is 512 dimensions, then the dimension of the video can be expressed as: 128 512; by calculating the weight (attention value) corresponding to each video segment, 128 attention values can be obtained, and by copying the 128 attention values along the 512 dimensions, the 128 512-dimensional video features can be aggregated to obtain a 512-dimensional feature.
[0225] After obtaining the predicted position information corresponding to the predicted video segment of the sample, the second loss value can be determined based on the predicted position information corresponding to each sample and the position label, and for a sample, the second loss value represents the difference between the predicted position information corresponding to the sample and the position information corresponding to the position label of the sample.
[0226] For the second loss value corresponding to a sample, refer to the following formula:
[0227]
[0228]
[0229] wherein, is the position deviation value corresponding to the i-th sample in the training data, i.e., the difference between the predicted position information corresponding to the sample and the position information corresponding to the position label, is a threshold parameter, represents the absolute value of the position deviation value, is a loss function corresponding to the position deviation value, is a least square error loss function (L2) corresponding to the position deviation value, is the second loss value corresponding to any sample in the training data, is the starting loss corresponding to the i-th sample, is the ending loss corresponding to the i-th sample, is the second loss value corresponding to the i-th sample, represents the number of samples in the training data.
[0230] When , i.e., when the absolute value of the position deviation value is less than the threshold parameter, the second loss value adopts the L2 loss function corresponding to the position deviation value determines the second loss value.
[0231] When the absolute value of the position bias value is not less than the threshold parameter, i.e., t determines the second loss value.
[0232] Based on the weight corresponding to each sample video segment in the multi-modal video feature corresponding to each sample and the position label corresponding to each sample, a third loss value is determined. For a sample, the third loss value represents the possibility that each sample video segment in the sample is a target video segment.
[0233] Optionally, the weight corresponding to each sample video segment can be represented by an attention mask corresponding to each sample video segment.
[0234] The third loss value corresponding to a sample can be represented as:
[0235]
[0236] wherein, is the third loss value, represents the weight (attention mask) corresponding to each sample video segment in each sample corresponding to the multi-modal video feature, which can be represented by a probability. The greater the probability, the greater the possibility that the sample video segment corresponding to the weight is a target video segment, represents the number of sample video segments, is an indicator function. If is within the interval corresponding to the position label, then is 1, otherwise 0. When is 1, it means that the greater the possibility that the sample video segment corresponding to is a target video segment, and when is 0, it means that the smaller the possibility that the sample video segment corresponding to is a target video segment. Within the interval corresponding to the position label, the greater the better.
[0237] As an example, for example, the target video segment corresponding to the position label is the 2nd to 4th second of the sample video, and the interval corresponding to the position label is the 2nd to 4th second. Based on the weight (attention mask) corresponding to each sample video segment of the sample video, when the weight is within the interval corresponding to the position label (the interval corresponding to the 2nd to 4th second), the weight is 1, and the greater the possibility that the sample video segment corresponding to the weight is a target video segment, otherwise, when the weight is not within the interval corresponding to the position label (the interval corresponding to the 2nd to 4th second), is 0, the smaller the possibility that the sample video clip corresponding to the weight is the target video clip.
[0238] Based on the first loss value, the second loss value and the third loss value, a value of a training loss function corresponding to the neural network model is determined, if the training loss function converges, the model corresponding to the convergence is taken as the final neural network model, if not, the model parameters of the neural network model are adjusted, and the neural network model is trained based on the training data.
[0239] In this example, after obtaining the neural network model, tests can be performed on the Charades-STA, ActivityNet Captions and TACoS three data sets. Corresponding to the above three data sets, optionally, the number of video clips of the sample video corresponding to the Charades-STA data set is 128, the number of video clips of the sample video corresponding to the ActivityNet Captions data set is 128, and the number of video clips of the sample video corresponding to the TACoS data set is 200.
[0240] Optionally, the maximum length of the sample text that can be processed by the neural network model corresponding to the Charades-STA data set at a time is 10, the maximum length of the sample text that can be processed by the neural network model corresponding to the ActivityNet Captions data set at a time is 25, and the maximum length of the sample text that can be processed by the neural network model corresponding to the TACoS data set at a time is 25.
[0241] Optionally, the threshold parameter corresponding to the Charades-STA data set is 0.1, the threshold parameter corresponding to the ActivityNet Captions data set is 0.4, and the threshold parameter corresponding to the TACoS data set is 0.2.
[0242] In this example, the number of samples that can be processed by the neural network model at a time is 100.
[0243] Optionally, the neural network model can select the Adam optimizer to improve the training speed of the model.
[0244] Optionally, the initial learning rate of the neural network model is 1e-3.
[0245] The application achieves the best performance on three datasets of Charades-STA, ActivityNet Captions and TACoS, where R1@m and mIoU are used as evaluation indexes, where R1@m represents the accuracy of IoU exceeding m in the Top1 recall, and the higher the value, the better the performance; mIoU represents the average IoU value of Top1 recall, and the higher the value, the better the performance. The specific experimental results are shown in the following table:
[0246] Table 1 Performance comparison of Charades-STA dataset
[0247]
[0248] Among them, Ours represents the method of the application, as shown in Table 1, through the method of the application, the evaluation indexes R1@0.3, R1@0.5, R1@0.7 and mIoU obtained are 72.53, 59.84, 37.74 and 51.45 respectively, based on the LGI[5] method, the evaluation indexes R1@0.3, R1@0.5, R1@0.7 and mIoU obtained are 72.96, 59.46, 35.48 and 51.38 respectively, comparison can know, compared with the evaluation indexes corresponding to other methods, the performance of the method of the application is higher than that of other methods.
[0249] Table 2 Performance comparison of ActivityNet Captions dataset
[0250]
[0251] Among them, Ours represents the method of the application, as shown in Table 2, through the method of the application, the evaluation indexes R1@0.3, R1@0.5, R1@0.7 and mIoU obtained are 60.26, 42.46, 24.09 and 42.51 respectively, based on the LGI[5] method, the evaluation indexes R1@0.3, R1@0.5, R1@0.7 and mIoU obtained are 58.52, 41.51, 23.07 and 41.13 respectively, comparison can know, compared with the evaluation indexes corresponding to other methods, the performance of the method of the application is higher than that of other methods.
[0252] Table 3 Performance comparison of TACoS dataset
[0253]
[0254] Wherein, Ours represents the method of the present application, as can be seen from Table 3, through the method of the present application, the evaluation index R1@0.3 obtained is 60.08, R1@0.5 is 45.81, R1@0.7 is 31.12, mIoU is 13.87, based on the 2D-TAN[1] method, the evaluation index R1@0.3 obtained is 47.59, R1@0.5 is 37.29, R1@0.7 is 25.32, and comparison shows that the performance of the method of the present application is higher than that of other methods.
[0255] In an embodiment of the present application, the method comprises:
[0256] The method further comprises:
[0257] The method further comprises:
[0258] The method further comprises:
[0259] If the target segment exists in the to-be-processed video, the video segment is sent to the user.
[0260] Wherein, the scheme for determining the target segment in the to-be-processed video that matches the to-be-processed text can be applied to any scenario that needs to determine the target segment, for example, a scenario of searching a video segment based on text.
[0261] Wherein, the search text indicates the relevant information of the video segment that the user wants to search, for example, the search text is: slam dunk, which means that the user wants to search for a video segment about slam dunk. Based on the manner described in the foregoing, the target segment corresponding to the search text can be determined from the video database by taking the search text as the to-be-processed text and any video in the video database as the to-be-processed video.
[0262] Wherein, the video search request can be initiated by the user based on the user's terminal device, and the terminal device can include at least one of the following: a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a smart television, and a smart vehicle device.
[0263] The video segment can be displayed through the searcher's terminal device, wherein the terminal device can run a client that provides a video display function, the client provides a video display function, and the specific form of the client is not limited, for example: a media player, a browser, etc., the client can be in the form of an application program or a web page, which is not limited herein.
[0264] In an embodiment of the present application, the method comprises:
[0265] The title information of the to-be-processed video is obtained, and the to-be-processed text is the title information of the to-be-processed video.
[0266] The method further includes:
[0267] If the target segment exists in the to-be-processed video, it is determined that the title information matches the to-be-processed video.
[0268] If the target segment does not exist in the to-be-processed video, it is determined that the title information does not match the to-be-processed video.
[0269] In another application scenario, for example, the to-be-processed text is the title information of the first video, and the to-be-processed video is the first video. In order to determine whether the title information is consistent with the video content of the first video, the scheme of the present application can also be used to determine whether the target segment corresponding to the title information exists in the first video. If it exists, it indicates that the title information is consistent with the video content of the first video (matching). If it does not exist, it indicates that the title information is not consistent with the video content of the first video (not matching).
[0270] Figure 5 A schematic diagram of an implementation environment of a media data processing method provided by an embodiment of the present application. The implementation environment in this example can include but is not limited to a search server 101, a network 102, and a terminal device 103. The search server 101 can communicate with the terminal device 103 through the network 102, send the received video search request to the search server 101, and the search server 101 can send the retrieved target image to the terminal device 103 through the network.
[0271] The terminal device 103 includes a human-computer interaction screen 1031, a processor 1032, and a memory 1033. The human-computer interaction screen 1031 is used to display the target image. The memory 1033 is used to store the retrieved image and the target image and other related data. The search server 101 includes a database 1011 and a processing engine 1012, and the processing engine 1012 can be used to train a neural network model. The database 1011 is used to store the trained neural network model and a video database. The terminal device 103 can upload the video search request to the search server 101 through the network. The processing engine 1012 in the search server 101 can obtain the video database corresponding to the video search request, determine the video segment corresponding to the search text from the video database according to the global text features of the search text and the local text features of each unit text contained in the search text, and the video features of the to-be-processed video, and provide the video segment to the terminal device 103 of the searcher for display.
[0272] The processing engine in the search server 101 has two main functions. The first function is to train a neural network model. The second function is to process a video search request based on the neural network model and a video database to obtain a video segment in the video database corresponding to the video search request and the search text (search function). It can be understood that the two functions can be implemented by two servers respectively, as shown in Figure 6 The two servers are a training server 201 and a search server 202. The training server 201 is used to train a neural network model. The search server 202 is used to implement the search function. The video database is stored in the search server 202.
[0273] In actual applications, the two servers can communicate with each other. After the training server 201 trains the neural network model, the neural network model can be stored in the training server 201 or sent to the search server 202. Alternatively, when the search server 202 needs to call the neural network model, a model calling request is sent to the training server 201. The training server 201 sends the neural network model to the search server 202 based on the request.
[0274] As an example, the terminal device 204 sends a video search request to the search server 202 through the network 203. The search server 202 calls the neural network model in the training server 201. Based on the neural network model, the search server 202 sends the video segment searched to the terminal device 204 through the network 203 after completing the search function, so that the terminal device 204 displays the video segment.
[0275] According to an optional solution of the present application, based on the video segment determined to match the text, a video segment recommendation can also be performed. For example, based on a search keyword (text) of a user, at least one video matching the search keyword is searched from a database and recommended to the user. The application of the video segment determined to match the text in the video is very extensive and will not be described here.
[0276] Based on the same principle as shown in Figure 1 The present application also provides a media data processing apparatus 20, as shown in Figure 7 The media data processing apparatus 20 can include a data acquisition module 210, a feature extraction module 220, a feature fusion module 230, and a target segment determination module 240, wherein:
[0277] The data acquisition module 210 is configured to acquire a to-be-processed text and a to-be-processed video.
[0278] The feature extraction module 220 is configured to extract global text features and local text features corresponding to the to-be-processed text and a first video feature of the to-be-processed video, wherein the global text features include phrase features corresponding to phrases included in the to-be-processed text, and the local text features include features corresponding to unit texts included in the to-be-processed text.
[0279] The feature fusion module 230 is configured to fuse the global text features into the first video feature to obtain a second video feature.
[0280] The target segment determination module 240 is configured to determine, according to the local text features and the second video feature, a target segment matching the to-be-processed text from the to-be-processed video.
[0281] Optionally, the first video feature includes features of a plurality of video segments in the to-be-processed video; and when the feature fusion module 230 fuses the global text features into the first video feature to obtain the second video feature, the feature fusion module 230 is specifically configured to:
[0282] fuse the global text features and the first video feature to obtain a fused video feature;
[0283] obtain position information features of the video segments in the to-be-processed video;
[0284] superimpose the features of each video segment in the fused video feature and the respective position information features of each video segment to obtain superimposed video features;
[0285] obtain the second video feature based on the superimposed video features.
[0286] Optionally, when the target segment determination module determines, according to the local text features and the second video feature, the target segment matching the to-be-processed text from the to-be-processed video, the target segment determination module is specifically configured to:
[0287] determine associated features of the local text features and the second video feature;
[0288] adjust the local text features according to the associated features and the second video feature to obtain adjusted local features;
[0289] adjust the second video feature according to the associated features and the local text features to obtain adjusted video features;
[0290] determine, based on the adjusted local features and the adjusted video features, the target segment matching the to-be-processed text from the to-be-processed video.
[0291] Optionally, the target segment determining module, when determining the target segment matching the to-be-processed text from the to-be-processed video according to the third video features, is specifically configured to:
[0292] determine the guidance information of the to-be-processed text to the to-be-processed video according to the adjusted local features;
[0293] adjust the features of each video segment in the adjusted video features according to the guidance information, to obtain third video features;
[0294] determine the target segment matching the to-be-processed text from the to-be-processed video according to the third video features.
[0295] Optionally, the target segment determining module, when determining the target segment matching the to-be-processed text from the to-be-processed video according to the third video features, is specifically configured to:
[0296] determine the weight corresponding to each video segment according to the correlation between the features of each video segment contained in the third video features;
[0297] weight the third video features of each video segment based on the weight corresponding to each video segment, to obtain fourth video features;
[0298] determine the position information of the target segment matching the to-be-processed text in the to-be-processed video based on the fourth video features;
[0299] determine the target segment matching the to-be-processed text in the to-be-processed video based on the position information.
[0300] Optionally, the feature extracting module, when extracting the global text features and the local text features of the to-be-processed text, is specifically configured to:
[0301] obtain each unit text in the to-be-processed text and the positional relationship between each unit text;
[0302] determine the features of each unit text based on each unit text and the positional relationship between each unit text, and the local text features include the features of each unit text;
[0303] determine the phrase features of each phrase contained in the to-be-processed text based on the features corresponding to each unit text;
[0304] fuse the phrase features of each phrase to obtain the global text features.
[0305] Optionally, the first video features include the features of a plurality of video segments in the to-be-processed video; and the feature fusing module, when obtaining the second video features based on the superimposed video features, is specifically configured to:
[0306] determine the weight corresponding to each video segment based on the association relationship between each video segment in the superimposed video features;
[0307] For each video segment, based on the weight corresponding to the video segment and the feature corresponding to the video segment in the superimposed video features, obtain the enhanced feature corresponding to the video segment;
[0308] Based on the enhanced feature corresponding to each video segment, the second video feature is extracted.
[0309] Optionally, the global text feature of the to-be-processed text is extracted, the global text feature is fused into the first video feature to obtain the second video feature, and the target segment matched with the to-be-processed text is determined from the to-be-processed video based on the local text feature and the second video feature based on the neural network model;
[0310] The neural network model includes a phrase feature extraction network, a multi-modal feature extraction network, and a video segment determination network, and the neural network model is obtained through the following model training module:
[0311] The model training module is used to:
[0312] Obtain training data, the training data including a plurality of samples, each sample including a sample video and a sample text, and each sample carrying a position label representing the position information of the corresponding target video segment in the sample video and the sample text;
[0313] For each sample in the training data, extract the global text feature and the local text feature of the sample text in the sample, and the video feature of the sample video;
[0314] For the sample, input the feature of each unit text in the sample text into the phrase feature extraction network to obtain the predicted phrase feature of each phrase in the sample text;
[0315] Based on the matching degree between the predicted phrase features of each phrase corresponding to each sample, a first loss value is determined, and for a sample, the first loss value represents the semantic difference between the phrases in the sample;
[0316] For the sample, input the global text feature and the local text feature of the sample text, and the video feature of the sample video into the multi-modal feature extraction network to obtain the multi-modal video feature corresponding to the sample video;
[0317] For the sample, the multi-modal video feature is input to the video segment determination network to obtain a weight corresponding to each sample video segment in the multi-modal video feature, and based on the multi-modal video feature and the weight corresponding to each sample video segment, a corresponding predicted position information of a predicted video segment of the sample text in the sample video is obtained.
[0318] Based on the predicted position information corresponding to each sample and the position label of each sample, a second loss value is determined, which represents the difference between the predicted position information corresponding to each sample and the position label of each sample.
[0319] Based on the weight corresponding to each sample video segment in the multi-modal video feature corresponding to each sample and the position label corresponding to each sample, a third loss value is determined, and for a sample, the third loss value represents the possibility that each sample video segment in the sample is a target video segment.
[0320] Based on the first loss value, the second loss value and the third loss value, the value of the training loss function corresponding to the neural network model is determined.
[0321] If the training loss function converges, the model corresponding to the convergence is taken as the final neural network model, and if it does not converge, the model parameters of the neural network model are adjusted and the neural network model is trained based on the training data.
[0322] Optionally, for a sample, the model training module, when determining the second loss value based on the predicted position information corresponding to each sample and the position label, is specifically configured to:
[0323] Determine a position deviation value based on the predicted position information corresponding to the sample and the position label.
[0324] If the absolute value of the position deviation value is less than a threshold parameter, the second loss value is determined based on the least square error loss function corresponding to the position deviation value.
[0325] If the absolute value of the position deviation value is not less than the threshold parameter, the second loss value is determined based on the loss function corresponding to the position deviation value, and the loss function includes the mean absolute error loss function and the threshold parameter.
[0326] Optionally, the data acquisition module, when acquiring the to-be-processed text and the to-be-processed video, is specifically configured to:
[0327] Acquire a video search request of a user, and the search text in the video search request is the to-be-processed text.
[0328] Acquire a video database corresponding to the video search request, and the search text is the to-be-processed text, and any video in the video database is the to-be-processed video.
[0329] The device further comprises:
[0330] The first video processing module is configured to send the video clip to the user when the target clip exists in the to-be-processed video.
[0331] Optionally, the data acquisition module is configured to:
[0332] acquire the to-be-processed video and the title information of the to-be-processed video, and the to-be-processed text is the title information of the to-be-processed video;
[0333] The apparatus further includes:
[0334] The second video processing module is configured to determine that the title information matches the to-be-processed video when the target clip exists in the to-be-processed video, and determine that the title information does not match the to-be-processed video when the target clip does not exist in the to-be-processed video.
[0335] The media data processing apparatus provided in the embodiments of the present application can execute the media data processing method provided in the embodiments of the present application, and the implementation principles are similar. The actions performed by each module and unit in the media data processing apparatus in the embodiments of the present application are corresponding to the steps in the media data processing method in the embodiments of the present application. The detailed function description of each module of the media data processing apparatus can be found in the description of the corresponding media data processing method provided in the foregoing description, and will not be repeated here.
[0336] The media data processing apparatus can be a computer program (including program code) running in a computer device, for example, the media data processing apparatus is an application software. The apparatus can be used to execute the corresponding steps in the method provided in the embodiments of the present application.
[0337] In some embodiments, the media data processing apparatus provided in the embodiments of the present application can be realized in a combination of software and hardware. For example, the media data processing apparatus provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the media data processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can be one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic elements.
[0338] In some embodiments, the media data processing apparatus provided by the embodiments of the present application can be implemented in software, Figure 7 The media data processing apparatus stored in the memory can be software in the form of programs, plug-ins, etc., and includes a series of modules, including a data acquisition module 210, a feature extraction module 220, a feature fusion module 230, and a target segment determination module 240, for implementing the media data processing method provided by the embodiments of the present application.
[0339] The modules described in the embodiments of the present application can be implemented in software or hardware. In some cases, the names of the modules do not limit the modules themselves.
[0340] Based on the same principles as the methods shown in the embodiments of the present application, the embodiments of the present application also provide an electronic device, which can include but is not limited to a processor and a memory; the memory is used to store a computer program; the processor is used to execute the media data processing method shown in any embodiment of the present application by calling the computer program.
[0341] The media data processing method provided by the present application can first perform preliminary processing on the first video feature based on the global text feature of the to-be-processed text and the first video feature of the to-be-processed video to obtain a second video feature, and then determine the target segment matching the to-be-processed text from each video segment in the to-be-processed video based on the local text feature of the to-be-processed text and the second video feature. Since the global text feature and the local text feature can describe all the information of the to-be-processed text from different granularities, the target segment determined by the text features of different granularities (global text feature and local text feature) in the present application is more matched with the to-be-processed text and has a closer semantic relationship.
[0342] In an optional embodiment, an electronic device is provided, as shown in Figure 8 The electronic device 4000 shown in Figure 8 The electronic device 4000 shown in
[0343] The processor 4001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in connection with the disclosure. The processor 4001 can also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0344] The bus 4002 can include a path for transmitting information between the above-mentioned components. The bus 4002 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, Figure 8 Only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0345] The memory 4003 can be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, an optical disk storage (including a compact disk, a laser disk, an optical disk, a digital versatile disk, a Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer, but not limited to this.
[0346] The memory 4003 is configured to store application code (computer program) for implementing the scheme of the present application, and the processor 4001 is configured to control the execution. The processor 4001 is configured to execute the application code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.
[0347] The electronic device can also be a terminal device, Figure 8 The electronic device shown is only an example, and should not impose any limitation on the functions and use range of the embodiments of the present application.
[0348] The computer readable storage medium provided by the embodiments of the present application has the computer program stored thereon, and when the computer program is run on the computer, the computer can execute the corresponding content in the foregoing method embodiments.
[0349] According to another aspect of the present application, a computer program product or computer program is also provided, which includes computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the media data processing method provided in the various embodiment implementation manners.
[0350] The computer program code for performing the operations of the present application can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as "C" or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0351] It should be understood that the flow diagrams and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of various embodiments of the present application. In this regard, each block in the flow diagrams and block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each of the blocks of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0352] The computer readable storage medium of embodiments of the present application may, for example, be— but is not limited to— an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present application, the computer readable storage medium can be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0353] The computer readable storage medium described above can bear one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.
[0354] The above description is merely illustrative of the application and the application of the principles of the application Thus, the above description should not be construed as limiting the scope of the application, which is defined by the appended claims.
Claims
1. A method of processing media data, the method comprising: The method comprises: acquiring a to-be-processed text and a to-be-processed video; extracting global text features corresponding to the to-be-processed text and local text features, and first video features of the to-be-processed video, wherein the global text features comprise phrase features corresponding to each phrase included in the to-be-processed text, the local text features comprise features corresponding to each unit text included in the to-be-processed text, and the first video features comprise features of multiple video clips in the to-be-processed video; fusing the global text features into the first video features to obtain second video features; determining associated features of the local text features and the second video features; acquiring first weights of each unit text in the local text features; adjusting the first weights of each unit text in the local text features according to the associated features and the second video features; and obtaining adjusted local features based on the features of each unit text and the adjusted weights corresponding to each unit text; adjusting second weights of each video clip in the to-be-processed video according to the associated features and the second video features; and obtaining adjusted video features based on the features of each video clip in the second video features and the adjusted weights corresponding to each video clip. determining target clips matching the to-be-processed text from the to-be-processed video based on the adjusted local features and the adjusted video features.
2. The method of claim 1, wherein, The method further comprises: fusing the global text features and the first video features to obtain fused video features; acquiring position information features corresponding to each video clip in the to-be-processed video; superimposing the features of each video clip in the fused video features and the position information features corresponding to each video clip to obtain superimposed video features; and obtaining the second video features based on the superimposed video features.
3. The method of claim 1, wherein, The method further comprises: fusing the adjusted local features to obtain guidance information of the to-be-processed text on the to-be-processed video; adjusting the features of each video clip in the adjusted video features according to the guidance information to obtain third video features; and determining target clips matching the to-be-processed text from the to-be-processed video based on the third video features.
4. The method of claim 3, wherein, The method further comprises: determining weights corresponding to each video clip according to the associated relationship between the features of each video clip included in the third video features; weighting the third video features of each video clip based on the weights corresponding to each video clip to obtain fourth video features; determining position information of target clips matching the to-be-processed text in the to-be-processed video based on the fourth video features. determine a target segment matching the to-be-processed text from the to-be-processed video based on the position information.
5. The method according to any one of claims 1 to 4, characterized in that, The extracting the global text feature and the local text feature of the to-be-processed text comprises: obtaining each unit text in the to-be-processed text and a position relationship between each unit text; determining a feature of each unit text based on each unit text and the position relationship between each unit text, wherein the local text feature comprises the feature of each unit text; determining a phrase feature of each phrase contained in the to-be-processed text based on the feature corresponding to each unit text; fusing the phrase feature of each phrase to obtain the global text feature.
6. The method of claim 2, wherein, The obtaining the second video feature based on the superimposed video feature comprises: determining a weight corresponding to each video segment based on the association relationship between each video segment in the superimposed video feature; for each video segment, obtaining an enhanced feature corresponding to the video segment based on the weight corresponding to the video segment and the feature corresponding to the video segment in the superimposed video feature; and extracting the second video feature based on the enhanced feature corresponding to each video segment.
7. The method according to any one of claims 1 to 4, characterized in that, The extracting the global text feature of the to-be-processed text, fusing the global text feature into the first video feature to obtain a second video feature, and determining a target segment matching the to-be-processed text from the to-be-processed video based on the local text feature and the second video feature are obtained based on a neural network model; The neural network model comprises a phrase feature extraction network, a multi-modal feature extraction network, and a video segment determination network, and the neural network model is obtained by training in the following manner: obtaining training data, wherein the training data comprises a plurality of samples, each sample comprises a sample video and a sample text, and each sample carries a position label representing position information of a corresponding target video segment in the sample video and the sample text; for each sample in the training data, extracting a global text feature and a local text feature of the sample text in the sample and a video feature of the sample video; for the sample, inputting a feature of each unit text in the sample text into the phrase feature extraction network to obtain a predicted phrase feature of each phrase in the sample text; determining a first loss value based on a matching degree between the predicted phrase feature of each phrase corresponding to each sample, wherein the first loss value represents a semantic difference between each phrase in the sample; for the sample, inputting the global text feature and the local text feature of the sample text and the video feature of the sample video into the multi-modal feature extraction network to obtain a multi-modal video feature corresponding to the sample video; For the sample, the multi-modal video feature is input into the video segment determination network to obtain a weight corresponding to each sample video segment in the multi-modal video feature, and based on the multi-modal video feature and the weight corresponding to each sample video segment, a predicted position information of a predicted video segment corresponding to the sample text in the sample video is obtained; Based on the predicted position information corresponding to each sample and the position label of each sample, a second loss value is determined, which represents the difference between the predicted position information corresponding to each sample and the position label of each sample; Based on the weight of each sample video segment in the multi-modal video feature corresponding to each sample and the position label of each sample, a third loss value is determined, which represents the possibility of each sample video segment in the sample being the target video segment; Based on the first loss value, the second loss value and the third loss value, the value of the training loss function corresponding to the neural network model is determined; If the training loss function converges, the model corresponding to the convergence is taken as the final neural network model, and if it does not converge, the model parameters of the neural network model are adjusted and the neural network model is trained based on the training data.
8. The method of claim 7, wherein, For a sample, the second loss value is determined based on the predicted position information corresponding to each sample and the position label, including: Based on the predicted position information corresponding to each sample and the position label, a position deviation value is determined; If the absolute value of the position deviation value is less than a threshold parameter, the second loss value is determined based on the least square error loss function corresponding to the position deviation value; If the absolute value of the position deviation value is not less than the threshold parameter, the second loss value is determined based on the loss function corresponding to the position deviation value, and the loss function includes the mean absolute error loss function and the threshold parameter.
9. The method according to any one of claims 1 to 4, characterized in that, The method further comprises: If the target segment exists in the processed video, the target segment is sent to the user. The method further comprises: If the target segment exists in the processed video, it is determined that the title information matches the processed video; If the target segment does not exist in the processed video, it is determined that the title information does not match the processed video.
10. The method according to any one of claims 1 to 4, characterized in that, The method further comprises: The data acquisition module is configured to acquire the processed text and the processed video. 11. A media data processing apparatus, characterized by comprising: The feature extraction module is used to extract global text features and local text features corresponding to the text to be processed, as well as the first video features of the video to be processed. The global text features include phrase features corresponding to each phrase contained in the text to be processed, and the local text features include features corresponding to each unit text contained in the text to be processed. The first video features include features of multiple video segments in the video to be processed. The feature fusion module is used to fuse the global text features into the first video features to obtain the second video features; The target segment determination module is used to determine the correlation features between the local text features and the second video features; Obtain the first weight of each unit text in the local text features; adjust the first weight of each unit text in the local text features according to the association features and the second video features; and obtain the adjusted local features based on the features of each unit text and the adjusted weights corresponding to each unit text. Based on the second weight of each video segment in the video to be processed; based on the association feature and the second video feature, the second weight of each video segment in the video to be processed is adjusted; based on the features of each video segment in the second video feature and the adjusted weight corresponding to each video segment, the adjusted video feature is obtained. Based on the adjusted local features and the adjusted video features, a target segment matching the text to be processed is determined from the video to be processed.
12. The apparatus of claim 11, wherein, When the feature fusion module fuses the global text features into the first video features to obtain the second video features, it is specifically used for: The global text features and the first video features are fused to obtain fused video features; Obtain the location information features of each video segment in the video to be processed; The features of each video segment and the corresponding location information features of each video segment in the fused video features are superimposed to obtain the superimposed video features. The second video feature is obtained based on the superimposed video features.
13. The apparatus of claim 11, wherein, When the target segment determination module determines the target segment matching the text to be processed from the video to be processed based on the adjusted local features and the adjusted video features, it is specifically used for: The adjusted local features are fused to obtain the guidance information of the text to be processed for the video to be processed; Based on the guidance information, the features of each video segment in the adjusted video features are adjusted to obtain the third video features; Based on the third video feature, a target segment matching the text to be processed is determined from the video to be processed.
14. The apparatus of claim 13, wherein, When the target segment determination module determines the target segment matching the text to be processed from the video to be processed based on the third video feature, it is specifically used for: Based on the correlation between the features of each video segment contained in the third video feature, the weight corresponding to each video segment is determined; Based on the weights corresponding to each video segment, the third video features of each video segment are weighted to obtain the fourth video features; determine position information of a target segment in the to-be-processed video that matches the to-be-processed text based on the fourth video feature; determine the target segment in the to-be-processed video that matches the to-be-processed text based on the position information.
15. The apparatus of any one of claims 11 to 14, wherein, In extracting the global text feature and the local text feature of the to-be-processed text, the feature extraction module is specifically configured to: obtain each unit text in the to-be-processed text and a positional relationship between each unit text; determine a feature of each unit text based on each unit text and the positional relationship between each unit text, the local text feature including the feature of each unit text; determine a phrase feature of each phrase included in the to-be-processed text based on the feature corresponding to each unit text; fuse the phrase features of each phrase to obtain the global text feature.
16. The apparatus of claim 12, wherein, In obtaining the second video feature based on the superimposed video feature, the feature fusion module is specifically configured to: determine a weight corresponding to each video segment based on the association relationship between each video segment in the superimposed video feature; for each video segment, obtain an enhanced feature corresponding to the video segment based on the weight corresponding to the video segment and the feature corresponding to the video segment in the superimposed video feature; extract the second video feature based on the enhanced features corresponding to each video segment.
17. The apparatus of any one of claims 11 to 14, wherein, extracting the global text feature of the to-be-processed text, fusing the global text feature into the first video feature to obtain a second video feature, and determining the target segment in the to-be-processed video that matches the to-be-processed text based on the local text feature and the second video feature are obtained based on a neural network model; the neural network model includes a phrase feature extraction network, a multi-modal feature extraction network, and a video segment determination network, and the neural network model is obtained by a model training module; the model training module is configured to: obtain training data, the training data including a plurality of samples, each sample including a sample video and a sample text, and each sample carrying a position label, the position label representing position information of a corresponding target video segment in the sample video that matches the sample text; for each sample in the training data, extract a global text feature and a local text feature of the sample text in the sample, and extract a video feature of the sample video; for the sample, input the features of each unit text in the sample text into the phrase feature extraction network to obtain predicted phrase features of each phrase in the sample text; determine a first loss value based on matching degrees between the predicted phrase features of each phrase corresponding to each sample, the first loss value representing semantic differences between the phrases in the sample for one sample; for the sample, input the global text feature and the local text feature of the sample text and the video feature of the sample video into the multi-modal feature extraction network to obtain a multi-modal video feature corresponding to the sample video; The multi-modal video feature is input into the video segment determination network to obtain a weight corresponding to each sample video segment in the multi-modal video feature, and based on the multi-modal video feature and the weight corresponding to each sample video segment, a predicted position information of a predicted video segment corresponding to the sample text in the sample video is obtained. A second loss value is determined based on the predicted position information corresponding to each sample and the position label of each sample, and the second loss value represents a difference between the predicted position information corresponding to each sample and the position label of each sample. A third loss value is determined based on the weight corresponding to each sample video segment in the multi-modal video feature corresponding to each sample and the position label corresponding to each sample, and for one sample, the third loss value represents a possibility that each sample video segment in the sample is the target video segment. Based on the first loss value, the second loss value and the third loss value, a value of a training loss function corresponding to the neural network model is determined. If the training loss function converges, the model corresponding to the convergence is taken as the final neural network model, and if the training loss function does not converge, the model parameters of the neural network model are adjusted and the neural network model is trained based on the training data.
18. The apparatus of claim 17, wherein, For one sample, when the model training module determines the second loss value based on the predicted position information corresponding to each sample and the position label of each sample, it is specifically used for: Based on the predicted position information corresponding to each sample and the position label, a position deviation value is determined. If the absolute value of the position deviation value is less than a threshold parameter, the second loss value is determined based on a least square error loss function corresponding to the position deviation value. If the absolute value of the position deviation value is not less than the threshold parameter, the second loss value is determined based on a loss function corresponding to the position deviation value, and the loss function includes a mean absolute error loss function and the threshold parameter.
19. The apparatus of any one of claims 11-14, wherein, The data acquisition module, when acquiring the to-be-processed text and the to-be-processed video, is specifically used for: Acquiring a video search request of a user, wherein the video search request includes a search text; Acquiring a video database corresponding to the video search request, wherein the search text is the to-be-processed text, and any video in the video database is the to-be-processed video. The device further comprises: A first video processing module configured to, if the target segment exists in the to-be-processed video, send the target segment to the user.
20. The apparatus of any one of claims 11-14, wherein, The data acquisition module, when acquiring the to-be-processed text and the to-be-processed video, is specifically used for: Acquiring a to-be-processed video and title information of the to-be-processed video, wherein the to-be-processed text is the title information of the to-be-processed video. The device further comprises: A second video processing module configured to, if the target segment exists in the to-be-processed video, determine that the title information and the to-be-processed video are matched; and if the target segment does not exist in the to-be-processed video, determine that the title information and the to-be-processed video are not matched.
21. An electronic device, comprising: A computer program product comprising a computer readable storage medium having stored thereon computer program means, the computer program means comprising computer program instructions executable by a processor to cause the processor to carry out the method of any of claims 1-10.
22. A computer-readable storage medium, characterized in that, A computer program product comprising a computer readable storage medium having stored thereon computer program means, the computer program means comprising computer program instructions executable by a processor to cause the processor to carry out the method of any of claims 1-10.