Video processing method and device
By performing multi-level feature extraction and cross-modal fusion on the positioning text and target video, and utilizing capsule routing network and pyramid network, the problem of semantic misalignment between video and text is solved, and the accuracy of temporal text positioning is improved.
Patent Information
- Application Number
- CN202110217613.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-26
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2041-02-26
AI Technical Summary
Existing technologies suffer from semantic misalignment when fusing a video with a given sentence, resulting in inaccurate temporal text localization.
The temporal text localization model is used to extract multi-level features of the localization text and the target video respectively, the capsule routing network is used for cross-modal multi-level fusion, and the pyramid network is used to obtain multi-level global features to achieve accurate localization of the video clip.
It improves the accuracy of temporal text positioning, solves the problem of video and text misalignment, and achieves better representation of the relationship between text and video.
Smart Images

Figure CN115049950B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing, in particular to a video processing method and device. BACKGROUND
[0002] The purpose of temporal text positioning is to locate the corresponding time sequence segment of the given sentence description in the unpruned long video. Since it has a wide range of applications in video understanding, video retrieval and human-computer interaction, it has attracted more and more attention from the industry and academia.
[0003] The key to solving the temporal text positioning task is how to extract the better relationship between the given text and the video. The existing scheme will have the problem of semantic misalignment between the video and the text when fusing the video and the given sentence, which will further lead to inaccurate temporal text positioning.
[0004] In view of the above problems, no effective solution has been proposed so far. SUMMARY
[0005] The embodiments of the present application provide a video processing method and device to at least solve the technical problem of low accuracy in temporal text positioning in the prior art.
[0006] According to an aspect of the embodiments of the present application, a video processing method is provided, comprising: receiving a video positioning instruction, wherein the video positioning instruction at least includes a positioning text used to describe a video segment in a target video; positioning in the target video based on the positioning text by a temporal text positioning model to obtain a positioning result, wherein the positioning result includes at least one target video segment associated with the positioning text; wherein the temporal text positioning model performs cross-modal multi-level fusion on multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video to obtain multi-level global feature, and positions the target video according to the multi-level global feature to obtain at least one target video segment associated with the positioning text.
[0007] According to another aspect of the embodiments of the present application, a video processing method is also provided, comprising: performing multi-level feature extraction on the positioning text and the target video respectively to obtain multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video; performing feature fusion based on the multi-level first feature information and the multi-level second feature information by a capsule routing network to generate multi-level fusion feature, wherein the multi-level fusion feature includes consistency information of the multi-level first feature information and the multi-level second feature information; obtaining multi-level global feature based on the multi-level fusion feature by a pyramid network; positioning the target video according to the multi-level global feature to obtain at least one target video segment associated with the positioning text.
[0008] According to another aspect of the embodiments of the present application, a video processing apparatus is also provided, which comprises: a receiving module configured to receive a video positioning instruction, wherein the video positioning instruction comprises at least a positioning text used for describing a video segment in a target video; and a positioning module configured to perform positioning in the target video based on the positioning text by using a time-series text positioning model to obtain a positioning result, wherein the positioning result comprises at least one target video segment associated with the positioning text; wherein the time-series text positioning model performs cross-modal multi-level fusion on multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video to obtain multi-level global feature, and performs positioning on the target video based on the multi-level global feature to obtain at least one target video segment associated with the positioning text.
[0009] According to another aspect of the embodiments of the present application, a video processing apparatus is also provided, which comprises: a first extracting module configured to perform multi-level feature extraction on the positioning text and the target video respectively to obtain multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video; a fusion module configured to perform feature fusion based on the multi-level first feature information and the multi-level second feature information by using a capsule routing network to generate multi-level fusion feature, wherein the multi-level fusion feature comprises consistency information of the multi-level first feature information and the multi-level second feature information; a second extracting module configured to obtain multi-level global feature based on the multi-level fusion feature by using a pyramid network; and a positioning module configured to perform positioning on the target video based on the multi-level global feature to obtain at least one target video segment associated with the positioning text.
[0010] According to another aspect of the embodiments of the present application, a storage medium is also provided, which comprises a stored program, wherein the program, when executed, controls a device where the storage medium is located to perform the video processing method.
[0011] According to another aspect of the embodiments of the present application, a processor is also provided, which is configured to execute a program, wherein the program, when executed, performs the video processing method.
[0012] According to another aspect of an embodiment of the present invention, a video processing method is also provided, including: receiving a product positioning instruction, wherein the product positioning instruction includes positioning text for describing at least one product in an e-commerce live broadcast video; performing positioning in the e-commerce live broadcast video based on the positioning text through a temporal text positioning model to obtain a positioning result, wherein the positioning result includes at least one live broadcast segment of the product associated with the positioning text; wherein the temporal text positioning model performs cross-modal and multi-level fusion on the multi-level first feature information corresponding to the positioning text and the multi-level second feature information corresponding to the target video to obtain multi-level global features, and locates the target video based on the multi-level global features to obtain at least one target video segment associated with the positioning text.
[0013] According to another aspect of an embodiment of the present invention, a video processing method is also provided, comprising: a cloud server receiving a text to be analyzed sent by a client, wherein the text to be analyzed is used to describe a video segment in a target video; the cloud server locates the text to be analyzed in the target video using a temporal text positioning model, obtaining a positioning result, wherein the positioning result includes at least one target video segment associated with the text to be analyzed; and the cloud server returns the positioning result to the client. The temporal text positioning model performs cross-modal multi-level fusion on the multi-level first feature information corresponding to the text to be analyzed and the multi-level second feature information corresponding to the target video to obtain a multi-level global feature, and locates the target video based on the multi-level global feature to obtain at least one target video segment associated with the text to be analyzed.
[0014] In an embodiment of the present invention, multi-level feature extraction is performed on the positioning text and the target video respectively to obtain multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video; based on the multi-level first feature information and the multi-level second feature information, feature fusion is performed through a capsule routing network to generate multi-level fusion features, wherein the multi-level fusion features include consistency information of the multi-level first feature information and the multi-level second feature information; based on the multi-level fusion features, multi-level global features are obtained through a pyramid network; the target video is positioned according to the multi-level global features to obtain at least one target video segment associated with the positioning text. The above scheme performs multi-level feature extraction on the positioning text and the target video respectively, and performs cross-modal multi-level fusion based on the extracted multi-level feature information to find the consistency between the target video and the positioning text, thereby obtaining a better text and video relationship, avoiding the problem of video and text being unable to align, thereby improving the accuracy of temporal text positioning, and solving the technical problem of low accuracy in temporal text positioning in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the application. In the drawings:
[0016] Figure 1 A hardware structure block diagram of a computing device (or mobile device) for implementing a video processing method is shown;
[0017] Figure 2 is a flowchart of a video processing method according to Embodiment 1 of the present application;
[0018] Figure 3 is a flowchart of another video processing method according to Embodiment 1 of the present application;
[0019] Figure 4 is a schematic diagram of a time sequence text positioning model according to Embodiment 2 of the present application;
[0020] Figure 5 is a schematic diagram of a video processing apparatus according to Embodiment 3 of the present application;
[0021] Figure 6 is a schematic diagram of a video processing apparatus according to Embodiment 4 of the present application;
[0022] Figure 7 is a structure block diagram of a computing device according to Embodiment 5 of the present application;
[0023] Figure 8 is a flowchart of another video processing method according to Embodiment 7 of the present application;
[0024] Figure 9 is a flowchart of another video processing method according to Embodiment 8 of the present application;
[0025] Figure 10a is a schematic diagram of video positioning on a client according to Embodiment 8 of the present application;
[0026] Figure 10b is a schematic diagram of displaying positioning results on a client according to Embodiment 8 of the present application;
[0027] Figure 11 is a schematic diagram of a video processing apparatus according to Embodiment 9 of the present application;
[0028] Figure 12 is a schematic diagram of a video processing apparatus according to Embodiment 10 of the present application. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0031] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:
[0032] Temporal text positioning: Based on a given sentence description, locate the starting time of the video clip described by the sentence in the video.
[0033] Capsule Routing Network: This network replaces neurons with directed vectors called capsules, and connects two layers of capsules in the network through dynamic routing. Capsule dynamic routing was proposed to extract consistency principles between multiple vectors, a feature that is well-suited to the multimodal fusion required for time-series text.
[0034] Dynamic routing: Assign weights to lower-level capsules to determine which higher-level capsule the lower-level capsule outputs to.
[0035] Trans-modality: refers to learning across multiple modalities such as images, videos, audio, and text. In this application, it refers to both video and text.
[0036] Example 1
[0037] According to an embodiment of the present invention, an embodiment of a video processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0038] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computing device or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computing device (or mobile device) for implementing a video processing method. Figure 1 As shown, the computing device 10 (or mobile device 10) may include one or more (illustrated as 102a, 102b, ..., 102n) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0039] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computing device 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0040] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the video processing method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the video processing method of the aforementioned application. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories may be connected to the computing device 10 via a network. Examples of the aforementioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0041] The transmission module 106 is configured to receive or send data via a network. The network can include a wireless network provided by a communication provider of the computing device 10. In one example, the transmission module 106 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In one example, the transmission module 106 can be a radio frequency (RF) module configured to communicate with the Internet via a wireless manner.
[0042] The display can be a liquid crystal display (LCD) that is touch screen, for example, which can enable a user to interact with a user interface of the computing device 10 (or mobile device).
[0043] It is noted that, in some alternative embodiments, the above-mentioned Figure 1 The computing device (or mobile device) can include hardware elements (including circuitry), software elements (including computer code stored on a computer readable medium), or a combination of both hardware and software elements. It should be noted that Figure 1 is merely one example of a particular implementation and is intended to illustrate the types of components that can be present in the above-described computing device (or mobile device).
[0044] In the above-described operating environment, the present application provides a video processing method as shown in Figure 2 Figure 2 is a flowchart of a video processing method according to an embodiment of the present application.
[0045] At step S21, a video positioning instruction is received, wherein the video positioning instruction includes at least a positioning text used to describe a video clip in a target video.
[0046] Specifically, the target video can be a TV series, an e-commerce live video, a teaching video, a traffic video, a news video, etc., and the positioning text can be a text used to position a desired video clip in the target video.
[0047] In an optional embodiment, taking a target video as an example, the positioning text can be a description text of a plot in the target video to find a segment of interest of the user; taking an e-commerce live video as an example, the positioning text can be a description text of one or more products in the live video to find an introduction to a product of interest in the live video; taking a teaching video as an example, the positioning text can be a description text of a knowledge point in the teaching video to find an explanation of a knowledge point of interest by a teacher; taking a traffic video as an example, the positioning text can be a description text of a traffic accident in the traffic video to understand the actual situation of the traffic accident; taking a news video as an example, the positioning text can be a description text of a news to find a report on a hot event of interest in the news video.
[0048] In step S23, the time-series text positioning model is used to position the target video based on the positioning text, and a positioning result is obtained, wherein the positioning result includes at least one target video segment associated with the positioning text.
[0049] In the above scheme, the time-series text positioning model is used to position the target video based on the positioning text, and a target video segment associated with the positioning text is obtained.
[0050] In the above scheme, the time-series text positioning model is used to position the target video based on the positioning text, and a target video segment associated with the positioning text is obtained.
[0051] In an optional embodiment, the above video positioning instruction can be issued by a user from a client, and the instruction is executed by a cloud server deployed in the cloud. The cloud server is deployed with a time-series text positioning model, and the positioning result is obtained through the time-series text positioning model and returned to the client for the user to view.
[0052] In an optional embodiment, the time-series text positioning model enters the target video and the positioning text into a multi-level video feature extraction module and a multi-level text feature extraction module to obtain multi-level semantic video (i.e., the above-mentioned multi-level second feature information) and multi-level text features (i.e., the above-mentioned multi-level first feature information), which are represented as and For each layer of semantic features corresponding to the positioning text and the target video, a capsule routing module is used for cross-modal multi-level fusion to obtain fused features Then, a cross-modal pyramid is used to obtain features aligned from local to global Finally, the is used for time-series positioning to obtain the final predicted start and end time, which is used to indicate the target video segment in the target video.
[0053] As can be seen from the above, the above embodiment of the present application receives a video positioning instruction, wherein the video positioning instruction includes at least a positioning text for describing a video segment in a target video; positioning is performed in the target video based on the positioning text by a temporal text positioning model to obtain a positioning result, wherein the positioning result includes at least one target video segment associated with the positioning text; wherein the temporal text positioning model performs cross-modal multi-level fusion on the multi-level first feature information corresponding to the positioning text and the multi-level second feature information corresponding to the target video to obtain multi-level global features, and locates the target video based on the multi-level global features to obtain at least one target video segment associated with the positioning text. The above scheme performs multi-level feature extraction on both the positioning text and the target video, and performs cross-modal multi-level fusion based on the extracted multi-level feature information to find the consistency between the target video and the positioning text, thereby obtaining a better text and video relationship, avoiding the problem of video and text being unable to align, thereby improving the accuracy of temporal text positioning, and solving the technical problem of low accuracy in temporal text positioning in the prior art.
[0054] As an optional embodiment, after positioning in the target video based on the positioning text through the temporal text positioning model and obtaining the positioning result, the above method also includes any one of the following: displaying the target video clip; displaying the start time and end time of the target video clip; marking the start time and end time of the target video clip on the timeline of the target video; and displaying the target video clip and the positioning text together according to the playback time of the target video.
[0055] After obtaining the positioning results, the target video clip in the positioning results can be directly displayed, and the start time and end time of the target video clip can also be displayed. The start time and end time of the target video clip can also be marked on the timeline of the target video to prompt the user the location of the target video clip, making it easier for the user to find the target video clip.
[0056] After obtaining the positioning results, the positioning text can be displayed when the target video clip appears during playback of the target video, so that the target video clip and the positioning text can be displayed together to serve as a prompt. For example, if the target video is a TV series, the positioning text is a description of a plot in the TV series. After determining the target video clip corresponding to the plot based on the positioning text, the positioning text is displayed in the form of a bullet screen when the target video clip is played during playback of the TV series.
[0057] As an optional embodiment, after locating the target video based on the positioning text through the temporal text positioning model and obtaining the positioning result, the above method also includes: receiving a correction instruction, wherein the correction instruction is used to perform at least one of the following processing on the positioning result: adjusting the start time and end time of the target video segment and deleting the target video segment; and correcting the positioning result according to the correction instruction.
[0058] If the positioning result is inaccurate, the user can modify it. For example, the user can change the start or end time of the target video segment, or delete the target video segment. After obtaining the modified positioning result, the target video segment in the positioning result can be output.
[0059] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0060] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0061] It should be noted that this embodiment may also include other steps in other embodiments if there is no conflict, which will not be repeated here.
[0062] Example 2
[0063] According to an embodiment of the present invention, another video processing method is provided. Figure 3 is a flowchart of another video processing method according to Example 1 of the present application, such as Figure 3 As shown, the method includes:
[0064] In step S31 , multi-level feature extraction is performed on the positioning text and the target video respectively to obtain multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video.
[0065] Specifically, the target video may be a TV series, an e-commerce live video, a teaching video, a traffic video, a news video, etc., and the positioning text may be a text used to locate the required video clip in the target video.
[0066] In an optional embodiment, taking the target video as a TV series as an example, the positioning text may be the description text of the plot therein, so as to find the segments that the user is interested in; taking the e-commerce live video as an example, the positioning text may be the description text of one or more products therein, so as to find the introduction of the products of interest in the live video; taking the teaching video as an example, the positioning text may be the description text of the knowledge points therein, so as to find the teacher's explanation of the knowledge points of interest; taking the traffic video as an example, the positioning text may be the description text of the traffic accident therein, so as to understand the actual situation of the traffic accident; taking the news video as an example, the positioning text may be the description text of a certain news, so as to find the report of the hot events of interest in the news video.
[0067] The above scheme performs multi-level feature extraction on both the positioning text and the target video. For the positioning text, multiple sub-texts containing different semantic information can be reconstructed based on the positioning text, and then feature extraction is performed on the multiple sub-texts respectively to obtain multi-level first feature information; for the target video, the multi-level feature extraction can be achieved through a multi-dimensional convolutional network and a cascaded multi-layer convolutional network, and the number of levels of the second feature information is the same as the number of levels of the cascaded convolutional network.
[0068] The target video is divided into a plurality of candidate video segments. In the above steps, corresponding multi-level second feature information can be extracted for each candidate video segment.
[0069] Step S33: Based on the multi-level first feature information and the multi-level second feature information, feature fusion is performed through a capsule routing network to generate a multi-level fusion feature, wherein the multi-level fusion feature includes consistency information of the multi-level first feature information and the multi-level second feature information.
[0070] The above scheme uses a capsule routing network to perform cross-modal fusion of the semantic features of each layer of multi-level first feature information and multi-level second feature information, and extracts the consistency information between the multi-level first feature information of the positioning text and the multi-level second feature information of the target video.
[0071] Similarly, the above steps may be to perform feature fusion based on the multi-level first feature information and the multi-level second feature information corresponding to each candidate video segment to obtain the multi-level fusion feature corresponding to each candidate video segment.
[0072] Step S35: obtaining multi-level global features through a pyramid network based on the multi-level fusion features.
[0073] The pyramid network is a feature pyramid network (FPN), which provides a top-down path to build higher resolution layers of semantically rich layers. The layers thus built have high resolution and rich semantics. However, due to the continuous up-sampling and down-sampling, the position of the object is already inaccurate, so the pyramid network builds a horizontal connection between the layers and the corresponding feature maps reconstructed, so that the detector can better predict the position.
[0074] By using the pyramid network, the multi-level first feature information of the positioned text and the multi-level second feature information of the target video are aligned from local to global.
[0075] Similarly, the above steps can be used to obtain the multi-level global feature corresponding to each candidate video segment through the pyramid network.
[0076] In step S37, the target video is positioned according to the multi-level global feature, and at least one target video segment associated with the positioned text is obtained.
[0077] Specifically, the multi-level global feature can be used to represent the consistency information of the positioned text and the candidate video segment, so that the consistency of the positioned text and the candidate video segment can be scored based on the multi-level global feature, and then the candidate video segment with the highest score is obtained, and the candidate video segment with the highest score is determined as the target video segment.
[0078] Figure 4 is a schematic diagram of a time sequence text positioning model according to Embodiment 2 of the present application, which is used to perform the method steps in the present embodiment. In an optional embodiment, the target video and the positioned text are input into a multi-level video feature extraction module and a multi-level text feature extraction module to obtain multi-level semantic video (i.e. the above-mentioned multi-level second feature information) and multi-level text features (i.e. the above-mentioned multi-level first feature information), which are represented as and For each layer of semantic features corresponding to the positioned text and the target video, a transmembrane state fusion is performed by using a capsule routing module to obtain the fused features Then, a transmembrane state pyramid is used to obtain the features aligned from local to global Finally, the is used for time sequence positioning to obtain the final predicted start and end time, which is used to indicate the target video segment in the target use.
[0079] In an optional embodiment, taking an educational scenario as an example, the target video can be a teaching video, and the positioning text can be a text description of a knowledge point in the teaching video. When a user wants to learn a certain knowledge point, they can generate positioning text by providing a text description of the knowledge point. Based on the positioning text, the above scheme can be used to locate the target video in the teaching video and obtain a video clip related to the knowledge point in the teaching video.
[0080] In another optional embodiment, taking a conference scenario as an example, the target video can be a conference video, and the anchor text can be a text description of a conference topic or a description of a speaker. When a user wants to watch videos related to a certain topic, the topic description can be used as the anchor text; when a user wants to watch a video of a speaker's speech, the speaker's description can be used as the anchor text to locate the desired part of the conference video.
[0081] As can be seen from the above, the above embodiments of the present application perform multi-level feature extraction on the positioning text and the target video respectively to obtain multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video; based on the multi-level first feature information and the multi-level second feature information, feature fusion is performed through a capsule routing network to generate multi-level fusion features, wherein the multi-level fusion features include consistency information of the multi-level first feature information and the multi-level second feature information; based on the multi-level fusion features, multi-level global features are obtained through a pyramid network; the target video is positioned according to the multi-level global features to obtain at least one target video segment associated with the positioning text. The above scheme performs multi-level feature extraction on the positioning text and target video respectively, and fuses the features of the two hierarchically through the capsule routing network, thereby learning the consistent representation of the target video and the positioning text, and aligns the multi-level fused features from local to global through the pyramid network, thereby realizing a multimodal capsule pyramid network, which can simultaneously perform local to global semantic alignment and fusion of text and video, thereby obtaining a better text and video relationship, avoiding the problem of video and text being unable to be aligned, and thus improving the accuracy of temporal text positioning, solving the technical problem of low accuracy in temporal text positioning in the existing technology.
[0082] As an optional embodiment, multi-level feature extraction is performed on the positioning text to obtain multi-level first feature information corresponding to the positioning text, including: creating multi-level sentences according to the structure of the positioning text, and encoding the multi-level sentences respectively to obtain the features of each level of sentences; using the features of each level of sentences to obtain text representations of the positioning text at multiple different granularities, wherein the multiple different granularities include at least the following two: word level, sentence level and paragraph level; splicing the representation texts at different granularities to obtain multi-level first feature information corresponding to the positioning text.
[0083] In an optional embodiment, the above-mentioned different granularities can be word level and sentence level. After the features of each level of sentences pass through the corresponding one-dimensional convolution layer, the output results of the one-dimensional convolution layer are spliced to obtain the representation of the located text at the word level, wherein the one-dimensional convolution layers corresponding to the features of each level of sentences are parallel to each other and have different convolution kernels; the features of each level of sentences pass through a multi-level bidirectional long short-term memory network or other natural language processing network (such as RNN) to obtain the representation of the located text at the sentence level; finally, the representation of the located text at the word level and the representation of the located text at the sentence level are spliced to obtain the multi-level first feature information corresponding to the located text.
[0084] In the above scheme, multiple levels of sentences are created based on the structure of the anchor text. This can be done by extracting multiple sentences from the anchor text, each containing a different amount of text information. In an optional embodiment, the first-level sentences only include the most basic text information in the anchor text. Each level of sentences can add some text information to the previous level. The last level of sentences includes all the information in the anchor text and can be the anchor text itself.
[0085] In the temporal text localization model, a corresponding one-dimensional convolutional layer is set for each level of sentences. The one-dimensional convolutional layers corresponding to each level of sentences are parallel to each other and have different convolution kernels. The Long Short-Term Memory (LSTM) network is a time-recursive neural network suitable for processing and predicting important events in time series with relatively long intervals and delays. In the above solution, through the multi-level bidirectional LSTM network, a sentence-level representation of each level of sentences can be obtained.
[0086] Finally, the word-level representation and sentence-level representation corresponding to each level of sentence are spliced together to obtain the first feature information of the text. Obviously, this first feature information is multi-level.
[0087] In an optional embodiment, three-level sentences can be created based on the positioning text, and the three-level sentences can be input into the GloVe (Global Vectors for Word Representation) model respectively. The GloVe model is a word representation tool based on count-based & overall statistics. It can express a word as a vector composed of real numbers. These vectors capture some semantic features between words, such as similarity and analogy. By encoding the three-level sentences through the GloVe model, we can get where i is used to represent the level. In this example, to obtain the word-level representation, for each S i Input three parallel one-dimensional convolutional layers with different convolution kernels, and input the obtained three vectors into a fully connected layer for dimension transformation, so as to obtain the final word-level representation Q w To obtain the sentence-level representation, S i Input three layers of bidirectional LSTM, and use the hidden units of the last layer as the representation Q of the entire sentence s Concatenate the word-level and sentence-level representations to obtain the final sentence representation Q. Perform the above operations on the multi-level sentence to obtain the multi-level sentence feature from local to global that is, the first feature information of the above multi-level.
[0088] As an optional embodiment, the multi-level sentence is created according to the structure of the positioning text, comprising: extracting a main sentence and multi-level subordinate sentences of the positioning text according to the structure of the positioning text; determining the main sentence as a first-level sentence, and determining other level sentences by adding a subordinate sentence to the last level sentence.
[0089] In an optional embodiment, the main sentence, the attributive clause and the adverbial clause can be extracted, and the main sentence is determined as a first-level sentence, the main sentence plus the attributive clause is determined as a second-level sentence, and the positioning text is determined as a third-level sentence.
[0090] In the above scheme, in order to obtain the semantic information from local to global, the sentence can be divided into three levels, i.e. the main sentence, the attributive clause and the adverbial clause, by using a preset nlp (Natural Language Processing) tool. The multi-level sentence information is obtained by sequentially adding the attributive clause and the adverbial clause to the main sentence. Therefore, the three-level text semantic information from local to global is respectively: (a) the main sentence; (b) the main sentence + the attributive clause; and (c) the complete sentence.
[0091] For example, the positioning text is "After explaining, the man mops the floor going around the furniture", the main sentence is "the man mops the floor", the attributive clause is "going around the furniture", and the adverbial clause is "After explaining". The three-level sentences are (a) the man mops the floor, (b) the man mops the floor going around the furniture, and (c) After explaining, the man mops the floor going around the furniture.
[0092] In an optional embodiment, the verbs and nouns of the positioning text can be extracted, and then the main clause, predicate clause and object adverbial clause can be determined. Finally, the main clause is determined to be a first-level sentence, the main clause plus the predicate clause is determined to be a second-level sentence, and the positioning text is determined to be a third-level sentence.
[0093] As an optional embodiment, multi-level feature extraction is performed on the target video to obtain multi-level second feature information corresponding to the target video, including: obtaining video feature information of the target video through a three-dimensional convolutional network; and obtaining multi-level second feature information corresponding to the target video through a cascaded multi-layer convolutional network based on the video feature information.
[0094] For video, the video input is first fed into the pre-trained C3D (3D ConvNets, three-dimensional convolutional network) network structure to obtain the entire video feature ∨, and then input into the cascaded multi-layer convolutional network to obtain the video features from local to global Here, i represents the level.
[0095] It can be seen that in order to solve the problem of semantic misalignment between the target video and the positioned text, the above scheme grades the video and text inputs from local to global levels respectively.
[0096] As an optional embodiment, feature fusion is performed through a capsule routing network based on multi-level first feature information and multi-level second feature information to generate multi-level fused features, including: extracting sentence capsule features of the positioning text based on the multi-level first feature information, and extracting video capsule features of the target video based on the multi-level second feature information; mapping the sentence capsule features and the video capsule features to the same space; splicing the mapped sentence capsule features and video capsule features to obtain fused low-level capsule features; transforming the fused low-level capsule features to high levels through dynamic routing in the capsule routing network to obtain fused high-level capsule features, and determining the fused high-level capsule features as multi-level fused features, wherein the dynamic routing in the capsule routing network is used to extract consistency information between the sentence capsule features and the video capsule features.
[0097] Specifically, the sentence capsule features of the located text can be extracted based on the multi-level first feature information through a one-dimensional convolutional network and an activation function in the capsule routing network, and the video capsule features of the target video can be extracted based on the multi-level second feature information through a one-dimensional convolutional network and an activation function in the capsule routing network.
[0098] Capsule dynamic routing is used to extract consistency information between multiple vectors, a characteristic that is well-suited to the multimodal fusion required for time-series text. Therefore, the above solution uses the capsule routing module to perform multimodal fusion of the features of the target video and the localized text. This process can discover the consistency relationship between the target video and the localized text pair.
[0099] The above scheme uses the capsule routing network to perform cross-modal fusion of the semantic features of each layer of multi-level first feature information and multi-level second information to extract the relationship between the two membrane states of the target video and the positioning text. First, the video capsule features are extracted by feeding V and Q into a one-dimensional convolutional network and a sigmoid activation function. and sentence capsule feature C s , where T represents the sequence length of the video feature. Unlike dynamic routing, here we also need to map the video and sentence capsules to the same space,
[0100] M v =W v ·C v
[0101] M s =W s ·C s ;
[0102] Where W v and W s is a learnable transformation matrix. Then M v and M sAfter splicing, the fused capsule M is obtained, and the high-level capsule F is obtained through dynamic routing. The high-level capsule F here is the above-mentioned multi-level fusion feature.
[0103] The following is an explanation of the capsule routing network. Given a low-level capsule l i and high-rise capsule h i The first step is to transform the dimensions of the low-level capsule to be consistent with the high-level capsule through the transformation matrix, which is specifically expressed as: μ j|i =w ij ·l i The second step is to cluster all capsules in the lower layers to the higher layers through the routing consensus algorithm. Specifically, the transformed capsule μ j|i By voting coefficient Aggregate to high-level capsules
[0104]
[0105] Where r represents the number of iterations, the voting coefficient By adding the routing factor b ij After softmax, we get b ij Update in an iterative manner. Specifically, given
[0106]
[0107] in By s (r) After squashing, The routing algorithm finds the most appropriate voting coefficient in an unsupervised iterative manner
[0108] As an optional embodiment, the multi-level fusion feature includes N levels, and the multi-level global feature also includes N levels, where N is an integer greater than 1. The multi-level global feature is obtained through a pyramid network based on the multi-level fusion feature, including: upsampling the multi-level global feature of the n+1th level to obtain a sampling result, and determining that the sum of the multi-level fusion feature of the nth level and the sampling result is the multi-level global feature of the nth level, wherein 0<n<N; determining the multi-level global feature of the Nth level to be the multi-level fusion feature of the Nth level.
[0109] In the above scheme, for non-highest levels, the global features of each level are obtained by upsampling the global features of the previous level and adding the fusion features of the current level; for the highest level, multi-level global features are equal to multi-level fusion features.
[0110] In an optional embodiment, taking N=3 as an example, the second-level multi-level global feature is up-sampled to obtain a first sampling result, and the sum of the first-level multi-level fusion feature and the first sampling result is determined as the first-level multi-level global feature; the third-level multi-level global feature is up-sampled to obtain a second sampling result, and the sum of the second-level multi-level fusion feature and the second sampling result is determined as the second-level multi-level global feature; and the third-level multi-level global feature is determined as the third-level multi-level fusion feature.
[0111] For example, after obtaining the multi-level global feature , the input pyramid network to obtain multi-level features fused from local to global Specifically, G3=F3, G2=F2+UP1(G3), and G1=F1+UP1(G2) to obtain multi-level first feature information of the positioning text and multi-level second feature information of the target video, which are aligned from local to global.
[0112] As an optional embodiment, the target video is positioned according to the multi-level global feature to obtain at least one target video segment associated with the positioning text, including: obtaining a two-dimensional time sequence diagram of the target video, wherein the two-dimensional time sequence diagram is used to represent candidate video segments in the target video, and the two-dimensional time sequence diagram includes a start time and an end time of each candidate video segment; obtaining a score corresponding to each candidate video segment in each two-dimensional time sequence diagram through multi-layer two-dimensional convolution based on the multi-level global feature; and determining that the candidate video segment with the highest score is the video segment associated with the positioning text.
[0113] Specifically, the above-mentioned two-dimensional time sequence diagram is used to model the temporal relationship between video segments. The two-dimensional time sequence diagram includes a plurality of two-dimensional time sequences, each of which represents a video segment in the target video. One dimension represents the start time of the video segment, and the other dimension represents the end time of the video segment. For example, the (i, j)th position on the time sequence diagram represents a candidate video segment starting at the ith time point and ending at the (j+1)th time point, where τ represents a preset time interval.
[0114] The above-mentioned candidate video segment can be a video segment pre-divided in the target video, or can be all possible video segments in the target video. For example, taking a television series as the target video, classic scenes in the television series can be taken as candidate video segments; and taking a live broadcast video of an e-commerce as an example, the introduction of each product in the live broadcast video can be taken as a candidate video segment.
[0115] In an alternative embodiment, the target video in step S31 is inputted through a two-dimensional time sequence diagram herein, which covers the time length of all possible video segments in the target video. Still taking the three-level example, the multi-level global feature Through steps S33-S35, the positioning text and each candidate video segment contained in the two-dimensional time sequence diagram can generate a corresponding multi-level global feature, and each multi-level global feature can be scored through the above multi-layer two-dimensional convolution to obtain the score corresponding to all possible candidate video segments wherein B is the number of candidate segments in the two-dimensional time sequence diagram, and the score is used to represent the consistency of each candidate video segment with the positioning text, so the candidate video segment with the highest score is the finally determined video segment associated with the positioning text.
[0116] As an alternative embodiment, the above method further comprises: obtaining the Intersection-over-Union (IoU) between the sample video segment and the ground truth; normalizing the IoU through a preset hyperparameter value to obtain a label value corresponding to the sample video segment; predicting the sample video segment through a preset initial model to obtain a prediction value; determining a target loss function according to the label value and the prediction value, and optimizing the initial model based on the target loss function.
[0117] In an alternative embodiment, the IoU value of each video candidate segment and the ground truth (ground-truth) is calculated and represented as u i . Then, the IoU value is normalized by two hyperparameter values t min and t max ,
[0118]
[0119] Then, the time sequence text positioning model is trained through the following cross-entropy loss function to obtain the final time sequence text positioning model, wherein the final time sequence text positioning model is used for time sequence text positioning according to the video processing method in the embodiment, and the cross-entropy loss function is the above target loss function:
[0120]
[0121] During prediction, the candidate segment with the highest value on the two-dimensional time sequence diagram is selected as the prediction result, and the two-dimensional time sequence diagram herein includes sample candidate video segments.
[0122] It should be noted that the present embodiment can also include other steps in other embodiments without conflict, which will not be described here.
[0123] Embodiment 3
[0124] The application provides a video processing device as shown in Figure 5 The application provides a video processing device as shown in Figure 5 The application provides a video processing device as shown in Figure 5 The device 500 comprises:
[0125] The receiving module 502 is configured to receive a video positioning instruction, wherein the video positioning instruction comprises at least a positioning text used for describing a video segment in a target video.
[0126] The positioning module 504 is configured to position the target video based on the positioning text by using a time-series text positioning model, to obtain a positioning result, wherein the positioning result comprises at least one target video segment associated with the positioning text.
[0127] The time-series text positioning model performs cross-modal multi-level fusion on multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video, to obtain multi-level global features, and positions the target video based on the multi-level global features, to obtain at least one target video segment associated with the positioning text.
[0128] It should be noted that the receiving module 502 and the positioning module 504 correspond to steps S21-S23 in Embodiment 1, and the two modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the modules as part of the device can run in the computing device 10 provided in Embodiment 1.
[0129] As an optional embodiment, the device further comprises any one of the following:
[0130] The first display module is configured to display the target video segment after the positioning module 504 positions the target video based on the positioning text by using the time-series text positioning model, to obtain the positioning result.
[0131] The second display module is configured to display the start time and the end time of the target video segment.
[0132] The labeling module is configured to label the start time and the end time of the target video segment on a time axis of the target video.
[0133] The third display module is configured to display the target video segment and the positioning text together according to a playing time of the target video.
[0134] As an optional embodiment, the device further comprises:
[0135] a receiving module configured to receive a correction instruction after obtaining a positioning result by performing positioning in a target video based on the positioning text using a temporal text positioning model, wherein the correction instruction is configured to perform at least one of the following processing on the positioning result: adjusting a start time or an end time of a target video segment, and deleting the target video segment;
[0136] The correction module is used to correct the positioning result according to the correction instruction.
[0137] Example 4
[0138] This application provides Figure 6 The video processing device is shown. Figure 6 is a schematic diagram of a video processing device according to embodiment 4 of the present application, combined with Figure 6 As shown, the apparatus 600 includes:
[0139] A first extraction module 602 is configured to perform multi-level feature extraction on the positioning text and the target video, respectively, to obtain multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video;
[0140] A fusion module 604 is configured to perform feature fusion based on the multi-level first feature information and the multi-level second feature information through a capsule routing network to generate a multi-level fusion feature, wherein the multi-level fusion feature includes consistency information of the multi-level first feature information and the multi-level second feature information;
[0141] The second extraction module 606 is used to obtain multi-level global features through a pyramid network based on the multi-level fusion features;
[0142] The positioning module 608 is configured to locate the target video according to the multi-level global features to obtain at least one target video segment associated with the positioning text.
[0143] It should be noted that the first extraction module 602, fusion module 604, second extraction module 606, and positioning module 608 described above correspond to steps S31 to S37 in Example 2. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules, as part of the apparatus, can be run in the computing device 10 provided in Example 1.
[0144] As an optional embodiment, the first extraction module includes:
[0145] The first encoding submodule is used to create multi-level sentences according to the structure of the positioning text, and encode the multi-level sentences respectively to obtain the features of each level of sentences;
[0146] A first acquisition submodule is configured to obtain text representations of the located text at multiple different granularities using the features of each level of sentences, wherein the multiple different granularities include at least two of the following: word level, sentence level, and paragraph level;
[0147] The first splicing submodule is used to splice the representation texts at different granularities to obtain multi-level first feature information corresponding to the positioning text.
[0148] As an optional embodiment, the first encoding submodule includes:
[0149] An extraction unit, configured to extract the main clause and multi-level subordinate clauses of the positioning text according to the structure of the positioning text;
[0150] The determination unit is used to determine that the main sentence is a first-level sentence, and to determine that other-level sentences are obtained by adding a first-level clause to the previous-level sentence.
[0151] As an optional embodiment, the first extraction module includes:
[0152] The second acquisition submodule is used to obtain video feature information of the target video through a three-dimensional convolutional network;
[0153] The third acquisition submodule is used to obtain multi-level second feature information corresponding to the target video through a cascaded multi-layer convolutional network based on the video feature information.
[0154] As an optional embodiment, the fusion module includes:
[0155] An extraction submodule, configured to extract sentence capsule features of the positioning text based on the multi-level first feature information, and to extract video capsule features of the target video based on the multi-level second feature information;
[0156] A mapping submodule for mapping sentence capsule features and video capsule features to the same space;
[0157] The second splicing submodule is used to splice the mapped sentence capsule features and video capsule features to obtain fused low-level capsule features;
[0158] The transformation submodule is used to transform the fused low-level capsule features to high levels through dynamic routing in the capsule routing network to obtain fused high-level capsule features, and determine the fused high-level capsule features as multi-level fusion features. The dynamic routing in the capsule routing network is used to extract consistency information between sentence capsule features and video capsule features.
[0159] As an optional embodiment, the multi-level fusion feature includes N levels, the multi-level global feature also includes N levels, N is an integer greater than 1, and the second extraction module includes:
[0160] The first determination submodule is used to upsample the multi-level global features of the n+1th level to obtain a sampling result, and determine the sum of the multi-level fusion features of the nth level and the sampling result as the multi-level global features of the nth level, wherein 0 <n<N;
[0161] The second determination submodule is used to determine the multi-level global features of the Nth level as the multi-level fusion features of the Nth level
[0162] As an optional embodiment, the positioning module includes:
[0163] A fourth acquisition submodule is configured to acquire a two-dimensional timing diagram of the target video, wherein the two-dimensional timing diagram is used to represent candidate video segments in the target video, and the two-dimensional timing diagram includes a start time and an end time of each candidate video segment;
[0164] The scoring submodule is used to obtain the score corresponding to each candidate video segment in each two-dimensional time sequence graph through multi-layer two-dimensional convolution based on multi-level global features;
[0165] The third determination submodule is configured to determine that the candidate video segment with the highest score is the video segment associated with the positioning text.
[0166] As an optional embodiment, the above device further includes:
[0167] An acquisition module is used to obtain the intersection-over-union ratio between the sample video clip and the true value;
[0168] A normalization module is used to normalize the intersection-over-union ratio using a preset hyperparameter value to obtain a label value corresponding to the sample video clip;
[0169] A prediction module is used to predict the sample video clip using a preset initial model to obtain a prediction value;
[0170] The optimization module is used to determine the target loss function according to the label value and the predicted value, and optimize the initial model based on the target loss function.
[0171] Example 5
[0172] An embodiment of the present invention may provide a computing device, which may be any computing device in a computing device group. Optionally, in this embodiment, the computing device may also be replaced by a terminal device such as a mobile terminal.
[0173] Optionally, in this embodiment, the computing device may be located in at least one network device among a plurality of network devices of a computer network.
[0174] In this embodiment, the computing device can execute the program code of the following steps in the method for processing a video of an application: performing multi-level feature extraction on the positioning text and the target video respectively to obtain multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video; performing feature fusion through a capsule routing network based on the multi-level first feature information and the multi-level second feature information to generate multi-level fusion features, wherein the multi-level fusion features include consistency information of the multi-level first feature information and the multi-level second feature information; obtaining multi-level global features through a pyramid network based on the multi-level fusion features; locating the target video according to the multi-level global features to obtain at least one target video segment associated with the positioning text.
[0175] Optionally, Figure 7 is a structural block diagram of a computing device according to an embodiment of the present invention. Figure 7 As shown, the computing device A may include: one or more (only one is shown in the figure) processors 702 , a memory 704 , and a peripheral interface 706 .
[0176] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the video processing method and device in the embodiments of the present invention. The processor executes the software programs and modules stored in the memory to perform various functional applications and data processing, thereby implementing the above-mentioned video processing method. The memory may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0177] The processor can call the information and application stored in the memory through the transmission device to execute the following steps: perform multi-level feature extraction on the positioning text and the target video respectively to obtain multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video; perform feature fusion through a capsule routing network based on the multi-level first feature information and the multi-level second feature information to generate multi-level fusion features, wherein the multi-level fusion features include consistency information of the multi-level first feature information and the multi-level second feature information; obtain multi-level global features through a pyramid network based on the multi-level fusion features; locate the target video according to the multi-level global features to obtain at least one target video segment associated with the positioning text.
[0178] Optionally, the processor can also execute the program code of the following steps: creating multi-level sentences according to the structure of the positioning text, and encoding the multi-level sentences separately to obtain the features of each level of sentences; obtaining the text representation of the positioning text at multiple different granularities using the features of each level of sentences, wherein the multiple different granularities include at least the following two: word level, sentence level and paragraph level; splicing the representation texts at different granularities to obtain the multi-level first feature information corresponding to the positioning text.
[0179] Optionally, the processor may also execute the program code of the following steps: extracting the main sentence and multi-level clauses of the positioning text according to the structure of the positioning text; determining that the main sentence is a first-level sentence, and determining that other-level sentences are obtained by adding a first-level clause to the previous-level sentence.
[0180] Optionally, the processor may also execute the program code of the following steps: obtaining video feature information of the target video through a three-dimensional convolutional network; and obtaining multi-level second feature information corresponding to the target video through a cascaded multi-layer convolutional network based on the video feature information.
[0181] Optionally, the processor may also execute the program code of the following steps: extracting sentence capsule features of the located text based on multi-level first feature information, and extracting video capsule features of the target video based on multi-level second feature information; mapping the sentence capsule features and the video capsule features to the same space; splicing the mapped sentence capsule features and the video capsule features to obtain fused low-level capsule features; transforming the fused low-level capsule features to high levels through dynamic routing in the capsule routing network to obtain fused high-level capsule features, and determining the fused high-level capsule features as multi-level fused features, wherein the dynamic routing in the capsule routing network is used to extract consistency information between the sentence capsule features and the video capsule features.
[0182] Optionally, the multi-level fusion feature includes N levels, and the multi-level global feature also includes N levels, where N is an integer greater than 1. The processor may further execute the following program code: upsampling the multi-level global feature of the n+1th level to obtain a sampling result, and determining the sum of the multi-level fusion feature of the nth level and the sampling result as the multi-level global feature of the nth level, wherein 0 <n<N;确定第N层级的多层级全局特征为第N层级的多层级融合特征。
[0183] The above-mentioned processor can also execute the program code of the following steps: obtaining a two-dimensional timing diagram of the target video, wherein the two-dimensional timing diagram is used to represent the candidate video segments in the target video, and the two-dimensional timing diagram includes the start time and end time of each candidate video segment; obtaining the score corresponding to each candidate video segment in each two-dimensional timing diagram through multi-layer two-dimensional convolution based on multi-level global features; determining that the candidate video segment with the highest score is the video segment associated with the positioning text.
[0184] An embodiment of the present invention provides a video processing solution. By performing multi-level feature extraction on the positioning text and the target video, respectively, multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video are obtained; based on the multi-level first feature information and the multi-level second feature information, feature fusion is performed through a capsule routing network to generate a multi-level fusion feature, wherein the multi-level fusion feature includes consistency information of the multi-level first feature information and the multi-level second feature information; based on the multi-level fusion feature, multi-level global features are obtained through a pyramid network; and the target video is located according to the multi-level global features to obtain at least one target video segment associated with the positioning text. The above scheme performs multi-level feature extraction on the positioning text and target video respectively, and fuses the features of the two hierarchically through the capsule routing network, thereby learning the consistent representation of the target video and the positioning text, and aligns the multi-level fused features from local to global through the pyramid network, thereby realizing a multimodal capsule pyramid network, which can simultaneously perform local to global semantic alignment and fusion of text and video, thereby obtaining a better text and video relationship, avoiding the problem of video and text being unable to be aligned, and thus improving the accuracy of temporal text positioning, solving the technical problem of low accuracy in temporal text positioning in the existing technology.
[0185] It can be understood by those skilled in the art that Figure 7 The structure shown is for illustration only, and the computing device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 7 It does not limit the structure of the above electronic device. For example, the computing device A may also include Figure 7 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 7 Different configurations shown.
[0186] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0187] Example 6
[0188] The embodiment of the present invention further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the video processing method provided in the first embodiment.
[0189] Optionally, in this embodiment, the above-mentioned storage medium may be located in any one of the computing devices in the computing device group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0190] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: performing multi-level feature extraction on the positioning text and the target video respectively to obtain multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video; performing feature fusion through a capsule routing network based on the multi-level first feature information and the multi-level second feature information to generate multi-level fusion features, wherein the multi-level fusion features include consistency information of the multi-level first feature information and the multi-level second feature information; obtaining multi-level global features through a pyramid network based on the multi-level fusion features; locating the target video according to the multi-level global features to obtain at least one target video segment associated with the positioning text.
[0191] Example 7
[0192] According to an embodiment of the present invention, another video processing method is provided. Figure 8 is a flowchart of another video processing method according to Example 7 of the present application, such as Figure 8 As shown, the method includes:
[0193] Step S81: Receive a product positioning instruction, wherein the product positioning instruction includes positioning text for describing at least one product in the e-commerce live video.
[0194] Specifically, the target video is an e-commerce live video, which can be a live e-commerce video, a recorded e-commerce video, or a video used to introduce products on an e-commerce platform.
[0195] In an optional embodiment, in an e-commerce scenario, in order to find an introduction video of a product of interest in a live e-commerce video, a text description of the product of interest can be made, and the text description can be used as positioning text to locate it in the e-commerce live video, so that an introduction video of the product of interest in the e-commerce live video can be obtained.
[0196] Step S83: Using a temporal text positioning model to locate the e-commerce live video based on the positioning text, obtaining a positioning result, wherein the positioning result includes at least one live broadcast segment of a product associated with the positioning text;
[0197] Among them, the temporal text positioning model performs cross-modal multi-level fusion on the multi-level first feature information corresponding to the positioning text and the multi-level second feature information corresponding to the target video to obtain multi-level global features, and locates the target video based on the multi-level global features to obtain at least one target video segment associated with the positioning text.
[0198] In the above solution, positioning is performed in the target video based on the positioning text using a temporal text positioning model to obtain a target video segment associated with the positioning text.
[0199] In an optional embodiment, the video positioning instruction can be issued by the user from the client, and the instruction is executed by a cloud server deployed in the cloud. The cloud server deploys a temporal text positioning model, obtains positioning results based on the temporal text positioning model, and returns the positioning results to the client for the user to view.
[0200] In an optional embodiment, the temporal text localization model inputs the target video and the localization text into the multi-level video feature extraction module and the multi-level text feature extraction module to obtain multi-level semantic video (i.e., the multi-level second feature information) and multi-level text features (i.e., the multi-level first feature information), which are respectively expressed as and For each layer of semantic features corresponding to the positioning text and the target video, the capsule routing module is used to perform cross-membrane multi-level fusion to obtain the fused features. Then, the features aligned from local to global are obtained through the transmembrane pyramid. Finally, use The final predicted start and end times are obtained by performing temporal positioning, and the start and end times are used to indicate the target video segment in the target use.
[0201] From the above, the above embodiments of the present application respectively perform multi-level feature extraction on the positioning text and the target video, and perform cross-modal multi-level fusion according to the extracted multi-level feature information, so as to find the consistency of the target video and the positioning text, thereby obtaining better text and video relationship, avoiding the problem that the video and the text cannot be aligned, and further improving the accuracy of the temporal text positioning, thereby solving the technical problem of low accuracy in the prior art when performing temporal text positioning.
[0202] It should be noted that the embodiments of the present application can also include other steps in other embodiments without conflict, which will not be described here.
[0203] Embodiment 8
[0204] According to the embodiments of the present application, another video processing method is also provided, Figure 9 is a flowchart of another video processing method according to Embodiment 8 of the present application, as shown in Figure 9 , the method comprises:
[0205] Step S91, the cloud server receives the text to be analyzed sent by the client, wherein the text to be analyzed is used to describe the video segment in the target video.
[0206] Specifically, the above target video can be a TV series, an e-commerce live video, a teaching video, a traffic video, a news video, etc., and the text to be analyzed can be a text used to position the required video segment in the target video.
[0207] In an optional embodiment, taking the target video as a TV series as an example, the text to be analyzed can be a description text of the plot therein, so as to find the segment of interest of the user; taking the target video as an e-commerce live video as an example, the text to be analyzed can be a description text of one or more products therein, so as to find the introduction of the product of interest in the live video; taking the target video as a teaching video as an example, the text to be analyzed can be a description text of a knowledge point therein, so as to find the explanation of the knowledge point of interest by the teacher; taking the target video as a traffic video as an example, the text to be analyzed can be a description text of a traffic accident therein, so as to understand the actual situation of the traffic accident; taking the target video as a news video as an example, the text to be analyzed can be a description text of a certain news, so as to find the report of the hot event of interest in the news video.
[0208] Step S93, the cloud server performs positioning in the target video based on the text to be analyzed through a temporal text positioning model, to obtain a positioning result, wherein the positioning result comprises at least one target video segment associated with the text to be analyzed;
[0209] Step S95, the cloud server returns the positioning result to the client.
[0210] Among them, the temporal text positioning model performs cross-modal multi-level fusion on the multi-level first feature information corresponding to the text to be analyzed and the multi-level second feature information corresponding to the target video to obtain multi-level global features, and locates the target video according to the multi-level global features to obtain at least one target video segment associated with the text to be analyzed.
[0211] In the above solution, the temporal text localization model is used to locate the text to be analyzed in the target video, thereby obtaining a target video segment associated with the text to be analyzed.
[0212] In an optional embodiment, the video positioning instruction can be issued by the user from the client, and the instruction is executed by a cloud server deployed in the cloud. The cloud server deploys a temporal text positioning model, obtains positioning results based on the temporal text positioning model, and returns the positioning results to the client for the user to view.
[0213] In an optional embodiment, the temporal text localization model enters the target video and the text to be analyzed into the multi-level video feature extraction module and the multi-level text feature extraction module to obtain multi-level semantic video (i.e., the multi-level second feature information) and multi-level text features (i.e., the multi-level first feature information), which are respectively expressed as and For each layer of semantic features corresponding to the text to be analyzed and the target video, the capsule routing module is used to perform trans-membrane multi-level fusion to obtain the fused features. Then, the features aligned from local to global are obtained through the transmembrane pyramid. Finally, use The final predicted start and end times are obtained by performing temporal positioning, and the start and end times are used to indicate the target video segment in the target use.
[0214] As can be seen from the above, the above embodiments of the present application perform multi-level feature extraction on the text to be analyzed and the target video respectively, and fuse the features of the two hierarchically through the capsule routing network, thereby learning the consistent representation of the target video and the text to be analyzed, and aligning the multi-level fused features from local to global through the pyramid network, thereby avoiding the problem of inability to align the video and text, thereby improving the accuracy of temporal text positioning, and solving the technical problem of low accuracy in temporal text positioning in the prior art.
[0215] Figure 10a This is a schematic diagram of video positioning on the client according to Example 8 of the present application, combined with Figure 10aFor example, in this example, the target video is a TV series video provided by the video client. The interface prompts the user to "enter the plot you are interested in and locate it for you immediately". The user enters the text to be analyzed in the input box to describe the plot. The cloud server locates the video based on the text to be analyzed and obtains the positioning result. Finally, the video can be located as follows: Figure 10b The positioning results are displayed in a small window floating on the top layer.
[0216] As an optional embodiment, after the cloud server returns the positioning result to the client, the above method also includes: the cloud server receives a correction instruction sent by the client, wherein the correction instruction is used to adjust the positioning result; the cloud server adjusts the positioning result according to the correction instruction; and the cloud server returns the adjusted correction instruction to the client.
[0217] After the cloud server returns the positioning results to the client, the user can modify them on the client. In the above solution, the user can modify the positioning results by sending a correction command to the cloud server from the client. This modification can include deleting the positioning results or changing the start or end time of the target video segment. After the cloud server adjusts the positioning results according to the correction command, it can return the adjusted positioning results to the client.
[0218] Example 9
[0219] This application provides Figure 11 The video processing device is shown. Figure 11 is a schematic diagram of a video processing device according to Example 9 of the present application, combined with Figure 11 As shown, the apparatus 1100 includes:
[0220] The receiving module 1102 is configured to receive a product location instruction, wherein the product location instruction includes location text for describing at least one product in the e-commerce live video;
[0221] A positioning module 1104 is configured to locate the e-commerce live video based on the positioning text using a temporal text positioning model, and obtain a positioning result, wherein the positioning result includes at least one live segment of a product associated with the positioning text;
[0222] Among them, the temporal text positioning model performs cross-modal multi-level fusion on the multi-level first feature information corresponding to the positioning text and the multi-level second feature information corresponding to the target video to obtain multi-level global features, and locates the target video based on the multi-level global features to obtain at least one target video segment associated with the positioning text.
[0223] Example 10
[0224] This application provides Figure 12 The video processing device is shown. Figure 12 is a schematic diagram of a video processing device according to embodiment 10 of the present application, combined with Figure 12 As shown, the apparatus 1200 includes:
[0225] Receiving module 1202, configured for the cloud server to receive the text to be analyzed sent by the client, wherein the text to be analyzed is used to describe a video segment in a target video;
[0226] A positioning module 1204 is configured to enable the cloud server to locate the text to be analyzed in the target video based on the text to be analyzed using a temporal text positioning model, and obtain a positioning result, wherein the positioning result includes at least one target video segment associated with the text to be analyzed;
[0227] The return module 1206 is used for the cloud server to return the positioning result to the client.
[0228] Among them, the temporal text localization model performs cross-modal multi-level fusion on the multi-level first feature information corresponding to the text to be analyzed and the multi-level second feature information corresponding to the target video to obtain multi-level global features, and locates the target video based on the multi-level global features to obtain at least one target video segment associated with the text to be analyzed.
[0229] As an optional embodiment, the above device further includes:
[0230] A correction instruction receiving module, configured to receive, from the cloud server, a correction instruction sent by the client after the cloud server returns the positioning result to the client, wherein the correction instruction is used to adjust the positioning result;
[0231] An adjustment module is used by the cloud server to adjust the positioning results according to the correction instructions;
[0232] The return module is used by the cloud server to return the adjusted correction instructions to the client.
[0233] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0234] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0235] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0236] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0237] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0238] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.
[0239] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A video processing method, characterized in that: include: receiving a video positioning instruction, wherein the video positioning instruction at least includes positioning text for describing a video segment in a target video; Locating the target video based on the positioning text using a temporal text positioning model to obtain a positioning result, wherein the positioning result includes at least one target video segment associated with the positioning text; The temporal text positioning model performs cross-modal multi-level fusion on the multi-level first feature information corresponding to the positioning text and the multi-level second feature information corresponding to the target video according to the hierarchy to obtain a multi-level global feature, wherein the multi-level global feature is used to represent the consistency information of the multi-level first feature information and the multi-level second feature information; the target video is positioned according to the multi-level global feature to obtain at least one target video segment associated with the positioning text; Among them, the multi-level first feature information is determined in the following way: reconstructing multiple sub-texts containing different semantic information based on the positioning text, and then performing feature extraction on the multiple sub-texts respectively to obtain multi-level first feature information; the multi-level second feature information is extracted through a multi-dimensional convolutional network and a cascaded multi-layer convolutional network.
2. The method according to claim 1, characterized in that After obtaining a positioning result by positioning the positioning text in the target video using a temporal text positioning model, the method further includes any one of the following: displaying the target video clip; Display the start time and end time of the target video segment; Marking the start time and end time of the target video segment on the time axis of the target video; The target video segment and the positioning text are displayed together according to the playing time of the target video.
3. The method according to claim 1, characterized in that After positioning the target video based on the positioning text using the temporal text positioning model to obtain a positioning result, the method further includes: receiving a correction instruction, wherein the correction instruction is used to perform at least one of the following processing on the positioning result: adjusting the start time or the end time of the target video segment, and deleting the target video segment; The positioning result is corrected according to the correction instruction.
4. A video processing method, characterized in that: include: Performing multi-level feature extraction on the positioning text and the target video respectively to obtain multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video; Based on the multi-level first feature information and the multi-level second feature information, feature fusion is performed according to the hierarchy through a capsule routing network to generate a multi-level fusion feature, wherein the multi-level fusion feature includes consistency information of the multi-level first feature information and the multi-level second feature information; Based on the multi-level fusion features, multi-level global features are obtained through a pyramid network; Position the target video according to multi-level global features, and obtain at least one target video segment associated with the positioning text. Among them, the multi-level first feature information is determined in the following way: reconstructing multiple sub-texts containing different semantic information based on the positioning text, and then performing feature extraction on the multiple sub-texts respectively to obtain multi-level first feature information; the multi-level second feature information is extracted through a multi-dimensional convolutional network and a cascaded multi-layer convolutional network.
5. The method according to claim 1, wherein Performing multi-level feature extraction on the positioning text to obtain multi-level first feature information corresponding to the positioning text includes: Creating multi-level sentences according to the structure of the positioning text, and encoding the multi-level sentences respectively to obtain features of each level of sentences; Using the features of each sentence level to obtain text representations of the located text at multiple different granularities, wherein the multiple different granularities include at least two of the following: word level, sentence level, and paragraph level; The representation texts at different granularities are spliced together to obtain multi-level first feature information corresponding to the positioning text.
6. The method according to claim 5, characterized in that Creating multi-level sentences based on the structure of the anchor text, including: Extracting the main sentence and multi-level subordinate clauses of the positioning text according to the structure of the positioning text; The main sentence is determined to be a first-level sentence, and other-level sentences are determined to be obtained by adding a first-level clause to the previous-level sentence.
7. The method according to claim 5, characterized in that Performing multi-level feature extraction on the target video to obtain multi-level second feature information corresponding to the target video includes: Acquire video feature information of the target video through a three-dimensional convolutional network; Based on the video feature information, multi-level second feature information corresponding to the target video is obtained through a cascaded multi-layer convolutional network.
8. The method according to claim 4, characterized in that Performing feature fusion based on the multi-level first feature information and the multi-level second feature information through a capsule routing network to generate a multi-level fusion feature, including: Extracting sentence capsule features of the positioning text according to the multi-level first feature information, and extracting video capsule features of the target video according to the multi-level second feature information; Mapping the sentence capsule features and the video capsule features to the same space; splicing the mapped sentence capsule features and the video capsule features to obtain fused low-level capsule features; The fused low-level capsule features are transformed to high levels through dynamic routing in the capsule routing network to obtain fused high-level capsule features, and the fused high-level capsule features are determined to be the multi-level fusion features, wherein the dynamic routing in the capsule routing network is used to extract consistency information between the sentence capsule features and the video capsule features.
9. The method according to claim 4, characterized in that The multi-level fusion feature includes N levels, and the multi-level global feature also includes N levels, where N is an integer greater than 1. The multi-level global feature is obtained through a pyramid network based on the multi-level fusion feature, including: The multi-level global features of the n+1th level are upsampled to obtain a sampling result, and the sum of the multi-level fusion features of the nth level and the sampling result is determined to be the multi-level global features of the nth level, wherein, 0 <n<N; The multi-level global features of the Nth level are determined as the multi-level fusion features of the Nth level.
10. The method according to claim 4, characterized in that Positioning the target video according to the multi-level global features to obtain at least one target video segment associated with the positioning text includes: Obtaining a two-dimensional timing graph of the target video, wherein the two-dimensional timing graph is used to represent candidate video segments in the target video, and the two-dimensional timing graph includes a start time and an end time of each candidate video segment; Obtaining a score corresponding to each candidate video segment in each two-dimensional time sequence graph through multi-layer two-dimensional convolution based on the multi-level global features; The candidate video segment with the highest score is determined to be the video segment associated with the positioning text.
11. A video processing device, characterized in that: include: A receiving module, configured to receive a video positioning instruction, wherein the video positioning instruction at least includes a positioning text for describing a video segment in a target video; a positioning module, configured to locate the target video based on the positioning text using a temporal text positioning model to obtain a positioning result, wherein the positioning result includes at least one target video segment associated with the positioning text; The temporal text positioning model performs cross-modal multi-level fusion on the multi-level first feature information corresponding to the positioning text and the multi-level second feature information corresponding to the target video according to the hierarchy to obtain multi-level global features, wherein the multi-level global features are used to represent the consistency information of the multi-level first feature information and the multi-level second feature information; and the target video is positioned according to the multi-level global features to obtain at least one target video segment associated with the positioning text. Among them, the multi-level first feature information is determined in the following way: reconstructing multiple sub-texts containing different semantic information based on the positioning text, and then performing feature extraction on the multiple sub-texts respectively to obtain multi-level first feature information; the multi-level second feature information is extracted through a multi-dimensional convolutional network and a cascaded multi-layer convolutional network.
12. A video processing device, characterized in that: include: A first extraction module is configured to perform multi-level feature extraction on the positioning text and the target video, respectively, to obtain multi-level first feature information corresponding to the positioning text and multi-level second feature information corresponding to the target video, wherein the multi-level first feature information is determined by reconstructing a plurality of subtexts containing different semantic information based on the positioning text, and then performing feature extraction on the plurality of subtexts to obtain the multi-level first feature information; the multi-level second feature information is extracted by a multi-dimensional convolutional network and a cascaded multi-layer convolutional network; A fusion module, configured to perform feature fusion according to the layers through a capsule routing network based on the multi-level first feature information and the multi-level second feature information to generate a multi-level fusion feature, wherein the multi-level fusion feature includes consistency information of the multi-level first feature information and the multi-level second feature information; A second extraction module is used to obtain multi-level global features through a pyramid network based on the multi-level fusion features; A positioning module is used to locate the target video according to multi-level global features to obtain at least one target video segment associated with the positioning text.
13. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the video processing method according to any one of claims 1 to 11.
14. A processor, characterized in that: The processor is configured to run a program, wherein the program, when running, executes the video processing method according to any one of claims 1 to 11.
15. A video processing method, characterized in that: Receive a product locating instruction, wherein the product locating instruction includes a locating text for describing at least one product in the e-commerce live video; Locating the e-commerce live video based on the positioning text using a temporal text positioning model to obtain a positioning result, wherein the positioning result includes at least one live broadcast segment of a product associated with the positioning text; The temporal text positioning model performs cross-modal multi-level fusion on the multi-level first feature information corresponding to the positioning text and the multi-level second feature information corresponding to the target video according to the hierarchy to obtain multi-level global features, and locates the target video according to the multi-level global features to obtain at least one target video segment associated with the positioning text, wherein the multi-level global features are used to represent consistency information of the multi-level first feature information and the multi-level second feature information; Among them, the multi-level first feature information is determined in the following way: reconstructing multiple sub-texts containing different semantic information based on the positioning text, and then performing feature extraction on the multiple sub-texts respectively to obtain multi-level first feature information; the multi-level second feature information is extracted through a multi-dimensional convolutional network and a cascaded multi-layer convolutional network.
16. A video processing method, characterized in that: include: The cloud server receives the text to be analyzed sent by the client, wherein the text to be analyzed is used to describe a video clip in a target video; The cloud server locates the text to be analyzed in the target video based on the temporal text positioning model to obtain a positioning result, wherein the positioning result includes at least one target video segment associated with the text to be analyzed; The cloud server returns the positioning result to the client; The temporal text localization model performs cross-modal multi-level fusion on the multi-level first feature information corresponding to the text to be analyzed and the multi-level second feature information corresponding to the target video according to the hierarchy to obtain multi-level global features, and locates the target video according to the multi-level global features to obtain at least one target video segment associated with the text to be analyzed, wherein the multi-level global features are used to represent consistency information of the multi-level first feature information and the multi-level second feature information; Among them, the multi-level first feature information is determined in the following way: reconstructing multiple sub-texts containing different semantic information based on the positioning text, and then performing feature extraction on the multiple sub-texts respectively to obtain multi-level first feature information; the multi-level second feature information is extracted through a multi-dimensional convolutional network and a cascaded multi-layer convolutional network.
17. The method according to claim 16, characterized in that After the cloud server returns the positioning result to the client, the method further includes: The cloud server receives a correction instruction sent by the client, wherein the correction instruction is used to adjust the positioning result; The cloud server adjusts the positioning result according to the correction instruction; The cloud server returns the adjusted correction instruction to the client.
Citation Information
Patent Citations
Video processing method and system, video player and cloud server
WO2017071227A1