Text processing method

By performing hierarchical processing and feature extraction on the target text, the problem of low retrieval accuracy in temporal text localization is solved, and the accurate localization and retrieval of target video segments are achieved.

CN115687701BActive Publication Date: 2026-07-31ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIBABA DAMO (HANGZHOU) TECH CO LTD
Filing Date
2021-07-23
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In time-series text localization, the accuracy of target video retrieval is relatively low because the segment described by the main clause may appear in multiple video segments simultaneously, causing the role of the main clause to be ignored and affecting the retrieval accuracy.

Method used

By performing layered processing on the target text, extracting multi-layered semantic features, and combining these with the video features of the target video, the target video segment can be identified, achieving precise positioning.

Benefits of technology

It improves the accuracy of target video retrieval, ensures that key information is effectively focused on and matched, and enhances the precision of retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115687701B_ABST
    Figure CN115687701B_ABST
Patent Text Reader

Abstract

This application discloses a text processing method. The method includes: responding to a first input command applied to an interface, inputting target text and a target video; responding to a text location command applied to the interface, displaying a target video segment in the target video that matches the target text on the interface, wherein the target video segment is obtained based on multi-layer semantic features of the target text and video features of the target video, the multi-layer semantic features are extracted from multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text. This application solves the technical problem of low accuracy in target video retrieval in related technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of text processing, and more specifically, to a text processing method. Background Technology

[0002] The purpose of temporal text localization is to locate temporal segments in undone long videos that correspond to a given sentence description. Due to its wide applications in video understanding, video retrieval, and human-computer interaction, it has attracted increasing attention from both industry and academia.

[0003] Currently, in time-series text localization, the segment corresponding to the main sentence description may appear simultaneously in multiple video segments. This phenomenon of the segment corresponding to the main sentence description appearing in multiple video segments will cause the main sentence to be ignored, and the focus will be on the other parts, resulting in low accuracy of target video retrieval.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This application provides a text processing method to at least address the technical problem of low accuracy in retrieving target videos in related technologies.

[0006] According to one aspect of the embodiments of this application, a text processing method is provided, comprising: responding to a first input instruction applied to an operation interface, inputting target text and a target video; responding to a text positioning instruction applied to the operation interface, displaying a target video segment in the target video that matches the target text on the operation interface, wherein the target video segment is obtained based on multi-layer semantic features of the target text and video features of the target video, the multi-layer semantic features are extracted from multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0007] According to another aspect of the embodiments of this application, another text processing method is also provided, including: responding to a second input instruction applied to a video display page, inputting target text, wherein a target video is displayed on the video display page; responding to a text positioning instruction applied to the video display page, displaying a target video segment in the target video that matches the target text on the video display page, wherein the target video segment is obtained based on multi-layer semantic features of the target text and video features of the target video, the multi-layer semantic features are extracted from multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0008] According to another aspect of the embodiments of this application, another text processing method is also provided, including: obtaining entertainment videos from an entertainment playback platform and displaying the entertainment videos on an operation interface; inputting target text in response to a third input instruction applied to the operation interface; and displaying a target video segment in the entertainment videos that matches the target text in response to a text positioning instruction applied to the operation interface, wherein the target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the entertainment videos, the multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0009] According to another aspect of the embodiments of this application, a text processing apparatus is also provided, comprising: a first input unit, configured to input target text and a target video in response to a first input instruction applied to an operation interface; and a first display unit, configured to display a target video segment in the target video that matches the target text on the operation interface in response to a text positioning instruction applied to the operation interface, wherein the target video segment is obtained based on multi-layer semantic features of the target text and video features of the target video, the multi-layer semantic features are extracted from multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0010] According to another aspect of the embodiments of this application, another text processing apparatus is also provided, including: a second input unit, configured to input target text in response to a second input instruction applied to a video display page, wherein a target video is displayed on the video display page; and a second display unit, configured to display a target video segment in the target video that matches the target text on the video display page in response to a text positioning instruction applied to the video display page, wherein the target video segment is obtained based on multi-layer semantic features of the target text and video features of the target video, the multi-layer semantic features are extracted from multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0011] In this embodiment, firstly, in response to a first input command applied to the operation interface, target text and target video are input; in response to a text positioning command applied to the operation interface, a target video segment matching the target text in the target video is displayed on the operation interface. The target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the target video. The multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text, thereby realizing the determination of the target video segment corresponding to the target text from the target video.

[0012] It is easy to note that by performing hierarchical processing on the target text, different levels of semantic text can be obtained. Each level of semantic text contains a main clause, thus enabling attention to text information at different levels. By extracting multi-level semantic features from the text information at different levels, and using the multi-level semantic features and the video features of the target video to determine the target video segment, the accuracy of target video retrieval can be improved.

[0013] Therefore, the solution provided in this application solves the technical problem of low accuracy in retrieving target videos in related technologies. Attached Figure Description

[0014] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0015] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a text processing method according to an embodiment of this application;

[0016] Figure 2 This is a flowchart of a text processing method according to an embodiment of this application;

[0017] Figure 3 This is a structural block diagram of a text processing method according to an embodiment of this application;

[0018] Figure 4 This is a flowchart of another text processing method according to an embodiment of this application;

[0019] Figure 5 This is a flowchart of another text processing method according to an embodiment of this application;

[0020] Figure 6 This is a schematic diagram of a text processing apparatus according to an embodiment of this application;

[0021] Figure 7 This is a schematic diagram of another text processing apparatus according to an embodiment of this application;

[0022] Figure 8 This is a schematic diagram of another text processing apparatus according to an embodiment of this application;

[0023] Figure 9 This is a structural block diagram of a computer terminal according to an embodiment of this application. Detailed Implementation

[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0026] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0027] Temporal text localization refers to locating the start time of a video segment described by a given sentence within a long video.

[0028] Cross-modal learning refers to learning across multiple modalities, such as images, videos, audio, and text. In this application, it specifically refers to the video and text modalities.

[0029] In English grammar, a main clause (also called an independent clause) is a combination of words consisting of a subject and a predicate that together represent a complete concept. In this application, the main clause is referred to as a predicate phrase.

[0030] The truth value represents the temporal segment in the video that corresponds to a given text.

[0031] Example 1

[0032] According to an embodiment of this application, an embodiment of a text processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0033] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware block diagram of a computer terminal (or mobile device) for implementing a text processing method is shown. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0034] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0035] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the text processing method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby implementing the above-mentioned application vulnerability detection method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0036] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0037] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0038] It should be noted here that, in some optional embodiments, the above... Figure 1 The computer device (or mobile device) shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 1 This is only one instance of a specific particular instance and is intended to illustrate the types of components that may exist in the aforementioned computer device (or mobile device).

[0039] In the aforementioned operating environment, this application provides, as follows: Figure 2 The text processing method shown Figure 2 This is a flowchart of a text processing method according to Embodiment 1 of this application. Figure 2 As shown, the method may include the following steps:

[0040] Step S202: In response to the first input command applied to the operation interface, input the target text and the target video.

[0041] The aforementioned user interface can be the user interface of a terminal device such as a computer.

[0042] In one alternative embodiment, a user can click on a pre-set input control in the operation interface to generate a first input instruction. At this time, the user can input target text and target video according to the first input instruction.

[0043] The target text mentioned above can be the text to be processed, and the target video mentioned above can be the video to be processed.

[0044] The target text mentioned above can be the target text of a road segment, and the target video mentioned above can be the surveillance video of the road segment. By acquiring the target text and the surveillance video, the target video segment that matches the target text can be determined from the surveillance video.

[0045] The target text mentioned above can be the target text of an instructional video, and the target video mentioned above can be an instructional video. By obtaining the target text and the instructional video, the target video segment that matches the target text can be identified in the instructional video.

[0046] The target text mentioned above can be target text in a live streaming platform, and the target video mentioned above can be target video in a live streaming platform. By obtaining the target text and the target video in the live streaming platform, the target video segment that matches the target text can be identified in the target video.

[0047] Step S204: In response to the text positioning command applied to the operation interface, display the target video segment in the target video that matches the target text on the operation interface. The target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the target video. The multi-layer semantic features are extracted from the multi-layer semantic text of the target text. The multi-layer semantic text is obtained by performing layered processing on the target text.

[0048] In one optional embodiment, after the user determines the input target text and target video, they can click a pre-set confirmation control on the operation interface to generate a text positioning instruction. At this time, the target video segment that matches the target text can be displayed on the operation interface.

[0049] Each layer of semantic text is a semantic text at a grammatical level, and includes the main sentence of the target text.

[0050] In one alternative embodiment, the target text can be processed in layers to obtain semantic text at different levels, thereby addressing the issue that the network ignores the role of the main clause because the main clause of the target text exists in both truth-value segments and non-truth-value segments.

[0051] In another alternative embodiment, the target text can be divided into three parts: the main clause, the relative clause, and the adverbial clause. Multi-level sentence information is obtained by sequentially adding the relative clauses and adverbial clauses to the main clause. The aforementioned multi-layered semantic text can be three levels of semantic text, where the three levels can be categorized from coarse to fine textual semantic information as the main clause, the main clause plus the relative clause, and the complete sentence. This allows for the acquisition of different levels of textual information, avoiding the problem of ignoring the main clause.

[0052] In another alternative embodiment, the semantic text of each layer can be input into a pre-trained model for encoding, resulting in... Then S iThe input is fed into a three-layer recurrent neural network (RNN), and the hidden units of its last layer are used as the semantic features Q of the entire sentence in the semantic text of that layer. The three-layer recurrent neural network can be a bidirectional long short-term memory (LSTM) network, and the pre-trained model used for each semantic text layer can be a 3D convolutional 3D (C3D) network model.

[0053] Furthermore, by performing the above encoding operations on multi-layered semantic text, multi-level sentence features can be obtained.

[0054] In another alternative embodiment, the target video can be input into a pre-trained model to obtain the video features of the entire target video. The pre-trained model used for the target video can be a Global Vectors for Word Representation (GloVe) model.

[0055] In another alternative embodiment, multi-layer semantic features and video features can be fused, and the target video segment that matches the target text can be determined in the target video based on the fused features.

[0056] In another optional embodiment, for video features, a two-dimensional temporal segment feature can first be generated. This two-dimensional temporal segment feature can be generated from multiple temporal segments in the video. Specifically, max pooling can be performed on any two segment features from the multiple temporal segments to obtain the two-dimensional temporal segment feature. After obtaining the two-dimensional temporal segment feature, the multi-layer semantic features and the two-dimensional temporal segment feature can be fused to obtain multiple features resulting from the fusion of multi-layer semantic features and video features. Finally, these multiple features are fused again to obtain a fused multi-layer feature, which is then input into the temporal localization module to predict the final temporal segment.

[0057] Through the above steps, firstly, in response to the first input command applied to the operation interface, the target text and the target video are input; in response to the text positioning command applied to the operation interface, the target video segment that matches the target text in the target video is displayed on the operation interface. The target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the target video. The multi-layer semantic features are extracted from the multi-layer semantic text of the target text. The multi-layer semantic text is obtained by performing layered processing on the target text, thus realizing the determination of the target video segment corresponding to the target text from the target video.

[0058] It is readily apparent that by performing hierarchical processing on the target text, different levels of semantic text can be obtained, each containing a main clause. This allows attention to textual information at different levels. By extracting multi-level semantic features from these different levels of textual information, and utilizing these multi-level semantic features along with the video features of the target video to determine the target video segment, the accuracy of target video retrieval can be improved. Therefore, the solution provided in this application addresses the technical problem of low accuracy in target video retrieval in related technologies.

[0059] In the above embodiments of this application, displaying a target video segment that matches the target text in the target video on the operation interface includes: displaying the target video segment and multi-layer semantic text on the operation interface; and / or the method further includes: in response to a annotation command applied to the operation interface, annotating the layers of the target text based on the multi-layer semantic text.

[0060] The aforementioned multi-layered semantic text includes first-layer semantic text, second-layer semantic text, and third-layer semantic text. The completeness of the first-layer semantic text is lower than that of the second-layer semantic text, and the completeness of the second-layer semantic text is lower than that of the third-layer semantic text.

[0061] The first layer of semantic text mentioned above can be the semantic text of the main clause.

[0062] The second layer of semantic text mentioned above can be the semantic text of the main clause and the semantic text of the relative clause.

[0063] The third layer of semantic text mentioned above can be the semantic text of a complete sentence.

[0064] In one alternative embodiment, the first layer of semantic text, the second layer of semantic text, and the third layer of semantic text can be semantic texts that progress from coarse to fine.

[0065] In another optional embodiment, after displaying the multi-layer semantic text, a pre-set annotation control in the operation interface can be clicked to generate annotation instructions. The main clause of the target text can be annotated as the first layer of semantic text, the main clause and relative clause of the target text can be annotated as the second layer of semantic text, and the completed sentences in the semantic text can be annotated as the third layer of semantic text.

[0066] In another alternative embodiment, the first statement is determined as the first-level semantic text, that is, the main clause is determined as the first-level semantic text; merging the first statement and the second statement yields the semantic text after merging the main clause and the relative clause, which is the second-level semantic text; merging the second-level semantic text with the third statement yields the complete sentence after merging the main clause, the relative clause, and the adverbial clause, which is the third-level semantic text.

[0067] In the above embodiments of this application, when there are multiple target video segments, the method further includes: displaying a first matching degree between each target video segment and the target text on the operation interface to obtain multiple first matching degrees; responding to a selection command applied to the operation interface, selecting a first target video segment from the multiple target video segments based on the multiple first matching degrees, and sending the first target video segment to the server, wherein the first target video segment is used by the server to adjust the target model, and the target model is used to process the multi-layer semantic features of the target text and the video features of the target video to obtain the target video segment.

[0068] The target video clip mentioned above can be the video clip that best matches the target text.

[0069] In one optional embodiment, the user interface can display the first matching degree between each target video segment and the target text, resulting in multiple first matching degrees. When the user sees multiple first matching degrees, they can press a pre-set selection control to generate a selection command. At this time, the user can select the first matching degree with the highest matching degree between the target video segment and the target text from the multiple first matching degrees according to the selection command, so as to determine the target video segment corresponding to the first matching pair as the first target video segment. Then, the first target video segment can be sent to the server, so that the server can use the first target video segment to adjust the target model to obtain a more accurate target model.

[0070] In the above embodiments of this application, the method further includes: responding to a title addition instruction applied to the operation interface, displaying the title of each target video segment on the operation interface, wherein the title is determined based on the target text and the target video.

[0071] In an alternative embodiment, the user can also press a pre-set title-adding control to generate a title-adding instruction, at which point a title can be added to each target video segment according to the title-adding instruction.

[0072] In another alternative embodiment, text corresponding to the target video segment can be extracted from the target text and used as the title of the target video segment.

[0073] In this embodiment of the application, each layer of semantic text is a semantic text at a grammatical level and includes the main clause of the target text. The multi-layer semantic text includes a first layer of semantic text, a second layer of semantic text, and a third layer of semantic text. The completeness of the first layer of semantic text is lower than that of the second layer of semantic text, and the completeness of the second layer of semantic text is lower than that of the third layer of semantic text.

[0074] In this embodiment of the application, the method further includes: responding to a semantic text generation instruction applied to the operation interface, dividing the target text according to the target syntax to obtain a first statement, a second statement and a third statement of the target text, determining the first statement as the first layer of semantic text, merging the first statement and the second statement to obtain the second layer of semantic text, and merging the second layer of semantic text with the third statement to obtain the third layer of semantic text.

[0075] In one alternative embodiment, the target text can be segmented according to the target grammar using Stanford Natural Language Processing (NLP) to obtain a first statement, a second statement, and a third statement, wherein the first statement is the main clause of the target text, the second statement is the relative clause of the target text, and the third statement is the adverbial clause of the target text.

[0076] In this embodiment of the application, the method further includes: responding to a first model generation instruction applied to the operation interface, determining a first loss function based on a first sample set and / or a second sample set of video samples, and training a first fusion model based on the first loss function; wherein, the first fusion model is used to fuse the first-layer semantic features of the first-layer semantic text and the temporal segment features of the video features to obtain a first sub-fusion feature, and is used to fuse the second-layer semantic features of the second-layer semantic text and the temporal segment features to obtain a second sub-fusion feature, the first sub-fusion feature, the second sub-fusion feature and the third sub-fusion feature are used to determine the target video segment in the target video, and the third sub-fusion feature is the fusion feature between the third-layer semantic features of the third-layer semantic text and the temporal segment features.

[0077] In one alternative embodiment, a first convolutional layer may be determined based on a first sample set, and a second convolutional layer may be determined based on a second sample set; a first loss function corresponding to the first and second convolutional layers may be determined; and the first and second convolutional layers may be trained based on the first loss function to obtain a first fusion model.

[0078] In one alternative embodiment, g:R can be defined. d →R is the convolutional layer in the first fusion model. Further, we can obtain X = g(M). Definition Given the loss function, we can further derive the first convolutional layer as X. p =g(M p The second convolutional layer is X. u =g(M u ).

[0079] The loss assessment of g above can be written in the following form and set as follows: Where, π p=P(Y=+1) is the class prior probability, that is, it represents the class prior probability that the sample is positive; π u =P(Y=-1)=1-π p .

[0080] In one optional embodiment, determining a target video segment that matches the target text in a target video based on multi-layer semantic features and video features includes: fusing multi-layer semantic features and video features to obtain fused features; and determining the target video segment in the target video based on the fused features.

[0081] In another alternative embodiment, the three semantic features can be fused with the video features respectively to obtain three fused features. Then, these three fused features are fused together to obtain a total fused feature, and the target video segment is determined in the target video based on the fused feature.

[0082] In another optional embodiment, the first-layer semantic features and temporal segment features of the first-layer semantic text can be fused based on the first fusion model to obtain the first sub-fusion feature. The first fusion model is trained based on a first sample set and a second sample set of video samples. The first sample set includes the first positive temporal segment feature samples of the video samples, and the second sample set includes the second positive temporal segment feature samples and negative temporal segment feature samples of the video samples. The second-layer semantic features and temporal segment features of the second-layer semantic text are fused based on the first fusion model to obtain the second sub-fusion feature. The third-layer semantic features and temporal segment features of the third-layer semantic text are fused based on the second fusion model to obtain the third sub-fusion feature. The second fusion model is trained based on semantic feature labels and temporal segment labels.

[0083] In another optional embodiment, the temporal segment features of each layer of semantic features and video features are fused to obtain multiple sub-fused features.

[0084] The first sample set mentioned above can be the positive sample set M. p The second sample set mentioned above can be a set of positive samples and a set of negative samples M. u .

[0085] The first sample set includes n p m positive samples were sampled from P(m|Y=+1). p The second sample set includes n u m positive samples were sampled from P(m). u Where Y∈{+1,-1} is the output random variable.

[0086] In one alternative embodiment, the first fusion model is trained through multi-instance positive sample unlabeled learning.

[0087] In this embodiment of the application, in response to a first model generation instruction applied to the operation interface, determining a first loss function based on a first sample set and / or a second sample set of video samples includes: in response to a first model generation instruction applied to the operation interface, determining a second loss function based on the start and end times of the target video segment samples, the similarity between the text samples and the first or second sample set, and determining the first loss function based on the second loss function.

[0088] The method further includes: responding to a second model generation instruction applied to the operation interface, determining a third loss function based on semantic feature labels and temporal segment labels, and training a first initial model based on the third loss function to obtain a second fusion model, wherein the second fusion model is used to fuse the third-layer semantic features and temporal segment features of the third-layer semantic text to obtain a third sub-fusion feature.

[0089] The second fusion model mentioned above can be a feature pyramid network, which is obtained through supervised learning training.

[0090] In one optional embodiment, multiple temporal segment features of video features can be obtained. Then, any two temporal segment features from these multiple temporal segment features are subjected to max pooling to obtain two-dimensional temporal segment features M. Then, the semantic features of each layer and the two-dimensional temporal segment features M are fused to obtain sub-fused features of the semantic features of each layer.

[0091] In another alternative embodiment, a feature pyramid network can be used to fuse the sub-fusion features of each layer of semantic features to obtain a fused three-layer fusion feature.

[0092] In another alternative embodiment, temporal segment features of semantic features and video features at each layer can be fused to obtain multiple sub-fused features, wherein each sub-fused feature corresponds to each layer of semantic features; the multiple sub-fused features are fused based on the feature pyramid network model to obtain fused features.

[0093] In another optional embodiment, the temporal segment features of each layer of semantic features and video features are fused to obtain multiple sub-fused features. Specifically, this can be done by obtaining the first product between the first learning parameter, each layer of semantic features and the target vector; obtaining the second product between the second learning parameter and the temporal segment features; obtaining the third product between the first product and the second product; and normalizing the third product to obtain the sub-fused features corresponding to each layer of semantic features.

[0094] The third product mentioned above can be the Hadamard product.

[0095] The first learning parameter mentioned above is the parameter for semantic feature learning, and the second learning parameter mentioned above is the parameter for temporal segment features of video features. The first and second learning parameters are used for feature mapping.

[0096] The target vector mentioned above can be a constant vector, which can be 1. T , of which 1 T The transpose of a vector of all elements with dimension 1, for example, a vector of dimension d with dimension 1. T It's just d 1s.

[0097] In an alternative embodiment, feature fusion can be performed using the following formula:

[0098] F i =||(w q ·Q i ·1 T )⊙(w m ·M)|| F ;

[0099] Among them, w q Q is the first learned parameter. i For each layer of semantic features, 1 T Let w be the target vector. m Here, M represents the temporal segment feature, ⊙ denotes the Hadamard product, and ||·|| F This represents the normalization of the norm (also known as Frobenius).

[0100] In another alternative embodiment, the aforementioned temporal segment features M can be obtained based on video features. Specifically, the video features can be V, where V represents the feature sequence of each time segment. For example, if a video has 10 segments, and each segment contains 16 frames, then feature extraction via a convolutional neural network will yield the following results. Among them, the middle v i ∈R d In other words, each segment is a d-dimensional feature. Then, by performing max-pooling on the segment features between any two time periods, that is, the time-series segment features, we obtain M∈R. 10×10×d .

[0101] In another alternative embodiment, a first loss function can be determined based on a first convolutional layer, a first prior probability associated with the first convolutional layer, a second convolutional layer, and a second prior probability associated with the second convolutional layer, wherein the first and second prior probabilities are determined by the number of multiple candidate video segment samples that have a similarity to the target video segment sample higher than a target threshold, the target video segment sample being matched with a text sample, and the video sample including multiple candidate video segment samples.

[0102] In another alternative embodiment, the number of candidate segments with high similarity in the ground truth segments can be used as a priori, since M is not labeled in the learning process in the positive samples. u If the label is unknown, it can be estimated directly. Because of π n P n (x)=P(x)-π p P p (x), where π n The probability prediction formula represents the probability that an input x is predicted as a positive sample. Therefore, R(g) can be directly estimated by the following formula:

[0103]

[0104] in, The loss function represents unlabeled data, which is all segments in the entire video, with each segment treated as a sample. This represents the loss when the sample is predicted as +1. This represents the loss when a sample is predicted as -1.

[0105] In one alternative embodiment, a candidate video segment sample package can be constructed for the target video segment sample, wherein the candidate video segment sample package includes multiple candidate video segment samples, and each candidate video segment sample is a positive candidate video segment sample.

[0106] In one optional embodiment, constructing a candidate video segment sample bag for the target video segment sample means forming a bag from all the positive samples in the labeled segments, and this bag must contain positive samples.

[0107] In the above embodiments of this application, the method further includes: determining a second loss function corresponding to the candidate video segment sample package based on the start and end times of the target video segment sample and the similarity between the text sample and the first sample set or the second sample set; and determining a first loss function based on the second loss function.

[0108] In an alternative embodiment, to avoid overfitting during training due to the possibility of negative numbers in the formula, a non-negative version is proposed:

[0109]

[0110] Unlike standard unlabeled learning with positive samples, labeled data contains both positive and negative samples. In other words, it is uncertain whether the segment described by the main sentence occurs in the video, but the main sentence may appear between labeled segments. Here, labeled segments refer to the segments in the video that are labeled to correspond to a sentence. For example, if the sentence describes the segment between 1 and 10 seconds in the video, then the video segment between 1 and 10 seconds is the labeled segment, which is mainly used during training.

[0111] Furthermore, a positive sample candidate fragment bag can be constructed for each labeled truth fragment, i.e., B = {(l s , l e )|l s , l e ∈[t s , t e ];l s ≤l e}. Among them, l s , l e These represent the start and end times of the truth segment, respectively. Then, the second loss function corresponding to the candidate video segment sample bag is obtained, i.e.

[0112]

[0113] Among them, M i,j Let a represent the (i, j)th position on M. i,j = <Q,M i,j >It's Q and M i,j The similarity between the sentences is Q, where Q is the representation of the sentences. The set of scores and similarities can be defined as follows: The set of scores and similarities, also known as the g function, is the probability value that the current segment is predicted as a positive sample.

[0114] Furthermore, the first loss function determined based on the second loss function can be: in, Even with the π of goods p The evaluation method is the same, where i represents the i-th layer, and the value of i can only be 1 or 2.

[0115] In another alternative embodiment, a third loss function can be determined based on semantic feature labels and temporal segment labels; the first initial model is trained based on the third loss function to obtain a second fusion model.

[0116] The third loss function mentioned above is This is represented as cross-entropy loss. For the π of the i-th layerp A, n u X n X p .

[0117] In another alternative embodiment, the first sub-fusion feature can be determined as the first fusion feature based on the feature pyramid network model; the first fusion feature and the second sub-fusion feature can be fused based on the feature pyramid network model to obtain the second fusion feature; and the second fusion feature and the third sub-fusion feature can be fused based on the feature pyramid network model to obtain the third fusion feature.

[0118] In another alternative embodiment, the input three-layer hierarchical network is based on fused features. The fused three-layer features can be obtained using a feature pyramid network. The first fusion feature can be G1 = F1; the second fusion feature can be G2 = F2 + UP1(G1); the third fusion feature can be G3 = F3 + UP2(G2).

[0119] UP1 and UP2 refer to two upsampling layers, respectively.

[0120] In this embodiment of the application, the method further includes at least one of the following: displaying the target start time and target end time of the target video segment on the operation interface; sending the target start time and target end time of the target video segment to the target client; and marking the target video segment in the target video.

[0121] In an optional embodiment, the third fusion feature can be processed based on a temporal localization model to obtain the target start time and target end time of the target video segment, wherein the temporal localization model is used to predict the temporal segment associated with the input feature.

[0122] In another alternative embodiment, the obtained The input timing localization module predicts the final timing segments, and then the scores of all valid candidate segments on the timing graph are defined as follows: Among them, l i This indicates the total number of candidate segments. Each This indicates the confidence level of the match between the corresponding candidate fragment and the query sentence.

[0123] In another optional embodiment, the confidence scores between multiple candidate video segment samples and text samples in the video samples can be determined to obtain multiple confidence scores; the intersection-union ratio (IU) between each candidate video segment sample and the target video segment sample can be obtained to obtain multiple IU values, wherein the target video segment sample matches the text sample; a fourth loss function is determined based on the multiple confidence scores and multiple IU values, and the initial model is trained based on the fourth loss function to obtain a temporal localization model.

[0124] In another alternative embodiment, normalized IoU values ​​can be used as ground truth during the training phase.

[0125] Specifically, firstly, the IoU value between each time series candidate segment and the ground truth is calculated and expressed as... Then, the IoU value is passed through two hyperparameter values ​​t. min and t max After normalization, the following results are obtained:

[0126]

[0127] The network can be trained using the following fourth loss function:

[0128]

[0129] in, The loss function for the last three layers can be expressed as follows: During the prediction phase, the candidate segment with the highest value can be selected as the segment that best matches the text. Among these, It represents the free loss of the target video, with 3 indicating the 3rd layer.

[0130] In the above embodiments of this application, the method further includes: determining multiple target candidate video segments in the target video that are associated with the fusion features between multi-layer semantic features and video features; obtaining a second matching degree between each target candidate video segment and the target text to obtain multiple second matching degrees; determining the target candidate video segment corresponding to the highest matching degree among the multiple second matching degrees as the target video segment; and / or encoding each layer of semantic text to obtain multiple encoding results, wherein each encoding result includes word vectors of each layer of semantic text; and extracting each layer of semantic features from each encoding result.

[0131] In one optional embodiment, after identifying multiple target candidate video segments associated with fusion features in the target video, matching pairs between each target candidate video segment and the target text can be obtained to obtain multiple second matching degrees. Then, the target candidate video segment corresponding to the highest matching degree among the second matching degrees can be selected as the target video segment to improve the accuracy of the target video segment.

[0132] In another alternative embodiment, each layer of semantic text can be encoded to obtain multiple encoding results, wherein each encoding result includes word vectors of each layer of semantic text; each encoding result is processed based on a temporal recursive network model to obtain multi-layer semantic features.

[0133] In one alternative embodiment, the GloVe model can be used to encode text in each layer, resulting in multiple encoding results, for example... Then, a three-layer bidirectional LSTM is used to process each encoding result separately, and the hidden units of the last layer are used as the representation Q of the entire sentence. Furthermore, the above operations are performed on multi-level sentences to obtain multi-level sentence features.

[0134] The following is combined Figure 3 A preferred embodiment of this application will be described in detail. The method can be executed by a mobile terminal or a server. In this embodiment, the method is executed by a server as an example.

[0135] Figure 3 This is a flowchart of a text processing method according to an embodiment of this application. The method includes the following steps:

[0136] Step S301, input video;

[0137] Step S302, Input text;

[0138] Step S303: Perform text layering on the input text to obtain multi-layer semantic text;

[0139] Step S304: Extract text features from the multi-layer semantic text to obtain multi-layer semantic features, and input the multi-layer semantic features into the hierarchical text network in step S306;

[0140] Step S305: Extract video features from the input video and input the video features into the hierarchical text network in step S306.

[0141] Step S306: Use a hierarchical text network to identify video segments in the input video that match the input text.

[0142] In one optional embodiment, monitoring videos of road segments can be obtained from a road monitoring platform and displayed on an operation interface; in response to a fourth input command applied to the operation interface, target text can be input; in response to a text positioning command applied to the operation interface, a target video segment matching the target text in the monitoring video can be displayed on the operation interface, wherein the target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the monitoring video, the multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0143] Furthermore, the target video clip and multi-layer semantic text can be sent to the road monitoring platform, where the multi-layer semantic text is used to annotate the hierarchy of the target text on the road monitoring platform.

[0144] The surveillance videos of the aforementioned road segments can be obtained from a pre-set surveillance video database, which may include surveillance videos of multiple road segments.

[0145] The target text mentioned above can be obtained from a pre-set vehicle information database, which may include vehicle information for multiple vehicles.

[0146] In another optional embodiment, teaching videos can be obtained from the teaching platform and displayed on the operation interface; in response to a fifth input command applied to the operation interface, target text can be input; in response to a text positioning command applied to the operation interface, a target video segment matching the target text in the teaching video can be displayed on the operation interface, wherein the target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the teaching video, the multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0147] Furthermore, the target video clip and multi-layered semantic text can be sent to the teaching video platform, where the multi-layered semantic text is used to annotate the hierarchy of the target text on the teaching video platform.

[0148] The teaching videos contain object information for different teaching objects, including teachers and / or different types of teaching content. The target text is used to describe the target object information of the target teaching object, including the target teacher and / or the target type of target teaching content.

[0149] The aforementioned teaching videos can be obtained from a pre-set teaching video database, which can store teaching videos from multiple teachers and various types of teaching videos.

[0150] The target text mentioned above can be obtained from a pre-set teaching object database, which includes information on multiple teachers and / or multiple types of teaching content.

[0151] In another alternative embodiment, a target video can be obtained from a live streaming platform and displayed on an operation interface; in response to a sixth input command applied to the operation interface, target text can be input; in response to a text positioning command applied to the operation interface, a target video segment matching the target text in the target video can be displayed on the operation interface, wherein the target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the target video, the multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0152] Furthermore, the target video clip and multi-layered semantic text can be sent to the live streaming platform, where the multi-layered semantic text is used to annotate the hierarchy of the target text on the live streaming platform.

[0153] The target video mentioned above can be a video that is being broadcast live on the live streaming platform, or a previously recorded video that the live streaming platform has saved.

[0154] In another alternative embodiment, the client can obtain the target text and target video to be processed, then upload the target text and target video to the server, and finally receive the target video segment that matches the target text from the target video returned by the server.

[0155] In another optional embodiment, to better process the target video, the acquired target video and target text can be transmitted to a corresponding processing device for processing. For example, they can be directly transmitted to the user's computer terminal (e.g., a laptop, personal computer, etc.) for processing, or transmitted to a cloud server through the user's computer terminal for processing. It should be noted that since the target video and target text require a large amount of computing resources, this embodiment uses a cloud server as an example for explanation.

[0156] For example, to facilitate users uploading target videos and text, an interactive interface can be provided. Users can select the target video by clicking the "Select Image and Text" control, and then upload the target video and text to the cloud server by clicking the "Upload" control. Furthermore, to help users confirm that the uploaded target video and text are correct, the selected target video and text can be displayed in the "Image and Text Display" area. After confirmation, the user can click the "Upload" control to upload the target video and text.

[0157] The target video segment is determined based on the multi-layer semantic features of the target text and the video features of the target video. The multi-layer semantic features are extracted from the multi-layer semantic text of the target text. The multi-layer semantic text is obtained by performing hierarchical processing on the target text. Each layer of semantic text is a semantic text at a grammatical level and includes the main sentence of the target text.

[0158] In another alternative embodiment, a medical video can be obtained from a medical platform and displayed on an operation interface; in response to a seventh input command applied to the operation interface, target text can be input; in response to a text positioning command applied to the operation interface, a target video segment in the medical video that matches the target text can be displayed on the operation interface, wherein the target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the medical video, the multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0159] Furthermore, the target video clip and multi-layer semantic text can be sent to the medical platform, where the multi-layer semantic text is used to annotate the hierarchy of the target text on the medical platform.

[0160] The aforementioned medical videos can be obtained from a pre-set medical video database, which may include multiple medical videos.

[0161] The target text mentioned above can be obtained from a pre-set medical information database, which may include multiple medical information entries.

[0162] In another alternative embodiment, the meeting video can be obtained from the meeting platform and displayed on the operation interface; in response to the eighth input command applied to the operation interface, target text can be input; in response to the text positioning command applied to the operation interface, a target video segment in the meeting video that matches the target text can be displayed on the operation interface, wherein the target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the meeting video, the multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0163] Furthermore, the target video clip and multi-layered semantic text can be sent to the conferencing platform, where the multi-layered semantic text is used to annotate the hierarchy of the target text on the conferencing platform.

[0164] The aforementioned meeting videos can be obtained from a pre-set meeting video database, which may include multiple meeting videos.

[0165] The target text mentioned above can be obtained from a pre-set database of meeting information, which may contain multiple meeting information entries.

[0166] Example 2

[0167] According to an embodiment of this application, a text processing method embodiment is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0168] Figure 4 This is a flowchart of a text processing method according to an embodiment of this application. Figure 4 As shown, the method may include the following steps:

[0169] Step S402: In response to a second input instruction applied to the video display page, input target text, wherein the video display page displays the target video.

[0170] In one alternative embodiment, the video display page can be a video website or a video application page, where users can input target text, such as a sentence, based on the video they want to watch. The system then finds the target video segment that corresponds to the input target text.

[0171] Step S404: In response to the text positioning instruction applied to the video display page, display the target video segment in the target video that matches the target text on the video display page. The target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the target video. The multi-layer semantic features are extracted from the multi-layer semantic text of the target text. The multi-layer semantic text is obtained by performing layered processing on the target text.

[0172] In the above embodiments of this application, the number of target texts is multiple. The method further includes: obtaining at least one target video segment in the target video that matches each target text to obtain multiple target video segments; determining the association information between the multiple target video segments based on the multiple target texts and / or target videos; and splicing the multiple target video segments based on the association information to obtain a spliced ​​video.

[0173] In one optional embodiment, at least one target video segment that matches each target text in the target video can be obtained to obtain multiple target video segments, the association information between the multiple target texts can be determined, and the multiple target video segments can be spliced ​​together according to the association information to obtain a spliced ​​video.

[0174] For example, multiple target texts can be elephant, tiger, lion, etc. Then, the association information between multiple target video clips can be determined to be animals. At this time, the target video clips corresponding to multiple target texts can be spliced ​​together to obtain a spliced ​​video about animals.

[0175] In another optional embodiment, at least one target video segment that matches each target text in the target video can be obtained to obtain multiple target video segments, the association information between the multiple target videos can be determined, and the multiple target video segments can be spliced ​​together according to the association information to obtain a spliced ​​video.

[0176] For example, multiple target videos can be multiple TV series. Then, the association information between multiple target video segments can be determined as TV series. At this time, multiple target video segments can be spliced ​​together according to TV series to obtain spliced ​​video related to TV series.

[0177] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0178] Example 3

[0179] According to an embodiment of this application, a text processing method embodiment is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0180] Figure 5 This is a flowchart of a text processing method according to an embodiment of this application. Figure 5 As shown, the method may include the following steps:

[0181] Step S502: Obtain entertainment videos from the entertainment playback platform and display the entertainment videos on the operation interface.

[0182] Step S504: In response to the third input command applied to the operation interface, input the target text.

[0183] Step S506: In response to the text positioning command applied to the operation interface, the target video segment that matches the target text in the entertainment video is displayed on the operation interface. The target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the entertainment video. The multi-layer semantic features are extracted from the multi-layer semantic text of the target text. The multi-layer semantic text is obtained by performing layered processing on the target text.

[0184] In the above embodiments of this application, the method further includes: sending the target video segment and multi-layer semantic text to an entertainment playback platform, wherein the multi-layer semantic text is used to annotate the layers of the target text on the entertainment playback platform.

[0185] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the scheme and implementation process provided in Embodiment 1, but are not limited to the scheme provided in Embodiment 1.

[0186] Example 4

[0187] According to embodiments of this application, a text processing apparatus for implementing the above-described text processing method is also provided, such as... Figure 6 As shown, the device 600 includes: a first input unit 602 and a first display unit 604.

[0188] The first input unit is used to respond to a first input command applied to the operation interface and input target text and target video; the first display unit is used to respond to a text positioning command applied to the operation interface and display a target video segment in the target video that matches the target text on the operation interface. The target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the target video. The multi-layer semantic features are extracted from the multi-layer semantic text of the target text. The multi-layer semantic text is obtained by performing layered processing on the target text.

[0189] It should be noted that the first input unit 602 and the first display unit 604 mentioned above correspond to steps S202 to S204 in Embodiment 1. The two units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0190] In the above embodiments of this application, the first display unit includes: a first display module.

[0191] The first display module is used to display the target video clip and multi-layered semantic text on the operation interface.

[0192] In the above embodiments of this application, the device further includes: a first labeling unit.

[0193] The first annotation unit is used to respond to annotation commands applied to the operation interface and to annotate the target text hierarchically based on multi-layer semantic text.

[0194] In the above embodiments of this application, the device further includes: a first display unit and a first selection unit.

[0195] The first display unit is used to display the first matching degree between each target video segment and the target text on the operation interface, thereby obtaining multiple first matching degrees; the first selection unit is used to respond to the selection command applied to the operation interface, select the first target video segment from multiple target video segments based on the multiple first matching degrees, and send the first target video segment to the server. The first target video segment is used by the server to adjust the target model, and the target model is used to process the multi-layer semantic features of the target text and the video features of the target video to obtain the target video segment.

[0196] In the above embodiments of this application, the device further includes a second display unit.

[0197] The second display unit is used to respond to the title addition command applied to the operation interface and display the title of each target video segment on the operation interface. The title is determined based on the target text and the target video.

[0198] In the above embodiments of this application, each layer of semantic text is a semantic text at a grammatical level and includes the main clause of the target text. The multi-layer semantic text includes a first layer of semantic text, a second layer of semantic text, and a third layer of semantic text. The completeness of the first layer of semantic text is lower than that of the second layer of semantic text, and the completeness of the second layer of semantic text is lower than that of the third layer of semantic text. The device also includes a first division unit.

[0199] The first division unit is used to respond to the semantic text generation command applied to the operation interface, divide the target text according to the target syntax, obtain the first statement, the second statement and the third statement of the target text, determine the first statement as the first layer of semantic text, merge the first statement and the second statement to obtain the second layer of semantic text, and merge the second layer of semantic text with the third statement to obtain the third layer of semantic text.

[0200] In the above embodiments of this application, the device further includes: a first determining unit.

[0201] The first determining unit is used to respond to the first model generation instruction applied to the operation interface, determine the first loss function based on the first sample set and / or the second sample set of video samples, and train the first fusion model based on the first loss function.

[0202] The first fusion model is used to fuse the first-level semantic features of the first-level semantic text and the temporal segment features of the video features to obtain the first sub-fusion feature, and is also used to fuse the second-level semantic features of the second-level semantic text and the temporal segment features to obtain the second sub-fusion feature. The first sub-fusion feature, the second sub-fusion feature, and the third sub-fusion feature are used to determine the target video segment in the target video. The third sub-fusion feature is the fusion feature between the third-level semantic features of the third-level semantic text and the temporal segment features.

[0203] In the above embodiments of this application, the first determining unit includes: a first determining module.

[0204] The first determining module is used to respond to the first model generation instruction applied to the operation interface, determine the second loss function based on the start and end times of the target video segment sample and the similarity between the text sample and the first or second sample set, and determine the first loss function based on the second loss function.

[0205] In the above embodiments of this application, the device further includes: a second determining unit.

[0206] The second determining unit is used to respond to the second model generation instruction applied to the operation interface, determine the third loss function based on semantic feature labels and temporal segment labels, and train the first initial model based on the third loss function to obtain the second fusion model. The second fusion model is used to fuse the third-layer semantic features and temporal segment features of the third-layer semantic text to obtain the third sub-fusion feature.

[0207] In the above embodiments of this application, the device further includes: a third display unit, a first sending unit, and a second annotation unit.

[0208] The third display unit is used to display the target start time and target end time of the target video segment on the operation interface; the first sending unit is used to send the target start time and target end time of the target video segment to the target client; and the second annotation unit is used to annotate the target video segment in the target video.

[0209] In the above embodiments of this application, the device further includes: a second determining unit, a first acquiring unit, and a third determining unit.

[0210] The second determining unit is used to determine multiple target candidate video segments in the target video that are associated with the fusion features between multi-layer semantic features and video features; the first obtaining unit is used to obtain the second matching degree between each target candidate video segment and the target text, thereby obtaining multiple second matching degrees; the third determining unit is used to determine the target candidate video segment corresponding to the highest matching degree among the multiple second matching degrees as the target video segment.

[0211] In the above embodiments of this application, the device further includes: a first encoding unit and a first extraction unit.

[0212] The first encoding unit is used to encode each layer of semantic text to obtain multiple encoding results, wherein each encoding result includes word vectors of each layer of semantic text; the first extraction unit is used to extract the semantic features of each layer from each encoding result.

[0213] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0214] Example 5

[0215] According to embodiments of this application, a text processing apparatus for implementing the above-described text processing method is also provided, such as... Figure 7 As shown, the device 700 includes: a second input unit 702 and a second display unit 704.

[0216] The second input unit is used to respond to a second input command applied to the video display page and input target text, wherein the video display page displays a target video; the second display unit is used to respond to a text positioning command applied to the video display page and display a target video segment in the target video that matches the target text, wherein the target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the target video, wherein the multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0217] It should be noted that the second input unit 702 and the second display unit 704 mentioned above correspond to steps S402 to S404 in Embodiment 2. The two units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0218] In the above embodiments of this application, the number of target texts is multiple, and the device further includes: a second acquisition unit, a fourth determination unit, and a first splicing unit.

[0219] The second acquisition unit is used to acquire at least one target video segment that matches each target text in the target video, thereby obtaining multiple target video segments; the fourth determination unit is used to determine the association information between multiple target video segments based on multiple target texts and / or target videos; and the first splicing unit is used to splice multiple target video segments based on the association information to obtain a spliced ​​video.

[0220] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0221] Example 6

[0222] According to embodiments of this application, a text processing apparatus for implementing the above-described text processing method is also provided, such as... Figure 8 As shown, the device 800 includes: a third acquisition unit 802, a third input unit 804, and a fourth display unit 806.

[0223] The third acquisition unit is used to acquire entertainment videos from the entertainment playback platform and display the entertainment videos on the operation interface; the third input unit is used to respond to the third input command applied to the operation interface and input the target text; the fourth display unit is used to respond to the text positioning command applied to the operation interface and display the target video segment in the entertainment video that matches the target text on the operation interface. The target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the entertainment video. The multi-layer semantic features are extracted from the multi-layer semantic text of the target text. The multi-layer semantic text is obtained by performing layered processing on the target text.

[0224] It should be noted that the third acquisition unit 802, the third input unit 804, and the fourth display unit 806 mentioned above correspond to steps S502 to S506 in Embodiment 3. Each unit and its corresponding step implements the same instance and application scenario, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules, as part of the device, can run on the computer terminal 10 provided in Embodiment 1.

[0225] In the above embodiments of this application, the device further includes a second transmitting unit.

[0226] The second sending unit is used to send the target video clip and multi-layer semantic text to the entertainment playback platform. The multi-layer semantic text is used to annotate the hierarchy of the target text on the entertainment playback platform.

[0227] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0228] Example 7

[0229] According to an embodiment of this application, a text processing system is also provided, comprising:

[0230] processor;

[0231] The memory, connected to the processor, is used to respond to a first input command applied to the operating interface, inputting target text and target video; and to respond to a text positioning command applied to the operating interface, displaying a target video segment in the target video that matches the target text on the operating interface, wherein the target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the target video, the multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0232] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0233] Example 8

[0234] Embodiments of this application may provide a computer terminal, which may be any computer terminal device in a group of computer terminals. Optionally, in this embodiment, the aforementioned computer terminal may also be replaced by a mobile terminal or other terminal device.

[0235] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.

[0236] In this embodiment, the computer terminal described above can execute the program code for the following steps in the text processing method: responding to a first input instruction applied to the operation interface, inputting target text and target video; responding to a text positioning instruction applied to the operation interface, displaying a target video segment in the target video that matches the target text on the operation interface, wherein the target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the target video, the multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0237] Optionally, Figure 9 This is a structural block diagram of a computer terminal according to an embodiment of this application. Figure 9 As shown, the computer terminal A may include one or more (only one is shown in the figure) processors 902 and memory 904.

[0238] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the text processing method and apparatus in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the aforementioned text processing method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0239] The processor can invoke information and application programs stored in memory via a transmission device to perform the following steps: responding to a first input instruction applied to the operation interface, inputting target text and target video; responding to a text positioning instruction applied to the operation interface, displaying a target video segment in the target video that matches the target text on the operation interface, wherein the target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the target video, the multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0240] Optionally, the processor may also execute program code that executes the following steps: displaying a target video segment in the target video that matches the target text on the operation interface, including: displaying the target video segment and multi-layer semantic text on the operation interface.

[0241] Optionally, the processor may also execute program code that performs the following steps: responding to annotation instructions applied to the user interface, and annotating the target text hierarchically based on multi-level semantic text.

[0242] Optionally, the processor may also execute program code for the following steps: when there are multiple target video segments, the method further includes: displaying a first matching degree between each target video segment and the target text on the operation interface to obtain multiple first matching degrees; responding to a selection command applied to the operation interface, selecting a first target video segment from the multiple target video segments based on the multiple first matching degrees, and sending the first target video segment to the server, wherein the first target video segment is used by the server to adjust the target model, and the target model is used to process the multi-layer semantic features of the target text and the video features of the target video to obtain the target video segment.

[0243] Optionally, the processor may also execute program code that performs the following steps: in response to a title addition instruction applied to the operation interface, displays the title of each target video segment on the operation interface, wherein the title is determined based on the target text and the target video.

[0244] Optionally, the processor may also execute program code that performs the following steps: responding to a semantic text generation instruction applied to the operating interface, dividing the target text according to the target syntax to obtain the first statement, the second statement and the third statement of the target text, determining the first statement as the first-level semantic text, merging the first statement and the second statement to obtain the second-level semantic text, and merging the second-level semantic text with the third statement to obtain the third-level semantic text.

[0245] Optionally, the processor may also execute program code that performs the following steps: responding to a first model generation instruction applied to the operation interface, determining a first loss function based on a first sample set and / or a second sample set of video samples, and training a first fusion model based on the first loss function; wherein the first fusion model is used to fuse the first-layer semantic features of the first-layer semantic text and the temporal segment features of the video features to obtain a first sub-fusion feature, and is used to fuse the second-layer semantic features of the second-layer semantic text and the temporal segment features to obtain a second sub-fusion feature, the first sub-fusion feature, the second sub-fusion feature and the third sub-fusion feature are used to determine the target video segment in the target video, and the third sub-fusion feature is the fusion feature between the third-layer semantic features of the third-layer semantic text and the temporal segment features.

[0246] Optionally, the processor may also execute program code that performs the following steps: in response to a first model generation instruction applied to the user interface, determines a second loss function based on the start and end times of the target video segment samples and the similarity between the text samples and the first or second sample set, and determines a first loss function based on the second loss function.

[0247] Optionally, the processor may also execute program code that performs the following steps: responding to a second model generation instruction applied to the operating interface, determining a third loss function based on semantic feature labels and temporal segment labels, and training a first initial model based on the third loss function to obtain a second fusion model, wherein the second fusion model is used to fuse the third-layer semantic features and temporal segment features of the third-layer semantic text to obtain a third sub-fusion feature.

[0248] Optionally, the processor may also execute program code that performs the following steps: displays the target start time and target end time of the target video segment on the operation interface; sends the target start time and target end time of the target video segment to the target client; and marks the target video segment in the target video.

[0249] Optionally, the processor may also execute program code that performs the following steps: identifying multiple target candidate video segments in the target video that are associated with the fusion features between multi-layer semantic features and video features; obtaining a second matching degree between each target candidate video segment and the target text to obtain multiple second matching degrees; and determining the target candidate video segment corresponding to the highest matching degree among the multiple second matching degrees as the target video segment.

[0250] Optionally, the processor may also execute program code that performs the following steps: encodes each layer of semantic text to obtain multiple encoding results, wherein each encoding result includes word vectors of each layer of semantic text; and extracts semantic features of each layer from each encoding result.

[0251] Those skilled in the art will understand that Figure 9 The structure shown is for illustrative purposes only. The computer terminal can also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, and other terminal devices. Figure 9 This does not limit the structure of the aforementioned electronic device. For example, computer terminal A may also include components that are more... Figure 9 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 9 The different configurations shown.

[0252] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0253] Example 9

[0254] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the text processing method provided in the above embodiments.

[0255] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0256] Optionally, the storage medium is configured to store program code for performing the following steps: in response to a first input instruction applied to the operation interface, inputting target text and target video; in response to a text positioning instruction applied to the operation interface, displaying a target video segment in the target video that matches the target text on the operation interface, wherein the target video segment is obtained based on multi-layer semantic features of the target text and video features of the target video, the multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text.

[0257] Optionally, the storage medium is further configured to store program code for performing the following steps: displaying a target video segment in the target video that matches the target text on the operation interface, including: displaying the target video segment and multi-layer semantic text on the operation interface.

[0258] Optionally, the aforementioned storage medium is also configured to store program code for performing the following steps: in response to annotation instructions applied to the operation interface, annotating the target text hierarchically based on multi-level semantic text.

[0259] Optionally, the storage medium is further configured to store program code for performing the following steps: when there are multiple target video segments, the method further includes: displaying a first matching degree between each target video segment and the target text on the operation interface to obtain multiple first matching degrees; responding to a selection instruction applied to the operation interface, selecting a first target video segment from the multiple target video segments based on the multiple first matching degrees, and sending the first target video segment to the server, wherein the first target video segment is used by the server to adjust the target model, and the target model is used to process the multi-layer semantic features of the target text and the video features of the target video to obtain the target video segment.

[0260] Optionally, the aforementioned storage medium is further configured to store program code for performing the following steps: in response to a title addition instruction applied to the operation interface, displaying a title for each target video segment on the operation interface, wherein the title is determined based on the target text and the target video.

[0261] Optionally, the aforementioned storage medium is further configured to store program code for performing the following steps: responding to a semantic text generation instruction applied to the operating interface, dividing the target text according to the target syntax to obtain a first statement, a second statement, and a third statement of the target text, determining the first statement as the first-level semantic text, merging the first statement and the second statement to obtain the second-level semantic text, and merging the second-level semantic text with the third statement to obtain the third-level semantic text.

[0262] Optionally, the aforementioned storage medium is further configured to store program code for performing the following steps: in response to a first model generation instruction applied to the operation interface, determining a first loss function based on a first sample set and / or a second sample set of video samples, and training a first fusion model based on the first loss function; wherein the first fusion model is used to fuse the first-layer semantic features of the first-layer semantic text and the temporal segment features of the video features to obtain a first sub-fusion feature, and is used to fuse the second-layer semantic features of the second-layer semantic text and the temporal segment features to obtain a second sub-fusion feature, the first sub-fusion feature, the second sub-fusion feature and the third sub-fusion feature are used to determine the target video segment in the target video, and the third sub-fusion feature is a fusion feature between the third-layer semantic features of the third-layer semantic text and the temporal segment features.

[0263] Optionally, the storage medium is further configured to store program code for performing the following steps: in response to a first model generation instruction applied to the user interface, determining a second loss function based on the start and end times of the target video segment samples and the similarity between the text samples and the first or second sample set, and determining a first loss function based on the second loss function.

[0264] Optionally, the aforementioned storage medium is further configured to store program code for performing the following steps: responding to a second model generation instruction applied to the operation interface, determining a third loss function based on semantic feature labels and temporal segment labels, and training a first initial model based on the third loss function to obtain a second fusion model, wherein the second fusion model is used to fuse the third-layer semantic features and temporal segment features of the third-layer semantic text to obtain a third sub-fusion feature.

[0265] Optionally, the storage medium is further configured to store program code for performing the following steps: displaying the target start time and target end time of the target video segment on the operation interface; sending the target start time and target end time of the target video segment to the target client; and marking the target video segment in the target video.

[0266] Optionally, the storage medium is further configured to store program code for performing the following steps: identifying multiple target candidate video segments in the target video that are associated with the fusion features between multi-layer semantic features and video features; obtaining a second matching degree between each target candidate video segment and the target text, thereby obtaining multiple second matching degrees; and determining the target candidate video segment corresponding to the highest matching degree among the multiple second matching degrees as the target video segment.

[0267] Optionally, the storage medium is further configured to store program code for performing the following steps: encoding each layer of semantic text to obtain multiple encoding results, wherein each encoding result includes word vectors of each layer of semantic text; and extracting semantic features of each layer from each encoding result.

[0268] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0269] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0270] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of units or modules may be electrical or other forms.

[0271] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0272] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0273] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0274] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A text processing method characterized by, include: Responding to the first input command applied to the user interface, input the target text and the target video; In response to a text location command applied to the operation interface, the target video segment in the target video that matches the target text is displayed on the operation interface; The target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the target video. The multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text. Each layer of semantic text is a semantic text at a grammatical level and includes the main sentence of the target text. The multiple layers of semantic text include a first layer of semantic text, a second layer of semantic text, and a third layer of semantic text. The completeness of the first layer of semantic text is lower than that of the second layer of semantic text, and the completeness of the second layer of semantic text is lower than that of the third layer of semantic text. The method further includes: determining the first statement of the target text as the first layer of semantic text; merging the first statement and the second statement of the target text to obtain the second layer of semantic text; and merging the second layer of semantic text and the third statement of the target text to obtain the third layer of semantic text.

2. The method according to claim 1, characterized in that, Displaying a target video segment that matches the target text in the target video on the operation interface includes: displaying the target video segment and multiple layers of the semantic text on the operation interface; and / or The method further includes: responding to annotation instructions applied to the operation interface, and annotating the target text hierarchically based on the multi-layer semantic text.

3. The method according to claim 1, In the case where the number of the target video segments is multiple, the method further comprises: The first matching degree between each target video segment and the target text is displayed on the operation interface, resulting in multiple first matching degrees; In response to a selection command applied to the user interface, a first target video segment is selected from multiple target video segments based on multiple first matching degrees, and the first target video segment is sent to the server. The first target video segment is used by the server to adjust a target model, and the target model is used to process the multi-layer semantic features of the target text and the video features of the target video to obtain the target video segment; and / or The method further includes: responding to a title addition instruction applied to the operation interface, displaying a title for each target video segment on the operation interface, wherein the title is determined based on the target text and the target video.

4. The method of claim 1, wherein, The method further includes: In response to a semantic text generation instruction applied to the operation interface, the target text is divided according to the target syntax to obtain the first statement, the second statement, and the third statement of the target text.

5. The method of claim 4, wherein, The method further includes: In response to a first model generation instruction applied to the operation interface, a first loss function is determined based on a first sample set and / or a second sample set of video samples, and a first fusion model is trained based on the first loss function. The first fusion model is used to fuse the first-layer semantic features of the first-layer semantic text and the temporal segment features of the video features to obtain a first sub-fusion feature, and is used to fuse the second-layer semantic features of the second-layer semantic text and the temporal segment features to obtain a second sub-fusion feature. The first sub-fusion feature, the second sub-fusion feature, and the third sub-fusion feature are used to determine the target video segment in the target video. The third sub-fusion feature is a fusion feature between the third-layer semantic features of the third-layer semantic text and the temporal segment features.

6. The method according to claim 5, characterized in that, In response to a first model generation instruction applied to the operation interface, determining a first loss function based on a first sample set and / or a second sample set of video samples includes: in response to a first model generation instruction applied to the operation interface, determining a second loss function based on the start and end times of the target video segment samples, the similarity between the text samples and the first sample set or the second sample set, and determining the first loss function based on the second loss function; The method further includes: responding to a second model generation instruction applied to the operation interface, determining a third loss function based on the semantic feature labels and the temporal segment labels, and training a first initial model based on the third loss function to obtain a second fusion model, wherein the second fusion model is used to fuse the third-layer semantic features of the third-layer semantic text and the temporal segment features to obtain the third sub-fusion feature.

7. The method of claim 1, wherein, The method further includes at least one of the following: The target start time and target end time of the target video segment are displayed on the operation interface; Send the target start time and target end time of the target video segment to the target client; The target video segment is marked in the target video.

8. The method of claim 1, wherein, The method further includes: In the target video, identify multiple target candidate video segments associated with the fusion features between the multi-layer semantic features and the video features; obtain a second matching degree between each target candidate video segment and the target text, resulting in multiple second matching degrees; determine the target candidate video segment corresponding to the highest matching degree among the multiple second matching degrees as the target video segment; and / or Encode the semantic text of each layer to obtain multiple encoding results, wherein each encoding result includes word vectors of the semantic text of each layer; extract the semantic features of each layer from each encoding result.

9. A text processing method characterized by, include: In response to a second input command applied to the video display page, target text is input, wherein the target video is displayed on the video display page; In response to a text location command applied to the video display page, a target video segment matching the target text in the target video is displayed on the video display page. The target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the target video. The multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text. Each layer of semantic text is a semantic text at a grammatical level and includes the main sentence of the target text. The multiple layers of semantic text include a first layer of semantic text, a second layer of semantic text, and a third layer of semantic text. The completeness of the first layer of semantic text is lower than that of the second layer of semantic text, and the completeness of the second layer of semantic text is lower than that of the third layer of semantic text. The first layer of semantic text is the first statement of the target text. The second layer of semantic text is obtained by merging the first statement and the second statement of the target text. The third layer of semantic text is obtained by merging the second layer of semantic text and the third statement of the target text.

10. The method of claim 9, wherein, The number of target texts is multiple, and the method further includes: Obtain at least one target video segment from the target video that matches each of the target texts, thereby obtaining multiple target video segments; Based on multiple target texts and / or target videos, determine the association information between multiple target video segments; Based on the association information, multiple target video segments are spliced ​​together to obtain a spliced ​​video.

11. A text processing method characterized by, include: Obtain entertainment videos from entertainment streaming platforms and display the entertainment videos on the user interface; In response to a third input command applied to the user interface, input the target text; In response to a text location command applied to the operation interface, a target video segment matching the target text in the entertainment video is displayed on the operation interface. The target video segment is obtained based on the multi-layer semantic features of the target text and the video features of the entertainment video. The multi-layer semantic features are extracted from the multi-layer semantic text of the target text, and the multi-layer semantic text is obtained by performing layered processing on the target text. Each layer of semantic text is a semantic text at a grammatical level and includes the main sentence of the target text. The multiple layers of semantic text include a first layer of semantic text, a second layer of semantic text, and a third layer of semantic text. The completeness of the first layer of semantic text is lower than that of the second layer of semantic text, and the completeness of the second layer of semantic text is lower than that of the third layer of semantic text. The first layer of semantic text is the first statement of the target text. The second layer of semantic text is obtained by merging the first statement and the second statement of the target text. The third layer of semantic text is obtained by merging the second layer of semantic text and the third statement of the target text.

12. The method of claim 11, wherein, The method further includes: The target video segment and the multi-layer semantic text are sent to the entertainment playing platform, wherein the multi-layer semantic text is used for marking the level of the target text on the entertainment playing platform.