Data matching method and device, product and equipment
By performing fine-grained segmentation and feature encoding matching of film and television drama videos and novel content, the problem of accuracy in matching film and television drama videos with novel content has been solved, enabling accurate display of novel content during video playback and improving user experience.
Patent Information
- Application Number
- CN202511041083.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-07
AI Technical Summary
How to accurately match the corresponding novel content to each episode of a TV series in order to provide richer user interactive content.
By segmenting video resources and text materials, multiple sub-videos and sub-texts are generated. Then, a feature encoding model is used to perform content matching between the sub-videos and sub-texts to generate matching pairs. Finally, the range of text materials that match the target video resources is determined.
It enables accurate matching and display of the corresponding novel content range during the playback of film and television drama videos, thereby improving the user interaction experience.
Smart Images

Figure CN120910307A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data matching, and in particular to a data matching method, device, product and equipment. BACKGROUND
[0002] A TV series is often obtained by shooting a novel, and the TV series can include multiple video episodes, each of which is obtained by shooting different parts of the novel content. Therefore, how to determine the respective novel content in the novel that each video episode in the TV series matches, so as to provide more novel interactive content for users in the playing scene of the TV series is a hot issue, and among them, how to accurately match the respective novel content in the novel for each video episode in the TV series is a key issue. SUMMARY
[0003] The present application provides a data matching method, device, product and equipment, which can improve the accuracy of the text range determined in the text material that matches the target video resource.
[0004] In one aspect, the present application provides a data matching method, which comprises:
[0005] Obtaining a text material associated with a video resource set and a target video resource, the text content described by the text material matching the video content presented by the video resource set, and the target video resource being any video resource in the video resource set;
[0006] Performing video segmentation processing on the target video resource to obtain multiple sub-videos of the target video resource, and performing text segmentation processing on the text material to obtain multiple sub-texts of the text material;
[0007] Performing content matching processing between the multiple sub-videos and the multiple sub-texts to obtain multiple matching pairs, each matching pair including a sub-video and a sub-text, and the sub-video and the sub-text in the same matching pair having content matching;
[0008] Determining a text range in the text material that matches the target video resource according to the multiple matching pairs.
[0009] In one embodiment, the text material includes multiple text chapters; the text segmentation processing on the text material to obtain the multiple sub-texts of the text material comprises:
[0010] Detecting a segmentation position in each text chapter of the text material;
[0011] Respectively performing text segmentation processing on each text chapter according to the detected segmentation position in each text chapter to obtain the multiple sub-texts;
[0012] Any subtext is any paragraph obtained by segmenting any text chapter of the text material.
[0013] In an implementation, content matching processing is performed between the plurality of sub-videos and the plurality of sub-texts to obtain a plurality of matching pairs, including:
[0014] The plurality of sub-videos are respectively subjected to feature encoding processing to generate video encoding features of each sub-video, and the video encoding features of any sub-video are used to represent the video content presented by any sub-video.
[0015] The plurality of sub-texts are respectively subjected to feature encoding processing to generate text encoding features of each sub-text, and the text encoding features of any sub-text are used to represent the text content described by any sub-text.
[0016] Based on the video encoding features of each sub-video and the text encoding features of each sub-text, content matching processing is performed between the plurality of sub-videos and the plurality of sub-texts to obtain a plurality of matching pairs.
[0017] In an implementation, any sub-video is a target sub-video; the plurality of sub-videos are respectively subjected to feature encoding processing to generate video encoding features of each sub-video, including:
[0018] The video frames in the target sub-video are subjected to extraction processing to obtain a plurality of extracted video frames in the target sub-video.
[0019] The plurality of extracted video frames are subjected to text extraction processing from a plurality of first text extraction dimensions to generate extracted texts of the target sub-video under each first text extraction dimension.
[0020] The extracted texts of the target sub-video under the plurality of first text extraction dimensions are subjected to feature encoding processing to generate video encoding features of the target sub-video.
[0021] In an implementation, the extracted texts of the target sub-video under the plurality of first text extraction dimensions are subjected to feature encoding processing to generate video encoding features of the target sub-video, including:
[0022] A feature encoding model is obtained, and the feature encoding model includes a feature encoding branch corresponding to each first text extraction dimension;
[0023] Each feature encoding branch is called to perform feature encoding processing on the extracted texts of the target sub-video under the corresponding first text extraction dimension to generate first text encoding features of the extracted texts of the target sub-video under each first text extraction dimension;
[0024] The plurality of first text encoding features of the extracted texts of the target sub-video under the plurality of first text extraction dimensions are subjected to feature fusion processing to generate video encoding features of the target sub-video.
[0025] In an implementation, any of the subtexts is a target subtext; the multiple subtexts are respectively subjected to feature coding processing to generate text coding features of each subtext, including:
[0026] The target subtext is subjected to text extraction processing from the multiple second text extraction dimensions to generate extracted texts of the target subtext under each second text extraction dimension;
[0027] The extracted texts of the target subtext under the multiple second text extraction dimensions are subjected to feature coding processing to generate text coding features of the target subtext.
[0028] In an implementation, the multiple second text extraction dimensions include a content extraction dimension, a dialogue extraction dimension, and a role extraction dimension; the target subtext is subjected to text extraction processing from the multiple second text extraction dimensions to generate extracted texts of the target subtext under each second text extraction dimension, including:
[0029] The target subtext is subjected to content extraction processing from the content extraction dimension to obtain content extracted texts of the target subtext;
[0030] The target subtext is subjected to dialogue extraction processing from the dialogue extraction dimension to obtain dialogue texts in the target subtext;
[0031] The target subtext is subjected to role extraction processing from the role extraction dimension to obtain second role identifiers of film and television roles described in the target subtext;
[0032] The extracted texts of the target subtext under the multiple second text extraction dimensions include the content extracted texts, the dialogue texts, and the second role identifiers.
[0033] In an implementation, the target subtext is subjected to content extraction processing from the content extraction dimension to obtain content extracted texts of the target subtext, including:
[0034] The text length of the target subtext is obtained;
[0035] If the text length is less than or equal to a set length threshold, the target subtext is taken as the content extracted texts;
[0036] If the text length is greater than the length threshold, the target subtext is subjected to content compression processing to generate the content extracted texts.
[0037] In an implementation, the extracted texts of the target subtext under the multiple second text extraction dimensions are subjected to feature coding processing to generate text coding features of the target subtext, including:
[0038] obtain a feature coding model, the feature coding model comprising a feature coding branch corresponding to each second text extraction dimension;
[0039] call each feature coding branch to perform feature coding processing on the extracted text of the target subtext in the corresponding second text extraction dimension, to generate a second text coding feature of the extracted text of the target subtext in each second text extraction dimension;
[0040] perform feature fusion processing on the plurality of second text coding features of the extracted text of the target subtext in the plurality of second text extraction dimensions, to generate a text coding feature of the target subtext.
[0041] In an embodiment, the video coding features of the plurality of sub-videos and the text coding features of the plurality of sub-texts are both generated by calling the feature coding model; the method further comprises:
[0042] obtain a feature coding model to be trained and a sample set, the sample set comprising positive sample pairs and negative sample pairs, the positive sample pairs comprising first sample texts and first sample videos, the text content described by the first sample texts matching the video content presented by the first sample videos, the negative sample pairs comprising second sample texts and second sample videos, the text content described by the second sample texts not matching the video content presented by the second sample videos;
[0043] call the feature coding model to be trained to perform feature coding processing on the positive sample pairs and the negative sample pairs respectively, to generate first sample text coding features of the first sample texts, first sample video coding features of the first sample videos, second sample text coding features of the second sample texts, and second sample video coding features of the second sample videos;
[0044] generate a first feature coding loss of the feature coding model to be trained for the positive sample pairs based on the first sample text coding features and the first sample video coding features, and generate a second feature coding loss of the feature coding model to be trained for the negative sample pairs based on the second sample text coding features and the second sample video coding features;
[0045] correct the model parameters of the feature coding model to be trained using the first feature coding loss and the second feature coding loss, to obtain the feature coding model.
[0046] In an embodiment, the model parameters of the feature coding model to be trained are corrected using the first feature coding loss and the second feature coding loss to obtain the feature coding model, comprising:
[0047] obtain a first training weight corresponding to the positive sample pairs and a second training weight corresponding to the negative sample pairs, the first training weight being greater than the second training weight;
[0048] The first feature encoding loss and the second feature encoding loss are weighted and summed by using the first training weight and the second training weight to obtain a total encoding loss of the feature encoding model to be trained for the positive sample pair and the negative sample pair;
[0049] The model parameters of the feature encoding model to be trained are corrected by using the total encoding loss to obtain the feature encoding model.
[0050] In an implementation, the method further includes:
[0051] When playing the target video resource in the video client, the reading link corresponding to the text range matched with the target video resource is pushed to the video client;
[0052] The video client is configured to display the reading link when playing the target video resource, and configured to display the text in the text range matched with the target video resource in the text material in response to a trigger operation on the reading link.
[0053] The present application provides a data matching device, which includes:
[0054] The acquisition module is configured to acquire a text material associated with a video resource set and a target video resource, the text content described by the text material being matched with the video content presented by the video resource set, and the target video resource being any video resource in the video resource set;
[0055] The segmentation module is configured to perform video segmentation processing on the target video resource to obtain a plurality of sub-videos of the target video resource, and perform text segmentation processing on the text material to obtain a plurality of sub-texts of the text material;
[0056] The matching module is configured to perform content matching processing between the plurality of sub-videos and the plurality of sub-texts to obtain a plurality of matching pairs, one matching pair including one sub-video and one sub-text, and the sub-video and the sub-text in the same matching pair having content matching property;
[0057] The determination module is configured to determine a text range in the text material matched with the target video resource according to the plurality of matching pairs.
[0058] In an implementation, the segmentation module performs video segmentation processing on the target video resource to obtain the plurality of sub-videos of the target video resource in the following manner:
[0059] The segmentation module performs scene detection processing on the video frames in the target video resource to obtain at least one transition video frame in the target video resource, one transition video frame being used to indicate that the video scene in the target video resource changes;
[0060] According to the position of at least one transition video frame in the target video resource, the target video resource is video segmented to obtain a plurality of sub-videos.
[0061] In an implementation, the text material includes a plurality of text chapters; and the text segmentation module segments the text material to obtain a plurality of sub-texts in the following manner:
[0062] detecting a segmentation position in each text chapter of the text material;
[0063] segmenting each text chapter according to the detected segmentation position in the text chapter to obtain a plurality of sub-texts;
[0064] wherein any sub-text is any paragraph segmented from any text chapter of the text material.
[0065] In an implementation, the matching module matches the content between the plurality of sub-videos and the plurality of sub-texts to obtain a plurality of matching pairs in the following manner:
[0066] performing feature encoding on the plurality of sub-videos to generate video encoding features of each sub-video, wherein the video encoding features of any sub-video represent the video content presented by the sub-video;
[0067] performing feature encoding on the plurality of sub-texts to generate text encoding features of each sub-text, wherein the text encoding features of any sub-text represent the text content described by the sub-text;
[0068] matching the content between the plurality of sub-videos and the plurality of sub-texts based on the video encoding features of each sub-video and the text encoding features of each sub-text to obtain a plurality of matching pairs.
[0069] In an implementation, any sub-video is a target sub-video; and the matching module performs feature encoding on the plurality of sub-videos to generate video encoding features of each sub-video in the following manner:
[0070] extracting video frames in the target sub-video to obtain a plurality of extracted video frames in the target sub-video;
[0071] performing text extraction on the plurality of extracted video frames from a plurality of first text extraction dimensions to generate extracted texts of the target sub-video in each first text extraction dimension;
[0072] performing feature encoding on the extracted texts of the target sub-video in the plurality of first text extraction dimensions to generate video encoding features of the target sub-video.
[0073] In an implementation, the plurality of first text extraction dimensions comprises a picture extraction dimension, a character detection dimension, and a content understanding dimension; the matching module performs text extraction processing on the plurality of extracted video frames from the plurality of first text extraction dimensions, and generates the extracted text of the target sub-video under each first text extraction dimension in the following manner, which comprises:
[0074] performing text extraction processing on the text appearing in the pictures of the plurality of extracted video frames from the picture extraction dimension to obtain picture extraction text in the plurality of extracted video frames;
[0075] performing face detection processing on the plurality of extracted video frames from the character detection dimension to obtain a first character identifier of the characters in the plurality of extracted video frames;
[0076] performing content understanding processing on the plurality of extracted video frames from the content understanding dimension to generate a content summary text of the video content presented by the plurality of extracted video frames;
[0077] The extracted text of the target sub-video under the plurality of first text extraction dimensions comprises the picture extraction text, the first character identifier, and the content summary text.
[0078] In an implementation, the matching module performs feature encoding processing on the extracted text of the target sub-video under the plurality of first text extraction dimensions to generate video encoding features of the target sub-video in the following manner, which comprises:
[0079] obtaining a feature encoding model, the feature encoding model comprising a feature encoding branch corresponding to each first text extraction dimension;
[0080] calling each feature encoding branch to perform feature encoding processing on the extracted text of the target sub-video under the corresponding first text extraction dimension to generate first text encoding features of the extracted text of the target sub-video under each first text extraction dimension;
[0081] performing feature fusion processing on the plurality of first text encoding features of the extracted text of the target sub-video under the plurality of first text extraction dimensions to generate the video encoding features of the target sub-video.
[0082] In an implementation, any sub-text is a target sub-text; the matching module performs feature encoding processing on the plurality of sub-texts to generate text encoding features of each sub-text in the following manner, which comprises:
[0083] performing text extraction processing on the target sub-text from the plurality of second text extraction dimensions to generate extracted text of the target sub-text under each second text extraction dimension;
[0084] performing feature encoding processing on the extracted text of the target sub-text under the plurality of second text extraction dimensions to generate text encoding features of the target sub-text.
[0085] In an implementation, the plurality of second text extraction dimensions comprises a content extraction dimension, a dialogue extraction dimension, and a character extraction dimension; the matching module performs text extraction processing on the target subtext from the plurality of second text extraction dimensions in the following manner:
[0086] performing content extraction processing on the target subtext from the content extraction dimension to obtain content extraction text of the target subtext;
[0087] performing dialogue extraction processing on the target subtext from the dialogue extraction dimension to obtain dialogue text in the target subtext;
[0088] performing character extraction processing on the target subtext from the character extraction dimension to obtain a second character identifier of a film character described in the target subtext;
[0089] wherein the extraction text of the target subtext in the plurality of second text extraction dimensions comprises the content extraction text, the dialogue text, and the second character identifier.
[0090] In an implementation, the matching module performs content extraction processing on the target subtext from the content extraction dimension in the following manner:
[0091] obtaining a text length of the target subtext;
[0092] if the text length is less than or equal to a set length threshold, regarding the target subtext as the content extraction text;
[0093] if the text length is greater than the length threshold, performing content compression processing on the target subtext to generate the content extraction text.
[0094] In an implementation, the matching module performs feature encoding processing on the extraction text of the target subtext in the plurality of second text extraction dimensions in the following manner to generate text encoding features of the target subtext:
[0095] obtaining a feature encoding model, the feature encoding model comprising a feature encoding branch corresponding to each second text extraction dimension;
[0096] calling each feature encoding branch to perform feature encoding processing on the extraction text of the target subtext in the corresponding second text extraction dimension to generate second text encoding features of the extraction text of the target subtext in each second text extraction dimension;
[0097] performing feature fusion processing on the plurality of second text encoding features of the extraction text of the target subtext in the plurality of second text extraction dimensions to generate the text encoding features of the target subtext.
[0098] In an implementation, any of the sub-videos is a target sub-video; the matching module performs content matching processing between the plurality of sub-videos and the plurality of sub-texts based on the video encoding features of each sub-video and the text encoding features of each sub-text, to obtain a plurality of matching pairs in the following manner:
[0099] obtaining feature similarities between the text encoding features of the plurality of sub-texts and the video encoding features of the target sub-video, respectively;
[0100] performing sorting processing on the plurality of sub-texts in descending order of the feature similarities between the text encoding features of each sub-text and the video encoding features of the target sub-video, to obtain sorted sub-texts;
[0101] determining K sub-texts arranged in the front of the sorted sub-texts as first candidate sub-texts, K being a positive integer;
[0102] determining a sub-text matched with the target sub-video from the K first candidate sub-texts;
[0103] wherein the target sub-video and the matched sub-text are used to constitute a matching pair.
[0104] In an implementation, the matching module determines the sub-text matched with the target sub-video from the K first candidate sub-texts in the following manner:
[0105] obtaining a first role identifier of a movie / TV role appearing in the target sub-video;
[0106] determining a first candidate sub-text including the first role identifier in the K first candidate sub-texts as a second candidate sub-text;
[0107] determining a sub-text matched with the target sub-video from at least one second candidate sub-text.
[0108] In an implementation, the matching module determines the sub-text matched with the target sub-video from at least one second candidate sub-text in the following manner:
[0109] calculating a comprehensive matching degree between each second candidate sub-text and the target sub-video, respectively;
[0110] determining a second candidate sub-text having the largest comprehensive matching degree between the target video from the at least one second candidate sub-text as the sub-text matched with the target sub-video.
[0111] In an implementation, the matching module calculates the comprehensive matching degree between each second candidate sub-text and the target sub-video, respectively, in the following manner:
[0112] obtain a feature similarity between the text encoding feature of each second candidate subtext and the video encoding feature of the target subvideo;
[0113] calculate a co-occurrence indication parameter between the text keyword associated with each second candidate subtext and the text keyword associated with the target subvideo;
[0114] weight and sum the feature similarity and the co-occurrence indication parameter between each second candidate subtext and the target subvideo to obtain a comprehensive matching degree between each second candidate subtext and the target subvideo.
[0115] In an implementation, the video encoding features of the plurality of subvideos and the text encoding features of the plurality of subtexts are generated by calling a feature encoding model; and the matching module is further configured to:
[0116] obtain the feature encoding model to be trained and a sample set, the sample set including positive sample pairs and negative sample pairs, the positive sample pairs including first sample texts and first sample videos, the text content described by the first sample texts matching the video content presented by the first sample videos, the negative sample pairs including second sample texts and second sample videos, the text content described by the second sample texts not matching the video content presented by the second sample videos;
[0117] call the feature encoding model to be trained to perform feature encoding processing on the positive sample pairs and the negative sample pairs respectively to generate first sample text encoding features of the first sample texts, first sample video encoding features of the first sample videos, second sample text encoding features of the second sample texts, and second sample video encoding features of the second sample videos;
[0118] generate a first feature encoding loss of the feature encoding model to be trained for the positive sample pairs based on the first sample text encoding features and the first sample video encoding features, and generate a second feature encoding loss of the feature encoding model to be trained for the negative sample pairs based on the second sample text encoding features and the second sample video encoding features;
[0119] correct the model parameters of the feature encoding model to be trained by using the first feature encoding loss and the second feature encoding loss to obtain the feature encoding model.
[0120] In an implementation, the matching module corrects the model parameters of the feature encoding model to be trained by using the first feature encoding loss and the second feature encoding loss to obtain the feature encoding model in the following manner:
[0121] obtain a first training weight corresponding to the positive sample pairs and a second training weight corresponding to the negative sample pairs, the first training weight being greater than the second training weight;
[0122] The first feature encoding loss and the second feature encoding loss are weighted and summed by using the first training weight and the second training weight to obtain a total encoding loss of the feature encoding model to be trained for the positive sample pair and the negative sample pair;
[0123] The model parameters of the feature encoding model to be trained are corrected by using the total encoding loss to obtain the feature encoding model.
[0124] In an implementation, the text material includes a plurality of text chapters; and the manner in which the determining module determines the text range in the text material that matches the target video resource according to the plurality of matching pairs includes:
[0125] The Z text chapters in the text material in which the subtexts in the plurality of matching pairs are hit and the number of hits of the subtexts in each hit text chapter are counted, any hit text chapter includes at least one subtext in the plurality of matching pairs, and Z is a positive integer;
[0126] Based on the respective number of hits of each hit text chapter, a point of divergence detection process is performed on each hit text chapter;
[0127] The hit text chapters that are detected as points of divergence in the Z hit text chapters are removed to obtain at least one retained text chapter;
[0128] The chapter range of the at least one retained text chapter in the text material is determined as the text range that matches the target video resource.
[0129] In an implementation, any hit text chapter is a target text chapter; and the manner in which the determining module performs a point of divergence detection process on each hit text chapter based on the respective number of hits of each hit text chapter includes:
[0130] A first hit density between the subtexts hit in the Z hit text chapters is calculated, and a density reference threshold for matching the subtexts is obtained based on the first hit density;
[0131] A second hit density between the subtexts hit in Z-1 hit text chapters other than the target text chapter in the Z hit text chapters is calculated;
[0132] The first number of the subtexts hit in the Z-1 hit text chapters and the second number of the subtexts hit in the Z hit text chapters are counted, and a number ratio between the first number and the second number is obtained;
[0133] If the number ratio is greater than or equal to a preset ratio threshold and the second hit density is greater than or equal to the density reference threshold, the target text chapter is determined to be a point of divergence;
[0134] If the quantity ratio is less than the ratio threshold, or the second hit density is less than the density reference threshold, it is determined that the target text chapter is detected as a non-divergence point.
[0135] In an embodiment, the determining module calculates the first hit density between the hit sub-texts in the Z hit text chapters in the following manner:
[0136] The missing text chapters are filled in between the Z hit text chapters to obtain L continuous text chapters corresponding to the Z hit text chapters, L being a positive integer and L being greater than or equal to Z;
[0137] The ratio between the total number of the hit sub-texts in the Z hit text chapters and L is calculated as the first hit density.
[0138] In an embodiment, the determining module obtains the density reference threshold for the sub-text matching based on the first hit density in the following manner:
[0139] The ratio between the total number of the hit sub-texts in the Z hit text chapters and Z is calculated as the maximum hit density;
[0140] The difference between the maximum hit density and the first hit density is obtained, and the product between the difference and a set density control parameter is calculated as a density fluctuation value;
[0141] The sum of the density fluctuation value and the first hit density is taken as the density reference threshold.
[0142] In an embodiment, the determining module is further configured to:
[0143] When playing the target video resource in the video client, the reading link corresponding to the text range matched with the target video resource is pushed to the video client;
[0144] The video client is configured to display the reading link when playing the target video resource, and configured to display the text in the text range matched with the target video resource in the text material in response to a trigger operation on the reading link.
[0145] In an aspect, the present application provides a computer device, comprising a memory and a processor, the memory storing a computer program, the computer program being executed by the processor to make the processor execute the method in the aspect of the present application.
[0146] In an aspect, the present application provides a computer readable storage medium, the computer readable storage medium storing a computer program, the computer program being executed by the processor to make the processor execute the method in the aspect.
[0147] In an aspect, the present application provides a computer program product, which comprises a computer program stored in a computer readable storage medium. A processor of a computer device reads the computer program from the computer readable storage medium, and the processor executes the computer program, so that the computer device executes the method provided in the above aspect or various optional manners.
[0148] The present application can obtain a text material associated with a video resource set and a target video resource, the text content described by the text material matches the video content presented by the video resource set, and the target video resource is any video resource in the video resource set; and the target video resource can be subjected to video segmentation processing to obtain multiple sub-videos of the target video resource, and the text material can be subjected to text segmentation processing to obtain multiple sub-texts of the text material; thus, content matching processing can be performed between the multiple sub-videos and the multiple sub-texts to obtain multiple matching pairs, one matching pair comprising one sub-video and one sub-text, and the sub-video and the sub-text in the same matching pair having content matching; thus, the text range in the text material that matches the target video resource can be determined according to the multiple matching pairs. As can be seen, the method provided by the present application can segment the target video resource into multiple sub-videos of smaller units, and can also segment the text material associated with the video resource set into multiple sub-texts of smaller units, so that accurate determination of the text range in the text material that matches the target video resource can be achieved through refined content matching between the multiple sub-videos and the multiple sub-texts of smaller units. BRIEF DESCRIPTION OF DRAWINGS
[0149] In order to more clearly illustrate the technical solutions of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0150] Figure 1 is a structural schematic diagram of a network architecture of a matching network for text and video provided by an embodiment of the present application;
[0151] Figure 2 is a scene schematic diagram of content matching processing of a text material and a video resource set provided by an embodiment of the present application;
[0152] Figure 3 is a flow schematic diagram of a data matching method provided by an embodiment of the present application;
[0153] Figure 4 is a principle schematic diagram of obtaining a text range that matches a target video resource provided by an embodiment of the present application;
[0154] Figure 5a is an interface schematic diagram of a video playing interface provided by an embodiment of the present application;
[0155] Figure 5b is an interface schematic diagram of a text reading interface provided by an embodiment of the present application;
[0156] Figure 6 is an interface schematic diagram of another video playing interface provided by an embodiment of the present application;
[0157] Figure 7 is a flow schematic diagram of content matching processing between a sub-video and a sub-text provided by an embodiment of the present application;
[0158] Figure 8 is a structure schematic diagram of a feature encoding model provided by an embodiment of the present application;
[0159] Figure 9 is a principle schematic diagram of obtaining a feature encoding loss of a feature encoding model to be trained provided by an embodiment of the present application;
[0160] Figure 10 is a structure schematic diagram of a data matching framework provided by an embodiment of the present application;
[0161] Figure 11 is a structure schematic diagram of a data matching apparatus provided by an embodiment of the present application;
[0162] Figure 12 is a structure schematic diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0163] The technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0164] All the data collected by the present application (such as video resource set, text material and other related data) are collected with the consent and authorization of the owner of the data (such as user, institution or enterprise), and the collection, use and processing of the related data need to comply with the relevant laws, regulations and standards of the relevant region.
[0165] Here, the related technical concepts involved in the present application are explained:
[0166] Video scene in a video: refers to a narrative unit with continuous space-time relationship, that is, with space-time unity, which can include time continuity (such as event development without interruption (such as dialogue process)) and space unity (such as consistent background environment (such as fixed indoor scenery or a specific outdoor scene)).
[0167] TF-IDF (Term Frequency-Inverse Document Frequency): a statistical method for evaluating the importance of a word to a document in a document set or corpus. The importance increases with the increase of the frequency of the word in the document, but decreases with the decrease of the frequency of the word in the entire corpus. The TF-IDF method can be used to extract keywords in the present application.
[0168] Multimodal Large Language Model (MLLM): an artificial intelligence model that integrates text, image, audio and other data types for joint training and inference. It is essentially a general artificial intelligence framework that extends the cross-modal understanding and generation capabilities based on large language models (LLM). It processes heterogeneous data (such as text, image, speech, etc.) through a unified architecture, achieving cross-modal semantic alignment and collaborative reasoning. It is essentially a model paradigm that deeply integrates the text processing capabilities of LLM with non-text modal information.
[0169] See Figure 1 , Figure 1 is a structural diagram of a network architecture of a matching network for text and video provided by an embodiment of the present application. As Figure 1 indicated, the network architecture can include a server 200 and a terminal device cluster, which can include one or more terminal devices, and the number of terminal devices will not be limited here. As Figure 1 indicated, the plurality of terminal devices can specifically include terminal device 1, terminal device 2, terminal device 3, …, terminal device n; as Figure 1 indicated, terminal device 1, terminal device 2, terminal device 3, …, terminal device n can all be connected to the server 200 in a network, so that each terminal device can interact with the server 200 through network connection.
[0170] As Figure 1The server 200 shown can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (content distribution network), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart television, a vehicle-mounted terminal, a smart home, and the like. The following describes the embodiments of the present application by taking the communication between the terminal device 1 and the server 200 as an example.
[0171] The terminal device 1 described above can include a video client, which can be used for video playing. The server 200 can be a background server of the video client, and the server 200 can obtain a text material whose text content matches (e.g., is consistent or nearly consistent with) the video content presented by a video resource set, such as a multi-episode TV series, and the text material can be a novel of the TV series. The server 200 can obtain a text range in the text material that matches each video resource in the video resource set, so that when playing any video resource in the video resource set on the video client, a reading control corresponding to the text range matching the video resource can also be displayed, and the user can click the reading control. The video client can display (e.g., in a fixed area in the video playing interface or in another reading interface) the text in the played text range matching the video resource for the user to read and view.
[0172] Please refer to Figure 2 , Figure 2 is a scene schematic diagram provided by the embodiments of the present application for content matching processing of a text material and a video resource set. As shown in Figure 2 , the server 200 in the above Figure 1 can divide the text material into multiple subtexts, and can also divide a target video resource (which can be any video resource) in the video resource set into multiple sub-videos. The server 200 can perform content matching processing between the multiple subtexts and the multiple sub-videos to obtain multiple matching pairs. One matching pair includes (i.e., corresponds to) one sub-video of the target video resource. The number of matching pairs can be the same as the number of sub-videos of the target video resource. One matching pair can also include a subtext matched to have content matching with the corresponding sub-video. Thus, the server 200 can determine a text range in the entire text material that matches the target video resource through the multiple matching pairs.
[0173] By adopting the method provided in the present application, the text material and the video resources in the video resource set can be divided into smaller granularity data (such as sub-text and sub-video), so that more accurate determination of the text range matched with the video resources can be realized through content matching between the smaller granularity data.
[0174] Please refer to Figure 3 , Figure 3 is a flowchart of a data matching method provided by an embodiment of the present application. The execution subject in the embodiment of the present application can be a data matching device (which can be referred to as a matching device), which can be a computer device or a computer device cluster composed of multiple computer devices. The computer device can be a server or a terminal device, or other devices, which are not limited herein. As shown in Figure 3 , the method can include the following steps.
[0175] In step S101, a text material associated with a video resource set and a target video resource are obtained. The text content described by the text material is matched with the video content presented by the video resource set. The target video resource is any video resource in the video resource set.
[0176] Specifically, the video resource set can include multiple video resources, and one video resource can be a video. The text material associated with the video resource set can be source material used to shoot or produce the video resource set. In other words, the video resource set can be shot or produced based on the text content of the text material. That is, the text content described by the text material is matched with the video content presented by the video resource set, such as consistent or nearly consistent.
[0177] For example, the video resource set can be a TV series, which can include multiple episode videos. One episode video can be a video of one episode of the TV series. One video resource in the video resource set can be an episode video of the TV series. On this basis, the text material associated with the video resource set can be a novel material associated with the TV series, such as the TV series being shot or adapted from the novel material.
[0178] In an implementation manner, the video resource in the video resource set can be a video shot by a real actor based on the text material, or the video resource in the video resource set can also be an animation video produced based on the text material. And so on. The specific type of the video resource in the video resource set can be determined according to the actual application scenario, which is not limited herein.
[0179] The matching device can obtain the text material associated with the video resource set and the target video resource. The text content described in the text material can be matched (e.g., consistent or nearly consistent) with the video content presented by the video resource set. The target video resource can be any video resource in the video resource set. The following process of the present application can be described in detail by taking the determination of the text range matched with the target video resource from the text material as an example. It can be known that the principle of determining the text range matched with each video resource in the video resource set from the text material is the same.
[0180] In step S102, the target video resource is subjected to video segmentation processing to obtain a plurality of sub-videos of the target video resource, and the text material is subjected to text segmentation processing to obtain a plurality of sub-texts of the text material.
[0181] Specifically, the matching device can perform video segmentation processing on the target video resource to obtain a plurality of sub-videos of the target video resource, i.e., the target video resource can be segmented into the plurality of sub-videos. Subsequently, the corresponding text range can be more accurately matched for the target video resource from the text material in the unit of the sub-video.
[0182] In an embodiment, the process of performing video segmentation processing on the target video resource to obtain a plurality of sub-videos of the target video resource can include that the target video resource can include a plurality of video frames, the matching device can perform scene detection processing on the video frames in the target video resource to obtain at least one transition video frame in the target video resource. One transition video frame can be used to indicate that the video scene in the target video resource changes (i.e., changes or switches).
[0183] For example, the matching device can call a pre-trained scene detection model to perform scene detection processing on the video frames in the target video resource to detect at least one transition video frame in the target video resource. Alternatively, the matching device can also use other suitable scene detection algorithms or scene detection tools (e.g., pyscenedetect, a video scene cutting tool) to perform scene detection processing on the video frames in the target video resource to obtain at least one transition video frame in the target video resource.
[0184] The matching device can perform video segmentation processing on the target video resource according to the positions (i.e., video frame positions) of the at least one transition video frame in the target video resource, to obtain a plurality of sub-videos of the target video resource. For example, the position of each transition video frame in the target video resource can be used as the position for segmenting the target video resource, to implement the video segmentation processing on the target video resource. Therefore, each sub-video obtained by segmentation can correspond to a video scene (i.e., a sub-scene) in the target video resource, and each sub-video can be referred to as an independent shot segment in the target video resource.
[0185] For example, the video scene corresponding to a sub-video can be a scene in which two people chat in a study, the video scene corresponding to a sub-video can be a scene in which clothes are washed at a lake, the video scene corresponding to a sub-video can be a scene in which a father meets a daughter, and the like.
[0186] Further, the matching device can also perform text segmentation processing on the text material to obtain a plurality of sub-texts of the text material, and the text range corresponding to the target video resource can also be accurately matched from the text material in units of the sub-texts.
[0187] In an embodiment, the process of performing text segmentation processing on the text material to obtain a plurality of sub-texts of the text material can include that the text material can include a plurality of text chapters, the matching device can detect a segmentation position in each text chapter of the text material, and can perform text segmentation processing on each text chapter according to the detected segmentation position in the text chapter, to obtain a plurality of sub-texts of the text material.
[0188] Therefore, any sub-text obtained by segmentation can be any paragraph obtained by segmentation of any text chapter of the text material. That is, the text segmentation processing on the text material can be performed in units of paragraphs in the present application, to obtain each paragraph included in the text material as a plurality of sub-texts obtained by segmentation of the text material.
[0189] Step S103: performing content matching processing between the plurality of sub-videos and the plurality of sub-texts to obtain a plurality of matching pairs, each matching pair including a sub-video and a sub-text, and the sub-video and the sub-text in the same matching pair having content matching.
[0190] Specifically, the matching device can perform content matching processing between the plurality of sub-videos obtained by segmentation of the target video resource and the plurality of sub-texts obtained by segmentation of the text material, that is, matching processing between the video content of the plurality of sub-videos and the text content of the plurality of sub-texts, to obtain a plurality of matching pairs.
[0191] One matching pair can include one sub-video and one sub-text, and the sub-video and the sub-text in the same matching pair have content matching, i.e., the video content of the sub-video in the same matching pair and the text content of the sub-text in the matching pair can be matched (e.g., consistent or nearly consistent), and it can be understood that the sub-video in one matching pair can be obtained by shooting or making the sub-text in the matching pair. The sub-videos included in different matching pairs can be different, i.e., for a sub-video obtained by splitting the target video resource, a corresponding matching pair can be matched, and the matching pair includes a sub-text whose text content is matched with the video content of the sub-video.
[0192] Specifically, how to perform content matching processing between the above-mentioned multiple sub-videos and multiple sub-texts can be seen from the following Figure 7 According to the description in the corresponding embodiment.
[0193] The present application performs matching between video and text by taking a lens segment (e.g., a sub-video) and a text segment (e.g., a sub-text) as the smallest matching unit, which has finer matching granularity and higher overall matching accuracy.
[0194] In step S104, a text range in the text material that is matched with the target video resource is determined according to the multiple matching pairs.
[0195] Specifically, the matching device can determine the text range in the text material that is matched with the target video resource by the multiple matching pairs obtained by matching, and the text range in the text material that is matched with the target video resource is a text range whose text content is matched (e.g., consistent or nearly consistent) with the video content of the target video resource, and it can be understood that the target video resource is obtained by shooting or making the text in the text range in the text material that is matched with the target video resource.
[0196] In an embodiment, the process of determining, by the matching device, the text range in the text material that is matched with the target video resource according to the multiple matching pairs can include:
[0197] The matching device can count the Z text chapters in the text material in which the sub-texts in the multiple matching pairs are hit and the number of hits of the sub-texts in each hit text chapter, and any hit text chapter can include at least one sub-text in the multiple matching pairs, and Z is a positive integer, and the specific value of Z can be determined according to the actual application scenario.
[0198] That is, the Z text chapters that are hit can be each text chapter in the text material that includes (that is, hits) at least one subtext in the plurality of matching pairs. If a text chapter in the text material does not include any subtext in the plurality of matching pairs, the text chapter is not a text chapter that is hit. The Z text chapters that are hit can be Z text chapters that are continuous with each other in the text material, or can be Z text chapters that are not continuous with each other (that is, only partially continuous) in the text material.
[0199] The number of hits of a subtext in a text chapter that is hit (that is, the number of hits corresponding to a text chapter that is hit) can be the number of subtexts included in the text chapter that is hit in the plurality of matching pairs. That is, the number of hits corresponding to a text chapter that is hit is the number of subtexts included in the text chapter that is hit in the plurality of matching pairs.
[0200] The matching device can perform isolated point detection processing on each text chapter that is hit by the number of hits corresponding to each text chapter, that is, detect whether each text chapter that is hit is an isolated point in the Z text chapters that are hit. The text chapter that is hit that is detected as an isolated point can be a text chapter that is not so high in tightness (or correlation) in the Z text chapters that are hit, and the text chapter that is hit that is detected as a non-isolated point can be a text chapter that is relatively high in tightness (or correlation) in the Z text chapters that are hit.
[0201] In an embodiment, the process of performing isolated point detection processing on each text chapter that is hit by the number of hits corresponding to each text chapter can include that any text chapter that is hit in the Z text chapters that are hit can be referred to as a target text chapter. Since the process of performing isolated point detection processing on each text chapter that is hit is the same, the following will be described in detail by taking the process of performing isolated point detection processing on the target text chapter as an example, as described in the following content.
[0202] The matching device can calculate the hit density (which can be referred to as a first hit density) between the subtexts hit in the Z text chapters that are hit. The process can include that since the Z text chapters that are hit are likely to be discontinuous text chapters, the matching device can perform a filling processing on the missing text chapters between the Z text chapters that are hit to obtain L continuous text chapters corresponding to the Z text chapters that are hit. L is a positive integer and L is greater than or equal to Z. When the Z text chapters that are hit are continuous text chapters themselves, L can be equal to Z.
[0203] For example, the Z number of matched text chapters includes the 5th text chapter, the 7th text chapter to the 10th text chapter, and the 12th text chapter in the text material, i.e. Z equals 6, and thus, the missing text chapters between the Z number of matched text chapters can be filled in as the 6th text chapter between the 5th text chapter and the 7th text chapter, and the 11th text chapter between the 10th text chapter and the 12th text chapter, so that the L number of continuous text chapters can include the 5th text chapter to the 12th text chapter in the text material, i.e. L equals 8.
[0204] The matching device can calculate the ratio between the total number of the matched subtexts in the Z number of matched text chapters and L as the first hit density, i.e. the first hit density can equal the value obtained by dividing the total number of the matched subtexts in the Z number of matched text chapters by L. The total number of the matched subtexts in the Z number of matched text chapters can equal the total number of the subtexts included in the plurality of matching pairs (the same subtexts can be de-duplicated).
[0205] The matching device can obtain the density reference threshold for the subtext matching (i.e. the subtext matching on the target video resource) through the first hit density calculated above, which can include that the matching device can calculate the ratio between the total number of the matched subtexts in the Z number of matched text chapters and Z as the maximum hit density, i.e. the maximum hit density can equal the value obtained by dividing the total number of the matched subtexts in the Z number of matched text chapters by Z.
[0206] The matching device can obtain the difference between the maximum hit density and the first hit density above, such as calculating the value obtained by subtracting the first hit density from the maximum hit density, and can calculate the product between the difference and the set density control parameter as the density fluctuation value. The density control parameter can be a preset relevant number for density fluctuation control, such as the density control parameter can be set to 60% (or other values), and the greater the density control parameter is set, the closer the subsequent matched subtexts can be, and the smaller the density control parameter is set, the less close the subsequent matched subtexts can be.
[0207] The matching device can add the calculated density fluctuation value and the first hit density as the density reference threshold, i.e. the sum of the density fluctuation value and the first hit density can be the density reference threshold.
[0208] For example, Z can equal 9, L can equal 12, the number of subtexts hit in the Z hit text chapters can equal 67, the density control parameter is set to 60% (indicating that more than 60% information density (i.e. the density between the finally matched subtexts) needs to be guaranteed), then the first hit density can equal 67 / 12=5.58, the maximum hit density can equal 67 / 9=7.44, the density fluctuation value can equal (7.44-5.58) x 60%=1.11, thus the density reference threshold can equal 1.11+5.58=6.69.
[0209] The matching device can also calculate the hit density (which can be referred to as the second hit density) of the subtexts hit in the Z-1 hit text chapters other than the target text chapter in the Z hit text chapters. The second hit density can be calculated by removing the target text chapter from the L text chapters that are padded to the target text chapter, and also removing the text chapters padded before and after the target text chapter, so that S text chapters are obtained, S is a positive integer, S equals L minus 1 minus the number of text chapters padded before and after the target text chapter. Thus, the matching device can divide the total number of subtexts hit in the Z-1 hit text chapters by S to obtain the second hit density.
[0210] For example, the Z hit chapters can include the 5th text chapter, the 7th-9th text chapters, and the 11th text chapter in the text material, thus the L continuous text chapters can include the 5th text chapter-11th text chapter in the text material, and if the target text chapter is the 5th text chapter, then the S text chapters can include the 7th text chapter-11th text chapter in the text material.
[0211] Further, the matching device can also count the number of subtexts hit in the Z-1 hit text chapters (which can be referred to as the first number) and the number of subtexts hit in the Z hit text chapters (which can be referred to as the second number), and the matching device can obtain the number ratio between the first number and the second number, i.e. the number ratio can equal the value obtained by dividing the first number by the second number.
[0212] The matching device can determine whether the target text chapter is a divergence point in the Z matched text chapters according to the quantity ratio, the second hit density and the density reference threshold. The process can include: if the quantity ratio is greater than or equal to a set ratio threshold (for example, the ratio threshold can be set to 95%, which can be understood as accepting 5% of divergence points), and the second hit density is greater than or equal to the calculated density reference threshold, it can be determined that the target text chapter is detected as a divergence point, indicating that the target text chapter is not so close between the Z matched text chapters, and the target text chapter is a matched text chapter with a lower association with the target video resource, which needs to be removed.
[0213] If the quantity ratio is less than the set ratio threshold, or the second hit density is less than the calculated density reference threshold, it can be determined that the target text chapter is detected as a non-divergence point, indicating that the target text chapter is relatively close between the Z matched text chapters, and the target text chapter is a matched text chapter with a higher association with the target video resource, which needs to be retained.
[0214] The matching device can perform divergence point detection on each matched text chapter according to the same principle of the above divergence point detection of the target text chapter, to obtain the divergence point detection result (which can be detected as a divergence point or a non-divergence point) of each matched text chapter. The matching device can remove the matched text chapters detected as divergence points in the above Z matched text chapters, and obtain at least one retained text chapter, which is a text chapter detected as a non-divergence point in the above Z matched text chapters.
[0215] The matching device can take the chapter range of the at least one retained text chapter in the text material as the text range matched with the target video resource. Alternatively, the at least one retained text chapter can be discontinuous text chapters. The matching device can also perform gap filling between the at least one retained text chapter to obtain each continuous text chapter corresponding to the at least one retained text chapter, so as to take the chapter range of each continuous text chapter corresponding to the at least one retained text chapter in the text material as the text range matched with the target video resource.
[0216] Please refer to Figure 4 , Figure 4 is a schematic diagram of a principle provided by an embodiment of the present application for obtaining a text range matched with a target video resource. As Figure 4As shown, the target video resource hits the text chapters corresponding to the black dots in the full text of the text material (i.e., the text chapters corresponding to the black dots can be the matched text chapters), and the text chapter corresponding to the first black dot is detected as a divergence point, while the text chapters corresponding to the other black dots are not detected as divergence points and thus can be used as reserved text chapters. Thus, the chapter range of each reserved text chapter can be used as a text range matched with the target video resource, or the chapter range of each continuous text chapter corresponding to each reserved text chapter can be used as a text range matched with the target video resource.
[0217] The above video resource set (e.g., a TV series) can be played in a video client, and the matching device can be a background device (e.g., a background server) of the video client. Thus, when the target video resource is played in the video client (e.g., when the target video resource is started or when the target video resource is about to be played to the end), the matching device can push a reading link (i.e., a URL) corresponding to the text range matched with the target video resource to the video client.
[0218] The video client can be configured to display (e.g., in a playing interface of the target video resource) the received reading link (which can be displayed through a corresponding reading control) when the target video resource is played (e.g., when the target video resource is started or when the target video resource is about to be played to the end), and can be configured to jump to and display the text in the text range matched with the target video resource in the text material in response to a triggering operation (e.g., a clicking operation) on the reading link (e.g., the reading control).
[0219] The display of the text in the text range matched with the target video resource in the text material can be display of only the text in the text range matched with the target video resource in the text material, or can be jump display to the text in the text range matched with the target video resource in the text material (e.g., jump display to the first text chapter in the text range matched with the target video resource) in an opening interface (e.g., a reading interface) of the text material, so as to allow the user to view and read the text in the text range matched with the target video resource in the text material, and facilitate the user to freely switch between the two media (i.e., the video resource and the text material) and explore the story content.
[0220] The matching device can also establish a mapping relationship between each sub-video in the target video resource and the sub-text matched therewith, and can record the start time and the end time of each sub-video in the target video resource. Thus, in the future, the user can be recommended to read the sub-text matched with each sub-video.
[0221] The matching device can determine the text range in the text material that matches each video resource in the video resource set according to the same principle of determining the text range in the text material that matches the target video resource.
[0222] In addition, the application can also record (e.g., in a database record) the text range (i.e., the matched text range, such as the corresponding novel chapter) corresponding to each video resource, the supplementary text range (e.g., the missing text between the text ranges corresponding to each sub-video of a video resource in the text material, or / and the missing text between the text range corresponding to the video resource and the text ranges corresponding to the previous and next video resources), and the text range corresponding to the next video resource, and can subsequently recommend these text contents to the user for viewing and reading.
[0223] Please refer to Figure 5a and Figure 5b , Figure 5a is an interface schematic diagram of a video playing interface provided by an embodiment of the application, Figure 5b is an interface schematic diagram of a text reading interface provided by an embodiment of the application. As shown in Figure 5a , the video playing interface a1 can include a video playing area, in which the 11th episode of the TV series “TV series 1” can be playing. The TV series “TV series 1” can have a total of 15 episodes, and the 15 episodes of videos can constitute the above-mentioned video resource set of the application. When the 11th episode playing in the video playing interface a1 is about to end, such as when the main content is played to the end, the end credits part is played, the video playing interface a2 can be displayed.
[0224] The video playing interface a2 can display a “original work” control a3, which is the above-mentioned reading control. When the user clicks the reading control, the video client can display Figure 5b the text reading interface, in which the text in the text range in the original novel (i.e., the text material) of the TV series “TV series 1” that matches the 11th episode of the TV series “TV series 1” can be displayed. Thus, the text at the beginning of the text range can be “The sun sets, the golden afterglow shines on the rippling lake, dyeing the entire water surface into a piece of orange”.
[0225] In addition, the above-mentioned video resource set can have an order between each video resource (e.g., the order can be the order between the video contents of each video resource, such as the coherent order or the connected order between the video contents), such as the playing order of each episode video in the TV series. Therefore, the application can also obtain the supplementary text between each video resource through the above-mentioned determined text range that matches each video resource in the video resource set, as described below.
[0226] For any two adjacent video resources (such as the target video resource and the next video resource of the target video resource), if the next paragraph of the last paragraph in the text range matching the target video resource in the text material is not the first (i.e., the first) paragraph in the text range matching the next video resource, a supplementary text can be generated between the target video resource and the next video resource thereof, which can be the next paragraph (or multiple paragraphs thereafter, or all paragraphs between the two text ranges corresponding to the two adjacent video resources, etc.) of the last paragraph in the text range matching the target video resource in the text material.
[0227] Therefore, when playing the set of video resources in the video client, a supplementary control can be displayed between the target video resource and the next video resource thereof, which can be a control for reading the link of the supplementary text between the target video resource and the next video resource thereof. The video client can be used to open the text material (such as displaying the reading interface of the text material) through the reading link in response to the triggering operation (such as the clicking operation) of the supplementary control, and automatically jump to display the position of the supplementary text between the target video resource and the next video resource thereof in the text material, so that the user can read the part of the text content in the text material that is not videoed between the target video resource and the next video resource thereof, to supplement the user's knowledge of the content of the text material before playing the next video resource, to improve the overall continuity of the user's subsequent viewing. Optionally, when automatically jumping to display the position of the supplementary text between the target video resource and the next video resource thereof in the text material, the supplementary text can also be highlighted (such as highlighted, etc.). Alternatively, the supplementary text can also be converted into audio (such as audio reading the supplementary text) or video (such as animation video made for the supplementary text, etc.) for output (such as broadcast or display).
[0228] That is, by using the method provided in the present application, it can also be screened which part of the text material (such as a novel) is not videoed or made into a video. This part of the text that is not videoed or made into a video (such as the text in the text material other than the text range matching each video resource in the set of video resources) can also be used to provide a reference for subsequent creation (such as secondary creation) of the text material, such as script modification, etc.
[0229] Please refer to Figure 6 , Figure 6 is another interface schematic diagram of a video playing interface provided by an embodiment of the present application. As shown in Figure 6The video playing area of the video playing interface b1 shown can be playing the 7th episode of the TV series "TV series 2", the TV series "TV series 2" can have a total of 16 episodes, and the 16 episodes of video can constitute the above-mentioned video resource set of the application. Here, there is supplementary text between the 2nd and 3rd episodes, so the control b2 (which can be the above-mentioned supplementary control) can be displayed between the 2nd and 3rd episodes. The user can click the control b2, and the video client can display a text reading interface, which can display the supplementary text between the 2nd and 3rd episodes.
[0230] Similarly, there is also supplementary text between the 10th and 11th episodes here, so the control b3 (which can also be the above-mentioned supplementary control) can be displayed between the 10th and 11th episodes. The user can click the control b3, and the video client can display a text reading interface, which can display the supplementary text between the 10th and 11th episodes.
[0231] The application can obtain text material associated with a video resource set and a target video resource, the text content described by the text material matches the video content presented by the video resource set, and the target video resource is any video resource in the video resource set; and the target video resource can be subjected to video segmentation processing to obtain a plurality of sub-videos of the target video resource, and the text material can be subjected to text segmentation processing to obtain a plurality of sub-texts of the text material; thereby, content matching processing can be performed between the plurality of sub-videos and the plurality of sub-texts to obtain a plurality of matching pairs, a matching pair includes a sub-video and a sub-text, and the sub-video and the sub-text in the same matching pair have content matching; thereby, the text range in the text material that matches the target video resource can be determined according to the plurality of matching pairs. As can be seen, the method proposed by the application can segment the target video resource into a plurality of sub-videos of smaller units, and can also segment the text material associated with the video resource set into a plurality of sub-texts of smaller units, thereby, through the refined content matching between the plurality of sub-videos and the plurality of sub-texts of smaller units, the text range in the text material that matches the target video resource can be accurately determined.
[0232] Please refer to Figure 7 , Figure 7 is a flowchart of content matching processing between sub-videos and sub-texts provided by an embodiment of the application. As Figure 7 shown, the flowchart can include:
[0233] Step S201, respectively encode a plurality of sub-videos to generate video encoding features of each sub-video, and the video encoding features of any sub-video are used to represent the video content presented by any sub-video.
[0234] Specifically, the matching device can perform feature coding processing on each of the plurality of sub-videos segmented from the target video resource to generate a video coding feature of each sub-video. One sub-video can have one corresponding video coding feature. The video coding feature of any sub-video can be used to represent the video content presented by the any sub-video, i.e., the video coding feature of any sub-video is used to embody the features of the video content of the any sub-video.
[0235] Any one of the plurality of sub-videos can be referred to as a target sub-video. Since the feature coding processing on each sub-video to generate the video coding feature of each sub-video is the same, the following will be described in detail by taking the process of performing feature coding processing on the target sub-video to generate the video coding feature of the target sub-video as an example, as described in the following content.
[0236] The matching device can perform extraction processing on the video frames in the target sub-video to obtain a plurality of extracted video frames in the target sub-video. For example, the plurality of extracted video frames can be obtained by extracting the target sub-video at a specified time interval or number of extractions. For example, the plurality of extracted video frames can be obtained by extracting the target sub-video at a rate of 3 frames per second. By extracting the video frames of the target sub-video and subsequently processing the plurality of extracted video frames, the computational load of feature coding processing on the target sub-video can be reduced, thereby saving computational overhead and improving computational speed.
[0237] The matching device can perform text extraction processing on the plurality of extracted video frames from a plurality of first text extraction dimensions to generate extracted text of the target sub-video under each first text extraction dimension. The target sub-video can have one extracted text under one first text extraction dimension.
[0238] In an embodiment, the matching device performs text extraction processing on the plurality of extracted video frames from a plurality of first text extraction dimensions to generate extracted text of the target sub-video under each first text extraction dimension. The manner in which the matching device performs text extraction processing on the plurality of extracted video frames from a plurality of first text extraction dimensions to generate extracted text of the target sub-video under each first text extraction dimension can include that the plurality of first text extraction dimensions can include a picture extraction dimension, a character detection dimension, and a content understanding dimension.
[0239] Therefore, the matching device can extract the text appearing in the pictures of the plurality of extracted video frames in the dimension from the pictures to obtain picture extraction text. For example, the matching device can perform OCR (Optical Character Recognition) detection on the plurality of extracted video frames to detect the text (i.e., the appearing words or / and characters) appearing in the plurality of extracted video frames, and can clean the appearing text to obtain picture extraction text. For example, the useless text can include the appearing title of the TV series (e.g., the name of the TV series), the episode number (e.g., the episode number of the target video resource in the TV series, such as the 5th episode), the platform name (e.g., the name of the platform for playing the video resource set in the picture, which can be the name of the video client where the video resource set is located), and / or stop words, etc. The stop words can be words in the stop word library collected and constructed. In actual application scenarios, the specific text included in the useless text can be flexibly set by itself, and the present application does not limit it. The picture extraction text can be the extraction text of the target sub-video in the picture extraction dimension.
[0240] The matching device can also perform face detection processing on the plurality of extracted video frames in the above-mentioned role detection dimension to obtain the role identification (which can be referred to as the first role identification) of the video role to which the face in the plurality of extracted video frames belongs. The first role identification can be the name of the video role (e.g., the name of the character) to which the face in the plurality of extracted video frames belongs. For example, the matching device can obtain a trained face detection model, which can be obtained by training the face images of the video roles (e.g., actors) in the video resource set and their corresponding role names. Therefore, the face detection model has the ability to detect and identify the role name of the video role to which the face in the video frame belongs. Therefore, the matching device can call the face detection model to perform face detection processing on the plurality of extracted video frames to obtain the first role identification of the video role to which the detected face belongs. The first role identification can include the role identification (e.g., the role name) of each detected face, such as "Xiaodu", "Xiatian", etc. The first role identification can be the extraction text of the target sub-video in the role detection dimension.
[0241] The matching device can also perform content understanding processing on the picture content of the plurality of extracted video frames in the above-mentioned content understanding dimension to generate a content summary text of the video content presented by the plurality of extracted video frames. For example, the plurality of extracted video frames can be used to construct prompt information (i.e., prompt words) for a trained large language model (which can be a multi-modal large language model, MLLM), so as to call and guide the large language model to generate and output a content summary text for the plurality of extracted video frames through the prompt information. For example, the large language model can be internVL3 (a multi-modal large language model).
[0242] In an implementation, the picture prompt text and the plurality of extracted video frames can be used to construct a prompt information for a large language model, such as the prompt information can be "combine the input text and the plurality of images, use a paragraph to describe all the events happened in the video represented by the images, do not describe the content of each image separately, do not describe the clothing of the characters, and output the corresponding video content", the input text can be the picture prompt text, and the input plurality of images can be the plurality of extracted video frames. Therefore, the matching device can guide and call the above trained large language model through the prompt information, generate and output the content summary text of the plurality of extracted video frames, the content summary text can be obtained by content understanding and summarizing the video represented by the plurality of extracted video frames, and the content summary text can be the extracted text of the target sub-video in the content understanding dimension. Optionally, the above picture prompt text in the plurality of extracted video frames can also be obtained by the trained large language model.
[0243] Therefore, the extracted text of the target sub-video in the above plurality of first text extraction dimensions can include the above picture extraction text, the first character identification, and the content summary text.
[0244] The matching device can perform feature encoding processing on the prompt text of the target sub-video in the above plurality of first text extraction dimensions to generate video encoding features of the target sub-video, and the process can include:
[0245] The matching device can obtain a feature encoding model, and the feature encoding model can include a feature encoding branch corresponding to each of the above first text extraction dimensions, that is, one first text extraction dimension can correspond to one feature encoding branch in the feature encoding model, and the number of feature encoding branches can be equal to the number of first text extraction dimensions.
[0246] The matching device can call each feature encoding branch to perform feature encoding processing on the extracted text of the target sub-video in the corresponding first text extraction dimension to generate text encoding features (which can be referred to as first text encoding features) of the extracted text of the target sub-video in each first text extraction dimension. Wherein, calling a feature encoding branch to perform feature encoding processing on the extracted text of the target sub-video in the first text extraction dimension corresponding to the feature encoding branch can generate a first text encoding feature of the target sub-video in the first text extraction dimension, that is, one first text extraction dimension can correspond to one first text encoding feature.
[0247] The matching device can perform feature fusion processing on the first text encoding features of the extracted text of the target sub-video in the plurality of first text extraction dimensions to generate the video encoding feature of the target sub-video. For example, an extraction weight corresponding to each first text extraction dimension can be set respectively, and the extraction weight corresponding to a more important first text extraction dimension can be set higher. The extraction weight corresponding to each first text extraction dimension can be used for feature weighting processing (which belongs to feature fusion processing) on the first text encoding features of the extracted text of the target sub-video in the plurality of first text extraction dimensions to obtain the video encoding feature of the target sub-video.
[0248] Optionally, the extraction weight corresponding to each first text extraction dimension can be used for multiplication on the first text encoding features of the extracted text of the target sub-video in each first text extraction dimension to obtain the weighted text encoding feature of the extracted text of the target sub-video in each first text extraction dimension. Thus, the weighted text encoding features of the extracted text of the target sub-video in each first text extraction dimension can be spliced or added (for example, the dimensions of the weighted text encoding features can be the same, and the addition here can mean adding the elements at the same positions in the weighted text encoding features) to obtain the video encoding feature of the target sub-video.
[0249] The matching model can generate and obtain the video encoding feature of each sub-video according to the same principle of obtaining the video encoding feature of the target sub-video. The video encoding feature of each sub-video can be an embedding.
[0250] In step S202, the feature encoding processing is performed on the plurality of sub-texts respectively to generate the text encoding feature of each sub-text. The text encoding feature of any sub-text is used to represent the text content described by the any sub-text.
[0251] Specifically, the matching device can perform feature encoding processing on the plurality of sub-texts obtained by cutting the text material to generate the text encoding feature of each sub-text. One sub-text can have one corresponding text encoding feature. The text encoding feature of any sub-text can be used to represent the text content described by the any sub-text, i.e., the video encoding feature of any sub-video is used to represent the features of the text content of the any sub-text.
[0252] Any of the plurality of sub-texts can be referred to as a target sub-text. Since the feature encoding processing is performed on each sub-text to generate the text encoding feature of each sub-text, the principle is the same. Therefore, the following describes the process of performing feature encoding processing on the target sub-text to generate the text encoding feature of the target sub-text as an example for specific description.
[0253] The matching device can perform text extraction processing on the target subtext from the plurality of second text extraction dimensions to generate extracted text of the target subtext under each second text extraction dimension. The target subtext can have one extracted text under one second text extraction dimension.
[0254] For example, the plurality of second text extraction dimensions can include a content extraction dimension, a dialogue extraction dimension, and a character extraction dimension. Therefore, the process of the matching device performing text extraction processing on the target subtext from the plurality of second text extraction dimensions to generate extracted text of the target subtext under each second text extraction dimension can include:
[0255] The matching device can perform content extraction processing on the target subtext from the content extraction dimension to obtain content extracted text of the target subtext. The process can include that the matching device can obtain the text length (e.g., the number of contained characters) of the target subtext. If the text length is less than or equal to a set length threshold (the specific value can be set according to the actual application scenario), it indicates that the target subtext is not too long and is within an acceptable range. In this case, the target subtext can be directly taken as the content extracted text.
[0256] If the text length is greater than the set length threshold, it indicates that the target subtext is too long and is not within an acceptable range, and it is very difficult to directly match the target subtext. In this case, the target subtext can be subjected to content compression processing to generate the content extracted text.
[0257] In an implementation, the target subtext can also be subjected to content compression processing by a trained large language model to generate the content extracted text. For example, a prompt information for the large language model can be generated from the target subtext to prompt and guide the large language model to perform content compression on the target subtext to generate and output the content extracted text. For example, the prompt information can be “Given a novel content, please extract the summary of the novel content. When the novel content contains dialogue or character speeches, the speech content needs to be preserved, and the text description content between speeches needs to be compressed. Requirements: Extract the summary within 100 words, and the input novel content is: target subtext”. The process of generating a summary from the target subtext is the process of performing content compression on the target subtext, and the length of the summary required in the prompt information can be flexibly set according to the actual application scenario. Therefore, the matching device can give the prompt information to the large language model, so that the large language model can generate and output the content extracted text of the input target subtext under the prompt and guidance of the prompt information. The content extracted text can be a summary generated from the target subtext. The content extracted text is the extracted text of the target subtext under the content extraction dimension.
[0258] For example, if the target subtext is: Since Aunt hates his relationship, Li family except uncle, others do not treat Li four, at least cousins will not show too close to him. Li three impatiently said: "I have been removed from office, but have the teacher's protection, do not need to be sent, so take care of yourself". Li three is famous in the capital school, and is highly valued, therefore, after uncle's incident, he was not jailed, but was not allowed to leave the capital, more than a day to keep running around. Li four was silent, he did not think that Li three would be better than himself, and he was afraid that he would not be able to turn over in the future. Two days later, the women of Li family will be sent to the teaching house. Li three is a scholar, how can he live in the capital, perhaps being sent to the frontier is a better choice.
[0259] Therefore, the content extraction text output by the content compression processing of the target subtext can be: Li three is removed from office, but is not sent because of the protection of the school, and responds to Li four with cold voice: "take care of yourself. Went to the border, can live for a year is a year". Li four speculates that he will not be able to turn over in the future, and the women of Li family will be sent to the teaching house two days later, and thinks that Li three as a scholar is difficult to bear.
[0260] More, the matching device can also extract the dialogue from the above dialogue extraction dimension to the target subtext to obtain the dialogue text in the target subtext. That is, the matching device can extract the dialogue in the target subtext as the dialogue text of the target subtext, and the dialogue text is the extraction text of the target subtext in the dialogue extraction dimension. For example, the dialogue text can include the dialogue and the aside in the target subtext. The dialogue text can include the text in the double quotation marks in the target subtext.
[0261] The matching device can also extract the character from the above character extraction dimension to the target subtext to obtain the character identification (such as character name, which can be called second character identification) of the film and television character described in the target subtext. For example, the matching device can obtain the character identification library composed of all the film and television characters in the video resource set, and can match the character identification in the target subtext through the character identification library, such as the character identification in the target subtext which is the same as the character identification in the character identification library can be taken as the second character identification. For example, the second character identification can include the character name mentioned and described in the target subtext. The second character identification is the extraction text of the target subtext in the character extraction dimension.
[0262] Therefore, the extracted text of the target subtext under the plurality of second text extraction dimensions can include the content extraction text, the dialogue text, and the second character identifier. When the second character identifier and / or the dialogue text do not exist in the target subtext, a default character (e.g., "none") can be used to represent the second character identifier and / or the dialogue text.
[0263] The matching device can perform feature encoding processing on the prompt text of the target subtext under the plurality of second text extraction dimensions to generate text encoding features of the target subtext. The process can include:
[0264] Similarly, the matching device can obtain a feature encoding model, which can be the same as the model used to generate the video encoding features of the target sub-video. The feature encoding model can include a feature encoding branch corresponding to each second text extraction dimension, i.e., one second text extraction dimension can correspond to one feature encoding branch in the feature encoding model. The number of feature encoding branches can be equal to the number of second text extraction dimensions, and the number of second text extraction dimensions can be equal to the number of first text extraction dimensions.
[0265] The first text extraction dimension and the second text extraction dimension can share a feature encoding branch in the feature encoding model. For example, the frame extraction dimension and the dialogue extraction dimension can correspond to the same feature encoding branch in the feature encoding model, the content understanding dimension and the content extraction dimension can correspond to the same feature encoding branch in the feature encoding model, and the character detection dimension and the character extraction dimension can correspond to the same feature encoding branch in the feature encoding model. That is, according to the dimension similarity between the first text extraction dimension and the second text extraction dimension, one first text extraction dimension and one second text extraction dimension with certain similarity can correspond to the same feature encoding branch in the feature encoding model.
[0266] The matching device can call each feature encoding branch to perform feature encoding processing on the extracted text of the target subtext under the corresponding second text extraction dimension to generate text encoding features (which can be referred to as second text encoding features) of the extracted text of the target subtext under each second text extraction dimension. Calling one feature encoding branch to perform feature encoding processing on the extracted text of the target subtext under the second text extraction dimension corresponding to the feature encoding branch can generate one second text encoding feature of the target subtext under the second text extraction dimension, i.e., one second text prompt dimension can correspond to one second text encoding feature.
[0267] The matching device can perform feature fusion processing on the second text coding features of the extracted text of the target subtext in the plurality of second text extraction dimensions to generate the text coding features of the target subtext. For example, an extraction weight corresponding to each second text extraction dimension can be set respectively, and the extraction weight corresponding to a more important second text extraction dimension can be set to be higher. The text coding features of the extracted text of the target subtext in the plurality of second text extraction dimensions can be weighted (which belongs to feature fusion processing) by using the extraction weight corresponding to each second text extraction dimension to obtain the text coding features of the target subtext.
[0268] Optionally, the second text coding features of the extracted text of the target subtext in each second text extraction dimension can be multiplied by using the extraction weight corresponding to each second text extraction dimension to obtain the weighted text coding features of the extracted text of the target subtext in each second text extraction dimension. Thus, the weighted text coding features of the extracted text of the target subtext in each second text extraction dimension can be spliced or added (for example, the dimensions of the weighted text coding features can be the same, and the addition here can mean adding the elements at the same positions in the weighted text coding features) to obtain the text coding features of the target subtext.
[0269] Alternatively, in another embodiment, the text coding features generated for the target subtext in the above manner can be used as the initial coding features of the target subtext, and the matching device can generate and obtain the initial coding features of each subtext according to the above principle. Since adjacent different subtexts (for example, adjacent different paragraphs) in the text material often have context continuity (also context dependency), the initial coding features of the target subtext and the initial coding features of the adjacent subtexts of the target subtext can be used together to generate the final text coding features of the target subtext, which can improve the accuracy of the text coding features obtained for the target subtext. For example, the adjacent subtexts of the target subtext can include at least one subtext before the target subtext (for example, two paragraphs adjacent to the target subtext before the target subtext) and at least one subtext after the target subtext (for example, two paragraphs adjacent to the target subtext after the target subtext). The order of the subtexts before and after can be the order of the subtexts in the text material (for example, the order of the paragraphs).
[0270] The subtext before the target subtext among the adjacent subtexts of the target subtext can be referred to as a first adjacent subtext, and the subtext after the target subtext among the adjacent subtexts of the target subtext can be referred to as a second adjacent subtext. The matching device can obtain a text weight (such as 0.6) set for the current target subtext, a text weight (such as 0.2) set for the first adjacent subtext of the target subtext, and a text weight (such as 0.2) set for the second adjacent subtext of the target subtext.
[0271] Therefore, the matching device can perform feature weighting summation on the initial encoding features of the target subtext, the initial encoding features of the first adjacent subtext, and the initial encoding features of the second adjacent subtext by using the text weight corresponding to the target subtext, the text weight corresponding to the first adjacent subtext, and the text weight corresponding to the second adjacent subtext, to obtain the final text encoding features of the target subtext.
[0272] For example, the text weight of the current target subtext can be 0.6, the text weight of the first adjacent subtext of the target subtext can be 0.2, the text weight of the second adjacent subtext of the target subtext can also be 0.2, the initial encoding features of the target subtext can be denoted as E1, the initial encoding features of the first adjacent subtext can be denoted as E2, and the initial encoding features of the second adjacent subtext can be denoted as E3. Therefore, the final text encoding features of the target subtext can be 0.6×E1+0.2×E2+0.2×E3. Each initial encoding feature can be a feature vector, and the dimensions can be the same. The addition "+" in the formula can mean adding elements at the same position in the feature vectors.
[0273] The matching device can generate and obtain the text encoding features of each subtext according to the same principle of obtaining the text encoding features of the target subtext. The text encoding features of each subtext can be a feature vector (embedding).
[0274] The present application fully acquires multiple information sources of the sub-video by extracting dimensions from the multiple first texts, so that by fusing the multiple information sources, the video coding features of the sub-video can be effectively generated; and the multiple information sources of the sub-text are fully acquired by extracting dimensions from the multiple second texts, so that by fusing the multiple information sources, the text coding features of the sub-text can be effectively generated, and by fully acquiring the multiple information sources of the sub-text from the multiple second texts, the interference brought by the complex situations such as character reference in the text material to the matching of the sub-text can be overcome, and the content of the sub-text can be accurately understood. Therefore, by the accurate video coding features of the sub-video and the accurate text coding features of the sub-text, the overall quality of the content matching processing between the sub-video and the sub-text can be greatly improved.
[0275] In addition, in the present application, by extracting dimensions from the multiple first texts to perform text extraction on the target sub-video to obtain multiple extracted texts associated with the target sub-video, and by extracting dimensions from the multiple second texts to perform text extraction on the sub-text to obtain multiple extracted texts associated with the sub-text, the cross-domain boundary (i.e. cross-modal interface) between the sub-video and the sub-text can be broken, the information domain (modal) between the sub-video and the sub-text can be unified, so that the same feature coding model can be used for feature coding processing of the sub-video and the sub-text, that is, the same feature coding model is used to generate the video coding features of the sub-video and the text coding features of the sub-text, so that the video coding features and the text coding features are located in the same feature space (such as the feature space of the feature coding model), that is, the feature space where the video coding features are located and the feature space where the text coding features are located are consistent, and the alignment relationship between the sub-video and the sub-text is improved, so that subsequent matching between the video coding features and the text coding features can achieve highly accurate matching between the sub-video and the sub-text.
[0276] In step S203, based on the video coding features of each sub-video and the text coding features of each sub-text, content matching processing is performed between the multiple sub-videos and the multiple sub-texts to obtain multiple matching pairs.
[0277] Specifically, the matching device can perform content matching processing between the multiple sub-videos and the multiple sub-texts based on the video coding features generated for each sub-video and the text coding features generated for each sub-text to obtain the multiple matching pairs, as described below.
[0278] Any one of the multiple sub-videos can be referred to as a target sub-video, and since the principle of obtaining the matching pair where each sub-video is located is the same, the process of obtaining the matching pair where the target sub-video is located is also described below as an example, as described below.
[0279] The matching device can obtain a feature similarity (e.g., a cosine similarity) between the text encoding feature of each of the plurality of subtexts and the video encoding feature of the target subvideo. The text encoding feature of one subtext and the video encoding feature of the target subvideo can have a feature similarity.
[0280] The matching device can sort the plurality of subtexts in descending order of the feature similarity between the text encoding feature of each subtext and the video encoding feature of the target subvideo to obtain sorted subtexts. In the sorted subtexts, the subtext with a greater feature similarity between the text encoding feature and the video encoding feature of the target subvideo can be arranged in the front.
[0281] The matching device can arrange the K subtexts in the front of the sorted subtexts as the first candidate subtexts, where K is a positive integer, and K can be less than the total number of the plurality of subtexts, e.g., K can be equal to 20. The process of obtaining the first candidate subtexts can be a process of rough screening of subtexts.
[0282] The matching device can determine the subtext matching the target subvideo from the K first candidate subtexts. This process can be a process of fine screening of subtexts, which can include: the matching device can obtain a first role identifier of a film and television role appearing in the target subvideo. The process of obtaining the first role identifier can refer to the related description in step S201 above. The matching device can arrange the first candidate subtext including the first role identifier in the K first candidate subtexts as the second candidate subtext. There can be at least one second candidate subtext. For example, the first role identifier can include the name of one or more film and television roles. When a first candidate subtext includes (i.e., describes) the name of the one or more film and television roles, the first candidate subtext can be arranged as the second candidate subtext.
[0283] The matching device can determine the subtext matching the target subvideo from the at least one second candidate subtext. This process can include: the matching device can calculate the comprehensive matching degree between each second candidate subtext and the target subvideo, and can arrange the second candidate subtext with the greatest comprehensive matching degree between the at least one second candidate subtext and the target subvideo as the subtext matching the target subvideo.
[0284] In an embodiment, the way of calculating the comprehensive matching degree between each second candidate subtext and the target subvideo can include: the matching device can obtain a feature similarity (e.g., a cosine similarity) between the text encoding feature of each second candidate subtext and the video encoding feature of the target subvideo (which can be obtained when the plurality of subtexts are sorted to obtain the first candidate subtexts).
[0285] The matching device can also calculate a co-occurrence indication parameter between the text keywords associated with each second candidate subtext and the text keywords associated with the target sub-video, respectively. The text keywords associated with a second candidate subtext can include keywords extracted from the second candidate subtext itself and keywords extracted from the dialogue text of the second candidate subtext (which can be the extracted text of the second candidate subtext under the dialogue extraction dimension described above). The text keywords associated with the target sub-video can be keywords extracted from the multiple extracted texts of the target sub-video under the multiple first text extraction dimensions described above.
[0286] The matching device can select a suitable keyword extraction method to extract keywords from the second candidate subtext and the dialogue text of the second candidate subtext to obtain the text keywords associated with the second candidate subtext, and use the keyword extraction method to extract keywords from the multiple extracted texts of the target sub-video under the multiple first text extraction dimensions described above to obtain the text keywords associated with the target sub-video. For example, the keyword extraction method can be a TF-IDF extraction method, or it can be extracted through a question and answer mode of a large language model.
[0287] Here, taking the keyword extraction using the TF-IDF extraction method described above as an example, and taking the process of keyword extraction from any second candidate subtext as an example, the process can include:
[0288] The matching device can collect a document set, which can be composed of each text chapter in each novel in the novel library. The document set can include multiple documents, and a document can be a text chapter of a novel.
[0289] The matching device can perform text preprocessing on the any second candidate subtext to obtain multiple valid words in the any second candidate subtext. For example, the matching device can perform word segmentation processing on the any second candidate subtext to obtain multiple words of the any second candidate subtext, and can remove invalid words (such as stop words and related meaningless adverbs, such as “of”, “of”, “and”, “and” and the like) in the multiple words, thereby obtaining multiple valid words in the any second candidate subtext.
[0290] The matching device can calculate the frequency (which can be denoted as TF) of each valid word in the any second candidate subtext, respectively. The TF of any valid word can be equal to the number of occurrences of the valid word in the any second candidate subtext divided by the total number of words in the any second candidate subtext (such as the number of the multiple words described above).
[0291] The matching device can also calculate the inverse document frequency (IDF) of each valid word in the document set (denoted as IDF), and the IDF of any valid word can be equal to: log (total number of documents in the document set / number of documents containing the valid word + 1), where the "+1" in the formula can prevent the denominator from being 0.
[0292] Thus, the TF-IDF value (which can be understood as an importance value, and a higher value indicates that the corresponding valid word is more important) of each valid word can be calculated, and the TF-IDF value of a valid word can be equal to the TF of the valid word multiplied by the IDF of the valid word. The matching device can sort the plurality of valid words according to the TF-IDF values of the plurality of valid words in descending order, and the first T valid words in the sorted valid words can be used as the keywords extracted from the any second candidate subtext. T is a positive integer, and the value of T can be flexibly set according to the actual application scenario, such as the value of T can be set as a target proportion (such as 20%) of the text length of the any second candidate subtext.
[0293] The matching device can extract keywords from the dialogue text of the any second candidate subtext through the document set according to the same principle of keyword extraction from the any second candidate subtext, so as to obtain all text keywords associated with the any second candidate subtext. In addition, the matching device can extract keywords from the plurality of extracted texts of the target sub-video through the document set according to the same principle of keyword extraction from the any second candidate subtext, so as to obtain each text keyword associated with the target sub-video.
[0294] For example, the calculation of the co-occurrence indication parameter between the text keywords associated with any second candidate subtext and the text keywords associated with the target sub-video can include: the matching device can obtain the number of text keywords that co-occur in the text keywords associated with any second candidate subtext and the text keywords associated with the target sub-video (which can be referred to as the co-occurrence number), i.e., the co-occurrence number can be equal to the number of text keywords that are contained in both the text keywords associated with any second candidate subtext and the text keywords associated with the target sub-video. The matching device can take the maximum number of the number of text keywords associated with any second candidate subtext and the number of text keywords associated with the target sub-video as the maximum keyword number, and can obtain the ratio of the co-occurrence number to the maximum keyword number as the co-occurrence indication parameter between the text keywords associated with any second candidate subtext and the text keywords associated with the target sub-video. The higher the value of the co-occurrence indication parameter, the more matched the text keywords associated with any second candidate subtext and the text keywords associated with the target sub-video, and vice versa. The lower the value of the co-occurrence indication parameter, the less matched the text keywords associated with any second candidate subtext and the text keywords associated with the target sub-video.
[0295] The matching device can perform weighted summation on the feature similarity between each second candidate subtext and the target sub-video (i.e., the feature similarity between the text encoding features of each second candidate subtext and the video encoding features of the target sub-video) and the co-occurrence indication parameter to obtain the comprehensive matching degree between each second candidate subtext and the target sub-video.
[0296] The matching device can obtain a first matching weight (such as 0.7) preset for the feature similarity, and can obtain a second matching weight (such as 0.3) preset for the co-occurrence indication parameter. The matching device can perform weighted summation on the feature similarity and the co-occurrence indication parameter between any second candidate subtext and the target sub-video using the first matching weight and the second matching weight, i.e., the comprehensive matching degree between the second candidate subtext and the target sub-video can be calculated.
[0297] For example, the feature similarity between a second candidate subtext and the target sub-video can be denoted as D, the co-occurrence indication parameter between the second candidate subtext and the target sub-video can be denoted as G, the first matching weight can be denoted as q1, and the second matching weight can be denoted as q2. Therefore, the comprehensive matching degree between the second candidate subtext and the target sub-video can be equal to: q1xD+q2xG.
[0298] The matching device can obtain the comprehensive matching degree between each second candidate subtext and the target sub-video according to the above-described principle. Each second candidate subtext has a comprehensive matching degree with the target sub-video.
[0299] The matching device can obtain the comprehensive matching degree between each second candidate subtext and the target sub-video according to the above-described principle. Each second candidate subtext has a comprehensive matching degree with the target sub-video.
[0300] More, the video encoding features of the plurality of sub-videos and the text encoding features of the plurality of subtexts are generated by calling the feature encoding model. The following exemplary description describes the process of training the feature encoding model.
[0301] The matching device can obtain the feature encoding model to be trained and a sample set. The sample set can include a positive sample pair and a negative sample pair. The positive sample pair can include a first sample text and a first sample video. The text content described by the first sample text matches (such as consistent or nearly consistent) the video content presented by the first sample video. The negative sample pair can include a second sample text and a second sample video. The text content described by the second sample text does not match (such as a large difference) the video content presented by the second sample video.
[0302] The matching device can call the feature encoding model to be trained to perform feature encoding processing on the positive sample pair and the negative sample pair, respectively, to generate a first sample text encoding feature of the first sample text, a first sample video encoding feature of the first sample video, a second sample text encoding feature of the second sample text, and a second sample video encoding feature of the second sample video. That is, the matching device can perform feature encoding processing on the first sample text and the first sample video in the positive sample pair, respectively, to generate the first sample text encoding feature of the first sample text and the first sample video encoding feature of the first sample video. In addition, the matching device can perform feature encoding processing on the second sample text and the second sample video in the negative sample pair, respectively, to generate the second sample text encoding feature of the second sample text and the second sample video encoding feature of the second sample video. The first sample text encoding feature of the first sample text, the first sample video encoding feature of the first sample video, the second sample text encoding feature of the second sample text, and the second sample video encoding feature of the second sample video can be feature vectors.
[0303] The manner of generating the sample text encoding features of the sample text (e.g., the first sample text and the second sample text) by the feature encoding model to be trained is the same as the manner of generating the text encoding features of the subtext by the feature encoding model, i.e., the sample text encoding features can also be generated by the extracted text of the sample text in the plurality of second text extraction dimensions. Similarly, the manner of generating the sample video encoding features of the sample video (e.g., the first sample video and the second sample video) by the feature encoding model to be trained is the same as the manner of generating the video encoding features of the sub-video by the feature encoding model, i.e., the sample video encoding features can also be generated by the extracted text of the sample video in the plurality of first text extraction dimensions.
[0304] The matching device can generate a first feature encoding loss of the feature encoding model to be trained for the positive sample pair by the first sample text encoding features and the first sample video encoding features obtained above, e.g., the first feature encoding loss can be a cosine similarity loss between the first sample text encoding features and the first sample video encoding features.
[0305] Similarly, the matching device can generate a second feature encoding loss of the feature encoding model to be trained for the negative sample pair by the second sample text encoding features and the second sample video encoding features obtained above, e.g., the second feature encoding loss can be a cosine similarity loss between the second sample text encoding features and the second sample video encoding features.
[0306] The matching device can correct the model parameters of the feature encoding model to be trained by the first feature encoding loss and the second feature encoding loss to obtain the feature encoding model, which can include that the matching device can obtain a first training weight corresponding to the positive sample pair and a second training weight corresponding to the negative sample pair, the first training weight can be greater than the second training weight to improve the learning bias and learning effect of the model for the positive sample pair. For example, the first training weight can be equal to 0.7, and the second training weight can be equal to 0.3.
[0307] The matching device can weight and sum the first feature encoding loss and the second feature encoding loss by the first training weight and the second training weight to obtain a total encoding loss of the feature encoding model to be trained for the positive sample pair and the negative sample pair. For example, the total encoding loss can be equal to the product of the first training weight and the first feature encoding loss plus the loss obtained by the product of the second training weight and the second feature encoding loss.
[0308] The matching device can correct the model parameters of the feature encoding model to be trained by using the total encoding loss, and when the correction is completed (e.g., the iteration training round of the feature encoding model to be trained is equal to the set round threshold or the model parameters of the feature encoding model to be trained are corrected to a convergence state), the above-mentioned feature encoding model (i.e., the trained feature encoding model) can be obtained. The goal of correcting the model parameters of the feature encoding model to be trained by using the total encoding loss can be to correct the model parameters of the feature encoding model to be trained so that the total encoding loss tends to a minimum value (e.g., tends to 0), so that the effect achieved is that the feature encoding model after training can generate encoding features with smaller differences for the sample text and the sample video in the positive sample pair, and can generate encoding features with larger differences for the sample text and the sample video in the negative sample pair, so that the feature encoding model obtained after training can subsequently generate encoding features with smaller differences between texts and videos with content matching, and can generate encoding features with larger differences between texts and videos without content matching, thereby achieving the effect of greater distinction between texts and videos without content matching.
[0309] The purpose of correcting the model parameters of the feature encoding model to be trained by using the total encoding loss can also include maximizing (e.g., approaching 1) the cosine similarity between the first sample text encoding features and the first sample video encoding features generated by the feature encoding model to be trained for the positive sample pair, and minimizing (e.g., approaching -1 or a certain lower threshold) the cosine similarity between the second sample text encoding features and the second sample video encoding features generated by the feature encoding model to be trained for the negative sample pair.
[0310] In an embodiment, in order to enable the feature encoding model to be trained to have a greater distinction between texts and videos without content matching, the ratio between the positive sample pair and the negative sample pair used in the present application for training (i.e., model parameter correction) of the feature encoding model to be trained can be set as b1:b2, and b1 can be less than b2, e.g., b1:b2 can be equal to 1:4. In this way, the feature encoding model to be trained can have a higher distinction between various texts and videos without content matching.
[0311] In an embodiment, the feature encoding model can include a shared network and three prediction heads, which can be three feature encoding branches in the feature encoding model, and each of the three prediction heads can be connected after the shared network, and the three prediction heads can be initialized by a normal distribution for model parameters. The shared network can be a basic NLP (Natural Language Processing) model (or other natural language processing models can also be selected), which can be composed of many transformers (the core module of natural language processing), such as the basic NLP model can be a pre-trained BGE M3 model (a general semantic vector model). Each prediction head can be composed of an NLP layer (which can include a small number of transformers, and can belong to a content analysis layer) and a prediction layer (which can be a fully connected layer or the first few layers (such as the first 3 layers) of a bert transformer model (a text embedding model), which belongs to a linear transformation output layer).
[0312] Please refer to Figure 8 , Figure 8 is a structural diagram of a feature encoding model provided by an embodiment of the present application, as shown in Figure 8 , the feature encoding model can include a shared module (i.e. a shared network) and three feature encoding branches (including feature encoding branch 1-feature encoding branch 3), the input of the feature encoding model can be input into the shared module, the shared module can perform feature learning on the input extracted text (which can be an extracted text of a subtext or an extracted text of a subvideo), generate a corresponding feature vector, and input the feature vector into the feature encoding branch corresponding to the extracted text for deep feature encoding, i.e. to generate text encoding features of the extracted text, and finally perform feature fusion on the text encoding features generated by each feature encoding branch, i.e. to generate text encoding features corresponding to the subtext or video encoding features corresponding to the subvideo.
[0313] Among them, the feature vector can be stored by a dense-vector (a field used to store the generated embedding) during model training, and the longest input length of the feature encoding model can be set to ensure that the text input into the feature encoding model for feature encoding processing is within the longest input length, and the shorter the input length, the faster the inference speed.
[0314] The output of the shared network for the cls (the vector for the first word of the input text) can be taken as the output of the shared network in this application. When training the feature encoding model to be trained, the parameters of the last layer transformer of the shared network (which can be referred to as the parameters of the last feature encoding block of the shared network) and the parameters of the three prediction heads can be modified (which can be referred to as fine-tuning), and the other model parameters of the shared network (such as the BGE M3 model) except the parameters of the last layer transformer can not be modified (such as can be frozen), so as to ensure that the original text processing capability of the shared network is retained.
[0315] When constructing the above-mentioned positive sample pairs and negative sample pairs, a large number of sample subtexts of sample novels and sample subvideos of associated video resources can be collected, and the shared network (such as the pre-trained BGE M3 model) not trained by this application can be used to match each sample subvideo with the corresponding sample subtext, and find the top 5 (or other number) of sample subtexts corresponding to each sample subvideo. For example, the shared network can generate video encoding features for each sample subvideo and text encoding features for each sample subtext, and the top 5 sample subtexts corresponding to a sample subvideo can be the 5 sample subtexts with the most similar (such as the maximum cosine similarity) text encoding features and video encoding features. Thus, a positive sample pair can be constructed by a sample subvideo and the most similar (which can be determined by human judgment) sample subtext in the top 5, and four negative sample pairs can be constructed by the sample subvideo and the other four sample subtexts in the top 5, so that the ratio between the positive sample pairs and the negative sample pairs constructed in this way is 1:4.
[0316] The above-mentioned process of this application can be the total encoding loss obtained by the cosine similarity loss of the constructed positive sample pairs and negative sample pairs, which is used to train the shared network and multiple prediction heads end-to-end.
[0317] In another embodiment, in addition to the total encoding loss generated by the cosine similarity loss of the above-mentioned positive sample pairs and negative sample pairs to modify the model parameters of the feature encoding model to be trained, this application can also construct training data for each feature encoding branch to modify the model parameters of the feature encoding model to be trained, as described below.
[0318] The application can also construct a sample pair for each prediction head, which can include a first type of sample pair and a second type of sample pair. The first type of sample pair can be a positive sample pair, and the second type of sample pair can be a negative sample pair. The first type of sample pair and the second type of sample pair can be constructed by extracting texts of sample sub-videos (such as sub-videos of episode videos of sample video episodes, the extraction principle is the same as the above-mentioned sub-videos) in the above-mentioned multiple first text extraction dimensions, and extracting texts of sample sub-texts (such as sub-texts of sample text materials, the extraction principle is the same as the above-mentioned sub-texts) in the above-mentioned multiple second text extraction dimensions.
[0319] For example, the first type of sample pair can include: a sample pair composed of extracting texts of sample sub-videos in the above-mentioned picture extraction dimension and correct answers corresponding to the extracting texts, a sample pair composed of extracting texts of sample sub-videos in the above-mentioned character detection dimension and correct answers corresponding to the extracting texts, a sample pair composed of extracting texts of sample sub-videos in the above-mentioned content understanding dimension and correct answers corresponding to the extracting texts; and, a sample pair composed of extracting texts of sample sub-texts in the above-mentioned content extraction dimension and correct answers corresponding to the extracting texts, a sample pair composed of extracting texts of sample sub-texts in the above-mentioned dialogue extraction dimension and correct answers corresponding to the extracting texts, and a sample pair composed of extracting texts of sample sub-texts in the above-mentioned character extraction dimension and correct answers corresponding to the extracting texts.
[0320] And the second type of sample pair can include: a sample pair composed of extracting texts of sample sub-videos in the above-mentioned picture extraction dimension and incorrect answers corresponding to the extracting texts, a sample pair composed of extracting texts of sample sub-videos in the above-mentioned character detection dimension and incorrect answers corresponding to the extracting texts, a sample pair composed of extracting texts of sample sub-videos in the above-mentioned content understanding dimension and incorrect answers corresponding to the extracting texts; and, a sample pair composed of extracting texts of sample sub-texts in the above-mentioned content extraction dimension and incorrect answers corresponding to the extracting texts, a sample pair composed of extracting texts of sample sub-texts in the above-mentioned dialogue extraction dimension and incorrect answers corresponding to the extracting texts, and a sample pair composed of extracting texts of sample sub-texts in the above-mentioned character extraction dimension and incorrect answers corresponding to the extracting texts.
[0321] In an implementation, a correct answer corresponding to an extracted text can be a text having the same semantics (e.g., the same meaning of the expressed content) but different expressions (e.g., different key words / phrases) as the extracted text, and a wrong answer corresponding to an extracted text can be a text having different (e.g., opposite or significantly different) semantics and different expressions as the extracted text.
[0322] For example, for a picture extraction text of a sample sub-video in the picture extraction dimension, a content summary text of the sample sub-video in the content understanding dimension, a content extraction text of a sample sub-text in the content extraction dimension, and a dialogue text of the sample sub-text in the dialogue extraction dimension, a large language model trained can be used to perform text rewriting (e.g., text compression or re-expression) on the picture extraction text, the content summary text, the content extraction text, or the dialogue text, respectively, to generate a text having the same semantics but different expressions (which can be achieved by performing synonymous key word replacement) as the picture extraction text, the content summary text, the content extraction text, or the dialogue text, as the correct answer corresponding to the picture extraction text, the content summary text, the content extraction text, or the dialogue text.
[0323] Similarly, for a picture extraction text of a sample sub-video in the picture extraction dimension, a content summary text of the sample sub-video in the content understanding dimension, a content extraction text of a sample sub-text in the content extraction dimension, and a dialogue text of the sample sub-text in the dialogue extraction dimension, a large language model trained can be used to perform text rewriting (e.g., text compression or re-expression) on the picture extraction text, the content summary text, the content extraction text, or the dialogue text, respectively, to generate a text having different semantics and different expressions (which can be achieved by performing antonymous or semantically different key word replacement) as the picture extraction text, the content summary text, the content extraction text, or the dialogue text, as the wrong answer corresponding to the picture extraction text, the content summary text, the content extraction text, or the dialogue text.
[0324] For a role identification (e.g., a first role identification of a sample sub-video in the role detection dimension or a second role identification of a sample sub-text in the role extraction dimension), the same role identification can be used as the correct answer corresponding thereto, and a different role identification (as long as there is a difference) can be used as the wrong answer corresponding thereto.
[0325] The present application can generate an encoding loss for each feature encoding branch through the first type of sample pair and the second type of sample pair (which can be collectively referred to as a sample pair), which can be a cosine similarity loss. The cosine similarity loss can be the cosine similarity loss between the two text encoding features (feature vectors) generated and output by the corresponding feature encoding branch for the two texts in the sample pair, respectively. The present application can sum the encoding losses of each feature encoding branch and the total encoding loss as the total feature encoding loss (which can be referred to as the overall encoding loss) of the feature encoding model to be trained, and can correct the model parameters of the feature encoding model to be trained through the overall encoding loss. The matching device can perform multiple rounds and multiple batches (one round can include multiple batches) of iterative training on the feature encoding model to be trained according to this principle, until the average training loss (such as the overall encoding loss) of the feature encoding model to be trained for each batch of iterative training in a certain round no longer decreases (such as decreases below a certain loss threshold). At this time, the trained feature encoding model can be obtained.
[0326] Please refer to Figure 9 , Figure 9 is a principle diagram for obtaining a feature encoding loss of a feature encoding model to be trained provided by an embodiment of the present application. As shown in Figure 9 , the feature encoding model to be trained can include feature encoding branch 1-feature encoding branch 3. The picture extraction text of the sample sub-video and the dialogue text of the sample sub-text can correspond to feature encoding branch 1. The first role identifier of the sample sub-video and the second role identifier of the sample sub-text can correspond to feature encoding branch 2. The content summary text of the sample sub-video and the content extraction text of the sample sub-text can correspond to feature encoding branch 3.
[0327] Therefore, the feature encoding branch 1 in the feature encoding model to be trained can be called to obtain the cosine similarity of the first type of sample pair or the second type of sample pair composed of the picture extraction text of the sample sub-video and its correct answer or / and incorrect answer, as the encoding loss 1 here. The encoding loss 1 can be the cosine similarity loss between the text encoding features generated by the feature encoding branch 1 for the picture extraction text of the sample sub-video and its correct answer or / and incorrect answer. Moreover, the feature encoding branch 1 in the feature encoding model to be trained can be called to obtain the cosine similarity of the first type of sample pair or the second type of sample pair composed of the dialogue text of the sample sub-text and its correct answer or / and incorrect answer, as the encoding loss 4 here. The encoding loss 4 can be the cosine similarity loss between the text encoding features generated by the feature encoding branch 1 for the dialogue text of the sample sub-text and its correct answer or / and incorrect answer.
[0328] Similarly, the feature encoding branch 2 in the feature encoding model to be trained can be called to obtain the cosine similarity of the first role label of the sample sub-video and the first type sample pair or the second type sample pair composed of the correct answer and / or the wrong answer thereof, as the encoding loss 2 here, which can be the cosine similarity loss between the text encoding features generated by the feature encoding branch 2 for the first role label of the sample sub-video and the correct answer and / or the wrong answer thereof. And the feature encoding branch 2 in the feature encoding model to be trained can be called to obtain the cosine similarity of the second role label of the sample sub-text and the first type sample pair or the second type sample pair composed of the correct answer and / or the wrong answer thereof, as the encoding loss 5 here, which can be the cosine similarity loss between the text encoding features generated by the feature encoding branch 2 for the second role label of the sample sub-text and the correct answer and / or the wrong answer thereof.
[0329] And the feature encoding branch 3 in the feature encoding model to be trained can be called to obtain the cosine similarity of the content summary text of the sample sub-video and the first type sample pair or the second type sample pair composed of the correct answer and / or the wrong answer thereof, as the encoding loss 3 here, which can be the cosine similarity loss between the text encoding features generated by the feature encoding branch 3 for the content summary text of the sample sub-video and the correct answer and / or the wrong answer thereof. And the feature encoding branch 3 in the feature encoding model to be trained can be called to obtain the cosine similarity of the content extraction text of the sample sub-text and the first type sample pair or the second type sample pair composed of the correct answer and / or the wrong answer thereof, as the encoding loss 6 here, which can be the cosine similarity loss between the text encoding features generated by the feature encoding branch 3 for the content extraction text of the sample sub-text and the correct answer and / or the wrong answer thereof.
[0330] And the matching device can also obtain the video encoding features generated by the feature encoding model to be trained for the sample sub-video through the picture extraction text, the first role label and the content summary text of the sample sub-video, and can obtain the text encoding features generated by the feature encoding model to be trained for the sample sub-text through the dialogue text, the second role label and the content extraction text of the sample sub-text, so as to obtain the encoding loss 7 through the cosine similarity loss between the video encoding features and the text encoding features. The sample sub-video and the sample sub-text can constitute the above-mentioned positive sample pair and negative sample pair, and the encoding loss 7 can be the total encoding loss for the positive sample pair and the negative sample pair.
[0331] The matching device can take the sum of the coding loss 1 to the coding loss 7 as an overall coding loss of the feature coding model to be trained, and can correct the model parameters of the feature coding model to be trained through the overall coding loss, so as to obtain the feature coding model. Wherein, the purpose of correcting the model parameters of the feature coding model to be trained through the overall coding loss is to maximize the cosine similarity (which can be the cosine similarity between the text coding features of the two contained texts) corresponding to the positive first type sample pair and the positive sample pair, and to minimize the cosine similarity corresponding to the negative second type sample pair and the negative sample pair.
[0332] By using the above method provided in the present application, the feature coding model to be trained can be accurately trained through the constructed positive sample pair and negative sample pair, so that the feature coding model trained can generate similar coding features for texts and videos with content matching, and generate coding features with large differences for texts and videos without content matching, so as to realize highly accurate content matching processing between sub-videos and sub-texts by using the coding features (such as the video coding features and the text coding features) generated by the feature coding model.
[0333] Please refer to Figure 10 , Figure 10 is a structural schematic diagram of a data matching framework provided by an embodiment of the present application. As shown in Figure 10 , the data matching framework can include a video processing part, a text processing part, and a front-end display part. The video processing part can include five modules, i.e., module m1 to module m5. The module m1 can be used to obtain episode videos that need to be subjected to content matching, such as the target video resource. The module m2 can be used to perform video segmentation processing on the episode videos to obtain the plurality of sub-videos. The module m3 can be used to convert the videos into texts, such as obtaining the extracted texts of the sub-videos in each first text extraction dimension. The module m4 can be used to perform feature extraction on the extracted texts of the sub-videos to generate video coding features of the sub-videos. The cross-theme alignment model here can be the feature coding model of the present application, which can be used to generate video coding features of the sub-videos and text coding features of the sub-texts. The module m5 can be used to aggregate the multi-episode videos, which can be a movie or a TV series, i.e., the video resource set. The aggregation here can refer to aggregation of the video coding features of each video resource in the video resource set. In addition, the module m10 can be a matching database, which can store the aggregated video coding features of each video resource in the matching database.
[0334] The text processing part can include four modules, i.e., a module m6, a module m7, a module m8 and a module m9. The module m6 can be configured to obtain each novel chapter (i.e., a text chapter) in a novel (i.e., a text material). The module m7 can be configured to perform paragraph segmentation on the novel chapter to obtain each subtext. One subtext can be one paragraph in one novel chapter. The module m8 can be configured to perform feature extraction on the obtained subtext to generate a text encoding feature of the subtext. The text encoding feature can also be generated by the cross-genre alignment model. The module m9 can be configured to perform multi-chapter summarization, such as summarizing the text encoding features of the subtexts obtained by dividing each chapter, and store the text encoding features of the subtexts obtained by summarization in the matching database of the module m10.
[0335] After obtaining the text encoding features of each subtext and the video encoding features of each video, the text range in the novel that matches each episode video can be obtained. The front-end display part can include four modules, i.e., a module m11, a module m12, a module m13 and a module m14. The module m11 can refer to user watching triggering in the front end (such as a video client), such as playing triggering of a target video resource. The module m12 can refer to that when the user triggers the playing of the target video resource, the video client can pre-load the reading link of the text range corresponding to the target video resource, and can display a reading control in the playing interface of the target video resource when the target video resource is about to end (such as playing the end of the main part and starting to play the end). The module m13 can refer to that the user clicks the reading control displayed in the playing interface of the target video resource to trigger the display of the reading interface. The module m14 can display the reading interface in response to the clicking of the reading control by the user. The reading interface can be an interface in the video client for displaying and reading the text in the text range matching the target video resource.
[0336] By using the above method provided in the present application, the user can trigger the reading of the text matching the played video when playing the video, so that the user can understand the corresponding plot from both the video dimension and the text dimension, and the interest of the user in watching the video is improved, and the reading experience of the user in watching the video is enriched.
[0337] Please refer to Figure 11 , Figure 11 is a structural schematic diagram of a data matching device provided by an embodiment of the present application. As shown in Figure 11 , the data matching device 110 can include an obtaining module 1101, a segmentation module 1102, a matching module 1103 and a determining module 1104.
[0338] The acquisition module 1101 is configured to acquire a text material associated with a video resource set and a target video resource, the text content described by the text material matches the video content presented by the video resource set, and the target video resource is any video resource in the video resource set.
[0339] The segmentation module 1102 is configured to perform video segmentation processing on the target video resource to obtain a plurality of sub-videos of the target video resource, and perform text segmentation processing on the text material to obtain a plurality of sub-texts of the text material.
[0340] The matching module 1103 is configured to perform content matching processing between the plurality of sub-videos and the plurality of sub-texts to obtain a plurality of matching pairs, one matching pair including one sub-video and one sub-text, and the sub-video and the sub-text in the same matching pair have content matching.
[0341] The determination module 1104 is configured to determine a text range in the text material that matches the target video resource according to the plurality of matching pairs.
[0342] In an implementation, the segmentation module 1102 performs video segmentation processing on the target video resource to obtain a plurality of sub-videos of the target video resource in the following manner:
[0343] Perform scene detection processing on the video frames in the target video resource to obtain at least one transition video frame in the target video resource, one transition video frame being used to indicate that a video scene in the target video resource changes;
[0344] According to the positions of the at least one transition video frame in the target video resource, perform video segmentation processing on the target video resource to obtain the plurality of sub-videos.
[0345] In an implementation, the text material includes a plurality of text chapters; the segmentation module 1102 performs text segmentation processing on the text material to obtain a plurality of sub-texts of the text material in the following manner:
[0346] Detect a segmentation position in each text chapter of the text material;
[0347] According to the segmentation position detected in each text chapter, respectively perform text segmentation processing on each text chapter to obtain the plurality of sub-texts;
[0348] Any sub-text is any paragraph obtained by segmenting any text chapter of the text material.
[0349] In an implementation, the matching module 1103 performs content matching processing between the plurality of sub-videos and the plurality of sub-texts to obtain a plurality of matching pairs in the following manner:
[0350] Multiple sub-videos are processed by feature encoding to generate video encoding features for each sub-video. The video encoding features of any sub-video are used to characterize the video content presented by any sub-video.
[0351] Multiple subtexts are processed by feature encoding to generate text encoding features for each subtext. The text encoding features of any subtext are used to characterize the text content described by any subtext.
[0352] Based on the video encoding features of each sub-video and the text encoding features of each sub-text, content matching is performed between multiple sub-videos and multiple sub-texts to obtain multiple matching pairs.
[0353] In one implementation, any sub-video is the target sub-video; the matching module 1103 performs feature encoding processing on multiple sub-videos respectively, generating video encoding features for each sub-video in the following ways:
[0354] The video frames in the target sub-video are extracted to obtain multiple extracted video frames in the target sub-video.
[0355] Text extraction processing is performed on multiple extracted video frames from multiple first text extraction dimensions to generate target sub-videos with extracted text under each first text extraction dimension.
[0356] The extracted text of the target sub-video under multiple first text extraction dimensions is subjected to feature encoding processing to generate video encoding features of the target sub-video.
[0357] In one implementation, the multiple first text extraction dimensions include a scene extraction dimension, a character detection dimension, and a content understanding dimension; the matching module 1103 performs text extraction processing on multiple extracted video frames from the multiple first text extraction dimensions to generate the extracted text of the target sub-video under each first text extraction dimension, including:
[0358] The text appearing in multiple extracted video frames is extracted from the image extraction dimension to obtain the image extracted text from multiple extracted video frames.
[0359] Face detection processing is performed on multiple extracted video frames from the perspective of character detection to obtain the first character identifier of the film and television character to which the face in the multiple extracted video frames belongs.
[0360] From the perspective of content understanding, multiple extracted video frames are processed to generate content summary text of the video content presented by the multiple extracted video frames;
[0361] Among them, the extracted text of the target sub-video under multiple first text extraction dimensions includes the image extraction text, the first character identifier, and the content summary text.
[0362] In an implementation, the matching module 1103 encodes features of the extracted text of the target sub-video in the plurality of first text extraction dimensions to generate video encoding features of the target sub-video in the following manner:
[0363] obtains a feature encoding model, the feature encoding model including a feature encoding branch corresponding to each first text extraction dimension;
[0364] invokes each feature encoding branch to encode features of the extracted text of the target sub-video in the corresponding first text extraction dimension, to generate first text encoding features of the extracted text of the target sub-video in each first text extraction dimension;
[0365] fuses the plurality of first text encoding features of the extracted text of the target sub-video in the plurality of first text extraction dimensions to generate the video encoding features of the target sub-video.
[0366] In an implementation, any subtext is a target subtext; the matching module 1103 encodes features of the plurality of subtexts respectively to generate text encoding features of each subtext in the following manner:
[0367] performs text extraction processing on the target subtext from a plurality of second text extraction dimensions to generate extracted text of the target subtext in each second text extraction dimension;
[0368] encodes features of the extracted text of the target subtext in the plurality of second text extraction dimensions to generate text encoding features of the target subtext.
[0369] In an implementation, the plurality of second text extraction dimensions include a content extraction dimension, a dialogue extraction dimension, and a character extraction dimension; the matching module 1103 performs text extraction processing on the target subtext from the plurality of second text extraction dimensions to generate extracted text of the target subtext in each second text extraction dimension in the following manner:
[0370] performs content extraction processing on the target subtext from the content extraction dimension to obtain content extraction text of the target subtext;
[0371] performs dialogue extraction processing on the target subtext from the dialogue extraction dimension to obtain dialogue text in the target subtext;
[0372] performs character extraction processing on the target subtext from the character extraction dimension to obtain a second character identifier of a film and television character described in the target subtext;
[0373] wherein the extracted text of the target subtext in the plurality of second text extraction dimensions includes the content extraction text, the dialogue text, and the second character identifier.
[0374] In an implementation, the matching module 1103 extracts the content of the target subtext according to the content extraction dimension to obtain the content extraction text of the target subtext in the following manner:
[0375] obtaining the text length of the target subtext;
[0376] if the text length is less than or equal to the set length threshold, regarding the target subtext as the content extraction text;
[0377] if the text length is greater than the length threshold, performing content compression processing on the target subtext to generate the content extraction text.
[0378] In an implementation, the matching module 1103 encodes the features of the extraction text of the target subtext in the plurality of second text extraction dimensions to generate the text encoding features of the target subtext in the following manner:
[0379] obtaining a feature encoding model, the feature encoding model including a feature encoding branch corresponding to each second text extraction dimension;
[0380] calling each feature encoding branch to encode the features of the extraction text of the target subtext in the corresponding second text extraction dimension to generate the second text encoding features of the extraction text of the target subtext in each second text extraction dimension;
[0381] performing feature fusion processing on the plurality of second text encoding features of the extraction text of the target subtext in the plurality of second text extraction dimensions to generate the text encoding features of the target subtext.
[0382] In an implementation, any sub-video is a target sub-video; the matching module 1103 performs content matching processing between the plurality of sub-videos and the plurality of sub-texts based on the video encoding features of each sub-video and the text encoding features of each sub-text to obtain a plurality of matching pairs in the following manner:
[0383] obtaining the feature similarity between the text encoding features of the plurality of sub-texts and the video encoding features of the target sub-video;
[0384] sorting the plurality of sub-texts according to the feature similarity between the text encoding features of each sub-text and the video encoding features of the target sub-video in descending order to obtain sorted sub-texts;
[0385] determining the K sub-texts arranged in the front of the sorted sub-texts as first candidate sub-texts, K being a positive integer;
[0386] determining the sub-text matched with the target sub-video from the K first candidate sub-texts;
[0387] The target sub-video and the matched sub-text are used to form a matching pair.
[0388] In an embodiment, the matching module 1103 determines the sub-text matched with the target sub-video from the K first candidate sub-texts in the following manner:
[0389] Obtain the first role identifier of the movie / TV role appearing in the target sub-video;
[0390] Determine the first candidate sub-text including the first role identifier in the K first candidate sub-texts as the second candidate sub-text;
[0391] Determine the sub-text matched with the target sub-video from the at least one second candidate sub-text.
[0392] In an embodiment, the matching module determines the sub-text matched with the target sub-video from the at least one second candidate sub-text in the following manner:
[0393] Calculate the comprehensive matching degree between each second candidate sub-text and the target sub-video respectively;
[0394] Determine the second candidate sub-text with the maximum comprehensive matching degree between the at least one second candidate sub-text and the target sub-video as the sub-text matched with the target sub-video.
[0395] In an embodiment, the matching module 1103 calculates the comprehensive matching degree between each second candidate sub-text and the target sub-video respectively in the following manner:
[0396] Obtain the feature similarity between the text encoding features of each second candidate sub-text and the video encoding features of the target sub-video respectively;
[0397] Calculate the co-occurrence indication parameter between the text keywords associated with each second candidate sub-text and the text keywords associated with the target sub-video respectively;
[0398] Weighted sum the feature similarity and the co-occurrence indication parameter between each second candidate sub-text and the target sub-video to obtain the comprehensive matching degree between each second candidate sub-text and the target sub-video respectively.
[0399] In an embodiment, the video encoding features of the plurality of sub-videos and the text encoding features of the plurality of sub-texts are generated by calling a feature encoding model; and the matching module 1103 is further configured to:
[0400] obtain a feature encoding model to be trained and a sample set, the sample set including positive sample pairs and negative sample pairs, the positive sample pairs including first sample texts and first sample videos, the text content described by the first sample texts matching the video content presented by the first sample videos, the negative sample pairs including second sample texts and second sample videos, the text content described by the second sample texts not matching the video content presented by the second sample videos;
[0401] call the feature encoding model to be trained to perform feature encoding processing on the positive sample pairs and the negative sample pairs respectively, to generate first sample text encoding features of the first sample texts, first sample video encoding features of the first sample videos, second sample text encoding features of the second sample texts, and second sample video encoding features of the second sample videos;
[0402] generate a first feature encoding loss of the feature encoding model to be trained for the positive sample pairs based on the first sample text encoding features and the first sample video encoding features, and generate a second feature encoding loss of the feature encoding model to be trained for the negative sample pairs based on the second sample text encoding features and the second sample video encoding features;
[0403] correct the model parameters of the feature encoding model to be trained by using the first feature encoding loss and the second feature encoding loss, to obtain the feature encoding model.
[0404] In an implementation manner, the matching module 1103 corrects the model parameters of the feature encoding model to be trained by using the first feature encoding loss and the second feature encoding loss, to obtain the feature encoding model in the following manner:
[0405] obtain a first training weight corresponding to the positive sample pairs and a second training weight corresponding to the negative sample pairs, the first training weight being greater than the second training weight;
[0406] weight and sum the first feature encoding loss and the second feature encoding loss by using the first training weight and the second training weight, to obtain a total encoding loss of the feature encoding model to be trained for the positive sample pairs and the negative sample pairs;
[0407] correct the model parameters of the feature encoding model to be trained by using the total encoding loss, to obtain the feature encoding model.
[0408] In an implementation manner, the text material includes a plurality of text chapters; the determining module 1104 determines the text range in the text material that matches the target video resource according to the plurality of matching pairs in the following manner:
[0409] count Z text chapters in the text material in which the subtexts in the plurality of matching pairs are hit and the number of hits of the subtexts in each hit text chapter, any hit text chapter including at least one subtext in the plurality of matching pairs, Z being a positive integer;
[0410] The determination module 1104 determines each hit text chapter as a non-divergent point based on the number of hits corresponding to each hit text chapter.
[0411] The determination module 1104 removes the hit text chapters determined as the divergent points from the Z hit text chapters to obtain at least one reserved text chapter.
[0412] The determination module 1104 determines the chapter range of the at least one reserved text chapter in the text material as the text range matched with the target video resource.
[0413] In an embodiment, any hit text chapter is a target text chapter; the determination module 1104 determines each hit text chapter as a non-divergent point based on the number of hits corresponding to each hit text chapter in the following manner:
[0414] The determination module 1104 calculates a first hit density between the hit subtexts in the Z hit text chapters, and obtains a density reference threshold for the subtext matching based on the first hit density.
[0415] The determination module 1104 calculates a second hit density between the hit subtexts in the Z-1 hit text chapters other than the target text chapter.
[0416] The determination module 1104 counts a first number of the hit subtexts in the Z-1 hit text chapters and a second number of the hit subtexts in the Z hit text chapters, and obtains a number ratio between the first number and the second number.
[0417] If the number ratio is greater than or equal to a preset ratio threshold and the second hit density is greater than or equal to the density reference threshold, the determination module 1104 determines that the target text chapter is determined as a divergent point.
[0418] If the number ratio is less than the ratio threshold or the second hit density is less than the density reference threshold, the determination module 1104 determines that the target text chapter is determined as a non-divergent point.
[0419] In an embodiment, the determination module 1104 calculates the first hit density between the hit subtexts in the Z hit text chapters in the following manner:
[0420] The determination module 1104 fills in the missing text chapters between the Z hit text chapters to obtain L continuous text chapters corresponding to the Z hit text chapters, L being a positive integer and L being greater than or equal to Z.
[0421] The determination module 1104 calculates a ratio between the total number of the hit subtexts in the Z hit text chapters and L as the first hit density.
[0422] In an embodiment, the determination module 1104 obtains the density reference threshold for the subtext matching based on the first hit density in the following manner:
[0423] The ratio between the total number of the hit subtexts in the Z number of text chapters and Z is calculated as the maximum hit density;
[0424] The difference between the maximum hit density and the first hit density is obtained, and the product between the difference and a set density control parameter is calculated as a density fluctuation value;
[0425] The sum of the density fluctuation value and the first hit density is taken as the density reference threshold.
[0426] In an embodiment, the determination module 1104 is further configured to:
[0427] When playing the target video resource, the reading link corresponding to the text range matched with the target video resource is pushed to the video client;
[0428] The video client is configured to display the reading link when playing the target video resource, and is configured to display the text in the text range matched with the target video resource in the text material in response to a trigger operation on the reading link.
[0429] According to an embodiment of the present application, Figure 3 The steps involved in the data matching method shown can be performed by Figure 11 The modules in the data matching apparatus 110 shown. For example, Figure 3 The step S101 shown in the method 100 can be performed by Figure 11 the obtaining module 1101 in the apparatus 100, Figure 3 The step S102 shown in the method 100 can be performed by Figure 11 the splitting module 1102 in the apparatus 100; Figure 3 The step S103 shown in the method 100 can be performed by Figure 11 the matching module 1103 in the apparatus 100, Figure 3 The step S104 shown in the method 100 can be performed by Figure 11 the determination module 1104 in the apparatus 100.
[0430] The application can obtain text material associated with a video resource set and a target video resource, the text content described by the text material matches the video content presented by the video resource set, and the target video resource is any video resource in the video resource set; and the target video resource can be subjected to video segmentation processing to obtain multiple sub-videos of the target video resource, and the text material can be subjected to text segmentation processing to obtain multiple sub-texts of the text material; thus, content matching processing can be performed between the multiple sub-videos and the multiple sub-texts to obtain multiple matching pairs, one matching pair includes one sub-video and one sub-text, and the sub-video and the sub-text in the same matching pair have content matching; thus, the text range in the text material that matches the target video resource can be determined according to the multiple matching pairs. As can be seen, the device proposed in the application can segment the target video resource into multiple sub-videos of smaller units, and can also segment the text material associated with the video resource set into multiple sub-texts of smaller units, so that the accurate determination of the text range in the text material that matches the target video resource can be realized through the refined content matching between the multiple sub-videos and the multiple sub-texts of smaller units.
[0431] According to an embodiment of the application, Figure 11 Each module in the data matching device 110 shown can be combined into one or several units to constitute, or some of the units can be further split into a plurality of sub-units with smaller functions, and the same operations can be realized without affecting the implementation of the technical effects of the embodiments of the application. The above modules are divided based on logical functions, and in actual applications, the functions of one module can also be realized by multiple units, or the functions of multiple modules can be realized by one unit. In other embodiments of the application, the data matching device 110 can also include other units, and in actual applications, these functions can also be realized by other units and can be realized by multiple units in cooperation.
[0432] In the embodiments of the application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of the module or unit.
[0433] According to an embodiment of the present application, the data matching apparatus 110 as shown in Figure 11 FIG. 36 can be constructed by running a computer program capable of performing each step involved in the corresponding method shown in each embodiment of the present application on a general computer device which can include a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM), and the like processing and storage elements.
[0434] Please refer to Figure 12 , Figure 12 FIG. 1 is a structural schematic diagram of a computer device provided by an embodiment of the present application. As shown in Figure 12 FIG. 1, the computer device 1000 can include a processor 1001, a network interface 1004, and a memory 1005, and in some embodiments, the computer device 1000 can further include a user interface 1003 and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 can include a display, a keyboard, and optionally the user interface 1003 can further include a standard wired interface and a wireless interface. The network interface 1004 can optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 can be a high-speed RAM memory or a non-volatile memory such as at least one disk memory. The memory 1005 can optionally be at least one storage device located away from the aforementioned processor 1001. As shown in Figure 12 FIG. 1, the memory 1005 as a computer storage medium can include an operating system, a network communication module, a user interface module, and a device control application program.
[0435] In the computer device 1000 as shown in Figure 12 FIG. 1, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an input interface for a user; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to execute the description of the aforementioned data matching method in each embodiment of the present application, and can also execute the description of the aforementioned data matching apparatus 110 in the aforementioned corresponding embodiment, which will not be described here again. In addition, the description of the beneficial effects of using the same method will also not be described here again. Figure 11
[0436] In addition, it should be noted that the present application also provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When a processor executes the computer program, the processor can execute the description of the data matching method in the embodiments of the present application. Therefore, the description will not be repeated. In addition, the beneficial effects of the same method will not be repeated. For technical details not disclosed in the computer storage medium embodiments of the present application, please refer to the description of the method embodiments of the present application.
[0437] As an example, the above computer program can be deployed on a computer device to execute, or deployed on multiple computer devices in a location to execute, or deployed on multiple computer devices distributed in multiple locations and interconnected through a communication network to execute. The multiple computer devices distributed in multiple locations and interconnected through a communication network can constitute a blockchain network.
[0438] The above computer readable storage medium can be an internal storage unit of the above computer device, such as a hard disk or a memory of the computer device. The computer readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the computer readable storage medium can include both the internal storage unit and the external storage device of the computer device. The computer readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer readable storage medium can also be used to temporarily store data that has been output or will be output.
[0439] The present application provides a computer program product including a computer program stored in a computer readable storage medium. The processor of the computer device reads the computer program from the computer readable storage medium. The processor executes the computer program, so that the computer device executes the description of the above data matching method in the embodiments of the present application. Therefore, the description will not be repeated. In addition, the beneficial effects of the same method will not be repeated. For technical details not disclosed in the computer readable storage medium embodiments of the present application, please refer to the description of the method embodiments of the present application.
[0440] The terms "first", "second", etc. in the specification and claims of the present application and the drawings are used to distinguish different objects, and are not used to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or apparatus that includes a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to such processes, methods, devices, products or apparatus.
[0441] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, each example has been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0442] The above disclosure is only the preferred embodiments of the present application, and of course cannot limit the scope of the rights of the present application, so the equivalent changes made in accordance with the claims of the present application are still within the scope of the present application.
Claims
1. A data matching method, characterized by, The method comprises: obtaining a text material associated with a video resource set and a target video resource, the text content described by the text material matching the video content presented by the video resource set, and the target video resource being any video resource in the video resource set; performing video segmentation processing on the target video resource to obtain a plurality of sub-videos of the target video resource, and performing text segmentation processing on the text material to obtain a plurality of sub-texts of the text material; performing content matching processing between the plurality of sub-videos and the plurality of sub-texts to obtain a plurality of matching pairs, one matching pair comprising one sub-video and one sub-text, and the sub-video and the sub-text in the same matching pair having content matching; determining a text range in the text material that matches the target video resource according to the plurality of matching pairs.
2. The method of claim 1, wherein, The video segmentation processing on the target video resource to obtain a plurality of sub-videos of the target video resource comprises: performing scene detection processing on the video frames in the target video resource to obtain at least one transition video frame in the target video resource, one transition video frame being used to indicate that a video scene in the target video resource changes; performing video segmentation processing on the target video resource according to the positions of the at least one transition video frame in the target video resource to obtain the plurality of sub-videos.
3. The method of claim 1, wherein, The content matching processing between the plurality of sub-videos and the plurality of sub-texts to obtain a plurality of matching pairs comprises: performing feature encoding processing on the plurality of sub-videos respectively to generate video encoding features of each sub-video, the video encoding features of any sub-video being used to represent the video content presented by any sub-video; performing feature encoding processing on the plurality of sub-texts respectively to generate text encoding features of each sub-text, the text encoding features of any sub-text being used to represent the text content described by any sub-text; based on the video encoding features of each sub-video and the text encoding features of each sub-text, performing content matching processing between the plurality of sub-videos and the plurality of sub-texts to obtain the plurality of matching pairs.
4. The method of claim 3, wherein, Any sub-video is a target sub-video; the feature encoding processing on the plurality of sub-videos respectively to generate video encoding features of each sub-video comprises: performing extraction processing on the video frames in the target sub-video to obtain a plurality of extracted video frames in the target sub-video; performing text extraction processing on the plurality of extracted video frames from a plurality of first text extraction dimensions to generate extracted texts of the target sub-video in each first text extraction dimension; performing feature encoding processing on the extracted texts of the target sub-video in the plurality of first text extraction dimensions to generate video encoding features of the target sub-video.
5. The method of claim 4, wherein, The first text extraction dimensions include a picture extraction dimension, a character detection dimension, and a content understanding dimension; the text extraction processing of the extracted video frames in the first text extraction dimensions generates the extraction text of the target sub-video in each of the first text extraction dimensions, including: The text extraction processing of the text appearing in the picture of the extracted video frames in the picture extraction dimension generates the picture extraction text of the extracted video frames; The face detection processing of the extracted video frames in the character detection dimension generates the first character identification of the character in the face of the extracted video frames; The content understanding processing of the extracted video frames in the content understanding dimension generates the content summary text of the video content presented by the extracted video frames; The extraction text of the target sub-video in the first text extraction dimensions includes the picture extraction text, the first character identification, and the content summary text.
6. The method of claim 3, wherein, Any of the sub-texts is a target sub-text; the feature coding processing of the sub-texts generates the text coding features of each of the sub-texts, including: The text extraction processing of the target sub-text in the second text extraction dimensions generates the extraction text of the target sub-text in each of the second text extraction dimensions; The feature coding processing of the extraction text of the target sub-text in the second text extraction dimensions generates the text coding features of the target sub-text.
7. The method of claim 6, wherein, The second text extraction dimensions include a content extraction dimension, a dialogue extraction dimension, and a character extraction dimension; the text extraction processing of the target sub-text in the second text extraction dimensions generates the extraction text of the target sub-text in each of the second text extraction dimensions, including: The content extraction processing of the target sub-text in the content extraction dimension generates the content extraction text of the target sub-text; The dialogue extraction processing of the target sub-text in the dialogue extraction dimension generates the dialogue text in the target sub-text; The character extraction processing of the target sub-text in the character extraction dimension generates the second character identification of the character described in the target sub-text; The extraction text of the target sub-text in the second text extraction dimensions includes the content extraction text, the dialogue text, and the second character identification.
8. The method of claim 3, wherein, Any of the sub-videos is a target sub-video; the content matching processing between the sub-videos and the sub-texts based on the video coding features of each of the sub-videos and the text coding features of each of the sub-texts generates the matching pairs, including: The feature similarity between the text coding features of the sub-texts and the video coding features of the target sub-video is obtained; The sub-texts are sorted in the order of the feature similarity between the text coding features of each of the sub-texts and the video coding features of the target sub-video from large to small, and the sorted sub-texts are obtained; Determine K subtexts arranged in front of the sorted subtexts as first candidate subtexts, K being a positive integer; Determine a subtext matching the target sub-video from the K first candidate subtexts; The target sub-video and the matching subtext are used to constitute a matching pair.
9. The method of claim 8, wherein, The determining a subtext matching the target sub-video from the K first candidate subtexts comprises: Obtain a first role identifier of a film and television role appearing in the target sub-video; Determine a first candidate subtext including the first role identifier in the K first candidate subtexts as a second candidate subtext; Determine a subtext matching the target sub-video from at least one second candidate subtext.
10. The method of claim 9, wherein, The determining a subtext matching the target sub-video from at least one second candidate subtext comprises: Calculate a comprehensive matching degree between each second candidate subtext and the target sub-video respectively; Determine a second candidate subtext with the largest comprehensive matching degree between the target video as the subtext matching the target sub-video from at least one second candidate subtext.
11. The method of claim 10, wherein, The calculating a comprehensive matching degree between each second candidate subtext and the target sub-video respectively comprises: Obtain a feature similarity between a text encoding feature of each second candidate subtext and a video encoding feature of the target sub-video respectively; Calculate a co-occurrence indication parameter between a text keyword associated with each second candidate subtext and a text keyword associated with the target sub-video respectively; Weighted sum the feature similarity and the co-occurrence indication parameter between each second candidate subtext and the target sub-video to obtain a comprehensive matching degree between each second candidate subtext and the target sub-video respectively.
12. The method of claim 3, wherein, The video encoding features of the plurality of sub-videos and the text encoding features of the plurality of subtexts are generated by calling a feature encoding model; the method further comprises: Obtain a feature encoding model to be trained and a sample set, the sample set comprising positive sample pairs and negative sample pairs, the positive sample pairs comprising a first sample text and a first sample video, the text content described by the first sample text matching the video content presented by the first sample video, the negative sample pairs comprising a second sample text and a second sample video, the text content described by the second sample text not matching the video content presented by the second sample video; Call the feature encoding model to be trained to perform feature encoding processing on the positive sample pairs and the negative sample pairs respectively to generate first sample text encoding features of the first sample text, first sample video encoding features of the first sample video, second sample text encoding features of the second sample text, and second sample video encoding features of the second sample video; generate a first feature encoding loss of the feature encoding model to be trained for the positive sample pair based on the first sample text encoding feature and the first sample video encoding feature, and generate a second feature encoding loss of the feature encoding model to be trained for the negative sample pair based on the second sample text encoding feature and the second sample video encoding feature; correct the model parameters of the feature encoding model to be trained by using the first feature encoding loss and the second feature encoding loss, to obtain the feature encoding model.
13. The method of claim 1, wherein, The text material includes a plurality of text chapters; determining the text range in the text material matched with the target video resource according to a plurality of matching pairs, comprises: statistically counting Z hit text chapters of subtexts in the text material in a plurality of matching pairs and the number of hits of subtexts in each hit text chapter, any hit text chapter includes at least one subtext in a plurality of matching pairs, and Z is a positive integer; based on the respective hit number of each hit text chapter, performing isolated point detection processing on each hit text chapter; removing the hit text chapter detected as an isolated point in Z hit text chapters to obtain at least one reserved text chapter; determining the chapter range of the at least one reserved text chapter in the text material as the text range matched with the target video resource.
14. The method of claim 13, wherein, Any hit text chapter is a target text chapter; based on the respective hit number of each hit text chapter, performing isolated point detection processing on each hit text chapter, comprises: calculating a first hit density between the hit subtexts in Z hit text chapters, and obtaining a density reference threshold for subtext matching based on the first hit density; calculating a second hit density between the hit subtexts in Z-1 hit text chapters other than the target text chapter in Z hit text chapters; statistically counting a first number of hit subtexts in the Z-1 hit text chapters and a second number of hit subtexts in Z hit text chapters, and obtaining a quantity ratio between the first number and the second number; if the quantity ratio is greater than or equal to a preset ratio threshold, and the second hit density is greater than or equal to the density reference threshold, it is determined that the target text chapter is detected as an isolated point; if the quantity ratio is less than the ratio threshold, or the second hit density is less than the density reference threshold, it is determined that the target text chapter is detected as a non-isolated point.
15. The method of claim 14, wherein, The calculation of the first hit density between the hit subtexts in Z hit text chapters comprises: filling in the missing text chapters between Z hit text chapters to obtain L continuous text chapters corresponding to Z hit text chapters, L is a positive integer and L is greater than or equal to Z; the ratio between the total number of hit subtexts in Z hit text chapters and L is calculated as the first hit density.
16. The method of claim 14, wherein, The density reference threshold based on the first hit density acquisition includes: calculating the ratio between the total number of hit subtexts in Z hit text chapters and Z as the maximum hit density; acquiring the difference between the maximum hit density and the first hit density, and calculating the product between the difference and the set density control parameter as the density fluctuation value; the sum of the density fluctuation value and the first hit density as the density reference threshold.
17. A data matching apparatus, characterized by The device includes: An acquisition module is configured to acquire a text material associated with a video resource set and a target video resource, wherein the text content described by the text material matches the video content presented by the video resource set, and the target video resource is any video resource in the video resource set. A segmentation module is configured to perform video segmentation processing on the target video resource to obtain a plurality of sub-videos of the target video resource, and perform text segmentation processing on the text material to obtain a plurality of subtexts of the text material. A matching module is configured to perform content matching processing between the plurality of sub-videos and the plurality of subtexts to obtain a plurality of matching pairs, wherein each matching pair includes one sub-video and one subtext, and the sub-video and the subtext in the same matching pair have content matching. A determination module is configured to determine a text range in the text material that matches the target video resource according to the plurality of matching pairs.
18. A computer program product, characterised in that, The computer program product includes a computer program stored in a computer readable storage medium, the computer program is suitable for being read and executed by a processor, so that the computer device with the processor executes the method of any one of claims 1-16.
19. A computer device, comprising: The computer program product includes a computer program stored in a computer readable storage medium, the computer program is suitable for being read and executed by a processor, so that the computer device with the processor executes the method of any one of claims 1-16.
20. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by a processor to execute the steps of the method of any one of claims 1-16.