Video searching method and device, equipment and medium
By identifying the target entity name in the video search request, and using nodes and edges in the graphical data structure to determine the target node that matches the target entity name, the problem of low accuracy of video search in the prior art is solved, and high-accurate video search is achieved.
Patent Information
- Application Number
- CN202510227295.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
The accuracy of video search in the prior art is not high, especially when it is necessary to search for a specific plot segment or a plot segment related to a certain main line.
By obtaining video search requests in natural language form, identify the target entity name, and use nodes and edges in the graphical data structure to determine the target node that matches the target entity name. Then, based on the target vector of the video search request and the feature vector of the description text, the description text that satisfies the preset conditions is determined from the target set, and the target video clip matching the video search request is finally determined.
Improve the accuracy of video search results and accurately search target video clips that match the video search request.
Smart Images

Figure CN120067391A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computers, and in particular, to a video search method, apparatus, device, and medium. Background Art
[0002] With the continuous development of technology, there is a large amount of media information in the network, such as videos. Some videos are plot segments in movies or TV dramas. During the viewing and analysis of movies and TV dramas, users often need to search for specific plot segments. For example, a user may hope to find all the plot segments of a certain character in a specific plot, or need to search for plot segments related to a certain main line according to the plot development.
[0003] The usual search method is keyword search. For example, corresponding videos are searched by keyword matching, but generally the accuracy of the search results obtained by this method is not high. Summary of the Invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a video search method, apparatus, device, and medium to improve the accuracy of video search results.
[0005] In a first aspect, an embodiment of the present disclosure provides a video search method, the method including:
[0006] Obtain a video search request;
[0007] Identify at least one target entity name in the video search request;
[0008] According to the at least one target entity name and the names of each node in the graphic data structure, determine at least one target node corresponding to the at least one target entity name. The at least one target node respectively forms a target set with a corresponding at least one description text. Nodes in the graphic data structure represent entities in the content description information corresponding to video segments, and edges in the graphic data structure represent relationships between entities;
[0009] According to the target vector of the video search request and the feature vectors of each description text in the target set, determine one or more description texts that meet preset conditions from the target set;
[0010] According to the one or more description texts that meet preset conditions, determine one or more target video segments matching the video search request. The description text is generated according to the content description information of the target video segment.
[0011] In a second aspect, an embodiment of the present disclosure provides a video search apparatus, the apparatus including:
[0012] An acquisition module, configured to acquire a video search request;
[0013] An identification module, configured to identify at least one target entity name in the video search request;
[0014] A first determination module, configured to determine at least one target node corresponding to the at least one target entity name according to the at least one target entity name and the names of each node in the graphic data structure, where the at least one target node respectively forms a target set with the corresponding at least one description text, nodes in the graphic data structure represent entities in the content description information corresponding to video segments, and edges in the graphic data structure represent relationships between entities;
[0015] A second determination module, configured to determine one or more description texts that meet a preset condition from the target set according to the target vector of the video search request and the feature vectors of each description text in the target set;
[0016] A third determination module, configured to determine one or more target video segments that match the video search request according to the one or more description texts that meet the preset condition, where the description text is generated according to the content description information of the target video segment.
[0017] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:
[0018] A memory;
[0019] A processor; and
[0020] A computer program;
[0021] Wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the method as described in the first aspect.
[0022] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the method as described in the first aspect.
[0023] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, and the computer program implements the method as described in the first aspect when being executed by a processor.
[0024] The video search method, apparatus, device, and medium provided by the embodiments of the present disclosure obtain a video search request in the form of natural language and identify at least one target entity name in the video search request. Further, according to each target entity name, a node search is performed on the graphic data structure to obtain at least one target node that matches the at least one target entity name. Since each target node corresponds to one or more description texts, each description text corresponds to a feature vector, and each description text is generated according to the content description information of a video segment. Therefore, according to the target vector of the video search request and the feature vectors of each description text, one or more description texts that are closer to the video search request in the vector space can be determined. Since the one or more description texts can accurately express the requirements of the video search request, one or more target video segments that match the video search request can be accurately searched according to the one or more description texts, thereby improving the accuracy of the video search results. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 It is a schematic diagram of the application scenario provided by the embodiments of the present disclosure;
[0028] Figure 2 It is a schematic diagram of constructing a graphic data structure provided by the embodiments of the present disclosure;
[0029] Figure 3 It is a schematic diagram of the attribute information of a node provided by the embodiments of the present disclosure;
[0030] Figure 4 It is a flowchart of the video search method provided by the embodiments of the present disclosure;
[0031] Figure 5 It is a flowchart of the video search method provided by the embodiments of the present disclosure;
[0032] Figure 6 It is a flowchart of the video search method provided by the embodiments of the present disclosure;
[0033] Figure 7 It is a flowchart of the video search method provided by the embodiments of the present disclosure;
[0034] Figure 8 Flowchart of the video search method provided by an embodiment of the present disclosure;
[0035] Figure 9 Schematic structural diagram of the video search device provided by an embodiment of the present disclosure;
[0036] Figure 10 Schematic structural diagram of an embodiment of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0037] In order to more clearly understand the above objects, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.
[0038] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some embodiments of the present disclosure, rather than all embodiments.
[0039] An embodiment of the present disclosure provides a video search method, which can be executed by a video search device. The device can be implemented in a software and / or hardware manner and can be configured in a server or a server cluster. Specifically, the method is applicable to an Figure 1 application scenario as shown in the figure. The application scenario includes a user terminal 11 and a server 12. Among them, the user terminal 11 can be a mobile phone, a personal digital assistant, a tablet computer, a wearable device with a display screen, a desktop computer, a notebook computer, an all-in-one computer, a smart home device, etc. The user can input a video search request in the form of natural language on the user interface provided by the user terminal 11. The user terminal 11 sends the video search request to the server 12. The server 12 searches for one or more target video segments matching the video search request according to the method described in the embodiment of the present disclosure and feeds back the one or more target video segments to the user terminal 11, so that the user terminal 11 displays the one or more target video segments to the user. During the process of searching for the target video segments, the server 12 needs to search the graphic data structure. Therefore, the generation process of the graphic data structure is introduced first. As Figure 2 shown, the graphic data structure is generated according to the following steps:
[0040] S201. Cut the original video into multiple video segments.
[0041] For example, the server 12 locally stores a TV drama, or the TV drama is stored in other servers, and the server 12 obtains the TV drama from other servers. This TV drama can be used as the original video. In other embodiments, the original video is not limited to TV dramas, and can also be movies, variety shows, animations, etc. The embodiments of the present disclosure take TV dramas as an example. Specifically, this TV drama includes multiple episodes of videos, and each episode of video is recorded as a long video. The server 12 can split each episode of video into multiple video segments according to the scene changes in each episode of video, and each video segment is recorded as a short video. For example, the server 12 uses a scene detection (scenedetect) tool to detect the scene change time points in each episode of video, and splits the episode of video at the scene change time point. In addition, it can be understood that splitting each episode of video according to scene changes is only an illustrative splitting method, and the embodiments of the present disclosure do not limit the specific splitting method. For example, it can also be a fixed-length splitting method.
[0042] S202. Generate content description information corresponding to each of the multiple video segments.
[0043] For example, the server 12 splits the entire TV drama into 100 video segments. For each video segment, the server 12 extracts M frame images from the video segment, for example, extracts M frame images at equal intervals. Further, the server 12 inputs the M frame images and prompt content into a large model. The prompt content can be an instruction sent to the large model. For example, this instruction is used to instruct the large model to understand the M frame images and describe the content of the M frame images, so as to obtain the content description information corresponding to the video segment, that is, the large model outputs the content description information corresponding to the video segment. By a similar method, the content description information corresponding to other video segments is obtained. That is to say, each of the 100 video segments corresponds to a content description information. The content description information can be a text of 300 - 500 words, and the content description information can also be referred to as the video segment content description text. For example, the content description information corresponding to the x-th video segment among the 100 video segments is as follows:
[0044] Gao and An had dinner together in a small restaurant with simple decoration and a warm environment. In the restaurant, Gao wore a dark suit and An wore a brown leather jacket. Gao took the initiative to praise the good business, and then mentioned to An that his father invited them to eat pig's trotter noodles when he was a child, expressing his love for this food. Then, An changed the subject and mentioned that Zhang had not been found yet. Gao hinted that if Zhang was not there, he could clear his suspicion. An asked Gao if he was familiar with a person named Cheng and asked if he had any information about her. Gao seemed unconcerned, saying that An saw the problem with Cheng at a glance and tried to change the subject to something else. He asked An if Zhang was still in City A, and revealed that all intersections and highways had been blocked, and even waterways could not escape. During this period, the two had been eating noodles, and their conversation was slightly tense. When Gao mentioned that Chen hoped to invite him and Cheng to go sea fishing together on the weekend to ease the relationship, he further asked An what he would do if she was Cheng. An seemed a little impatient and asked not to beat around the bush. Gao continued to eat noodles and happily said that the noodles were delicious and asked for another bowl. In the second half of the story, when An stood up to leave, Gao still lowered his head to eat noodles and casually said, "Another bowl please." An stood up and confirmed with the person next to him that he had paid the bill before leaving, and Gao also confirmed it once. Throughout the process, Gao tried to paralyze and test An through constant conversation, while An remained vigilant and thought about the complex relationship between them under this calm surface, revealing the true intentions and purposes of both parties in this dinner.
[0045] S203. Generate attribute information of multiple nodes and attribute information of multiple edges according to the content description information respectively corresponding to the multiple video clips, wherein the attribute information of the nodes includes a node name, a node type, a node description set, and source information of each description text in the node description set; the edge is used to associate a first node and a second node, and the attribute information of the edge includes a first node name, a second node name, a relationship metric set between the first node and the second node, a relationship description set, and source information of each description text in the relationship description set.
[0046] In the embodiments of the present disclosure, nodes in the graphic data structure represent entities in the content description information corresponding to video segments, and edges in the graphic data structure represent relationships between entities. The edges between nodes represent the association relationships between nodes. Specifically, the relationship between every two different entities is recorded as a relationship, each entity corresponds to a node, and each relationship corresponds to an edge. Each entity corresponds to entity information, and each relationship corresponds to relationship information. Specifically, the entity information corresponding to any entity includes an entity name, an entity type, and an entity description. The relationship information corresponding to any relationship, such as the relationship between a first entity and a second entity, includes a first entity name, a second entity name, a relationship description, and a relationship metric. In addition, the same entity may appear in multiple different content description information, and the entity descriptions of the same entity generated according to different content description information may be different, that is, the same entity corresponds to multiple entity descriptions. Similarly, the same relationship may appear in multiple different content description information, and the relationship descriptions and relationship metrics of the same relationship generated according to different content description information may be different, that is, the same relationship corresponds to multiple relationship descriptions and multiple relationship metrics.
[0047] Furthermore, attribute information of the node corresponding to the entity is generated according to the entity information of the same entity. The attribute information of the node includes a node name, a node type, a set of node descriptions, and source information of each description text in the set of node descriptions. Optionally, the source information of the description text is the video segment identifier corresponding to the content description information used to generate the description text. Specifically, the entity name in the entity information can be used as the node name, the entity type can be used as the node type, and the multiple entity descriptions of the entity constitute the set of node descriptions. Since each entity description is text information, each entity description is recorded as a description text, that is, the set of node descriptions includes multiple description texts, and the source information of each description text can be the video segment identifier corresponding to the content description information used to generate the description text, or the number of the content description information. In addition, if the entity descriptions of the same entity generated according to different content description information are the same, or the same entity has only one entity description, the set of node descriptions may include one description text.
[0048] In addition, the attribute information of the edge corresponding to the relationship is generated according to the relationship information of the same relationship. For example, taking the relationship between the first entity and the second entity as an example, the first entity corresponds to the first node, the second entity corresponds to the second node, and the attribute information of the edge includes the first node name, the second node name, the relationship metric set between the first node and the second node, the relationship description set, and the source information of each description text in the relationship description set. Among them, the first entity name can be used as the first node name, the second entity name as the second node name, the multiple relationship metrics corresponding to the relationship constitute the relationship metric set between the first node and the second node, and the multiple relationship descriptions corresponding to the relationship constitute the relationship description set. Since each relationship description is text information, each relationship description is recorded as a description text, that is, the relationship description set includes multiple description texts. The source information of each description text can be the video segment identifier corresponding to the content description information for generating the description text, or the number of the content description information. In addition, if the relationship descriptions of the same relationship generated according to different content description information are the same, or there is only one relationship description for the same relationship, the relationship description set may include one description text.
[0049] Taking the content description information corresponding to the x-th video segment as described above as an example, the server 12 inputs the content description information corresponding to the x-th video segment and the prompt content into the large model. The prompt content is used to instruct the large model to extract entity information of types such as person, action, status, mission, event, object, organization, location, emotion, etc., and to instruct the large model to extract the relationship information between these types of entities. The output of the large model is as follows:
[0050] ("entity" **"Gao Moumou"** "person" **"Gao Moumou is wearing a dark suit and having dinner with An Moumou. During the conversation, he tries to lull and test An Moumou, appearing nonchalant."**) **{record_delimiter}** ("entity" **"An Moumou"** "person" **"An Moumou is wearing a brown leather jacket and having dinner with Gao Moumou, remaining vigilant and thinking about their complex relationship."**) **{record_delimiter}** ("entity" **"Zhang Moumou"** "person" **"Zhang Moumou is a person who has not been found yet. Mentioning him triggers a conversation between Gao Moumou and An Moumou."**) **{record_delimiter}** ("entity" **"Cheng Moumou"** "person" **"Cheng Moumou is a person mentioned by An Moumou. Gao Moumou tries to change the topic and imply her problems."**) **{record_delimiter}** ("entity" **"Chen Moumou"** "person" **"Chen Moumou hopes to invite Gao Moumou and Cheng Moumou to go sea fishing together on the weekend to ease the relationship."**) **{record_delimiter}** ("entity" **"Pig's Trotter Noodles"** "object" **"Pig's Trotter Noodles is a kind of food that Gao Moumou's father treated them to when he was a child. Gao Moumou expressed his love for this food."**) **{record_delimiter}** ("entity" **"Small Restaurant"** "location" **"The small restaurant is the place where Gao Moumou and An Moumou have dinner. It is simply decorated and has a warm environment."**) **{record_delimiter}** ("entity" **"City A"** "location" **"City A is the place where Gao Moumou asks if Zhang Moumou is still there."**) **{record_delimiter}** ("entity" **"Blockade"** "action" **"Gao Moumou reveals that all the intersections and highways have been blocked, and even the waterways cannot be escaped."**) **{record_delimiter}** ("entity" **"Sea Fishing"** "action" **"Chen Moumou hopes to ease the relationship with Gao Moumou and Cheng Moumou through sea fishing."**)\n**{record_delimiter}**\n(\"entity\"**\"nervous\"**\"emotion\"**\"Gao and An were slightly nervous during the conversation.\"**)\n**{record_delimiter}**\n(\"relationship\"**\"Gao\"**\"An\"**\"Gao tried to lull and test An through the conversation, while An remained vigilant and pondered their complex relationship.\"**8**)\n**{record_delimiter}**\n(\"relationship\"**\"Gao\"**\"Zhang\"**\"Gao implied that if Zhang was not around, he could clear himself of suspicion.\"**7**)\n**{record_delimiter}**\n(\"relationship\"**\"An\"**\"Zhang\"**\"An mentioned that Zhang had still not been found and asked Gao if he was familiar with Cheng.\"**6**)\n**{record_delimiter}**\n(\"relationship\"**\"Gao\"**\"Cheng\"**\"Gao tried to change the topic and hinted at Cheng's issues.\"**5**)\n**{record_delimiter}**\n(\"relationship\"**\"Gao\"**\"Chen\"**\"Gao mentioned that Chen hoped to invite him and Cheng to go sea fishing on the weekend to ease the relationship.\"**6**)\n**{record_delimiter}**\n(\"relationship\"**\"Gao\"**\"pork knuckle noodles\"**\"Gao expressed his love for pork knuckle noodles and mentioned the past event of his father treating them to pork knuckle noodles when he was a child.\"**7**)\n**{record_delimiter}**\n(\"relationship\"**\"Gao\"**\"small restaurant\"**\"Gao and An had dinner together at a small restaurant. The restaurant had simple decoration and a warm environment.\"**6**)\n**{record_delimiter}**\n(\"relationship\"**\"An\"**\"small restaurant\"**\"An and Gao had dinner together at a small restaurant. The restaurant had simple decoration and a warm environment.\"**6**)\n**{record_delimiter}**\n(\"relationship\"**\"Gao\"**\"City A\"**\"Gao asked if Zhang was still in City A."**5**)\n**{record_delimiter}**\n(\"relationship\"**\"Gao Moumou\"**\"blockade\"**\"Gao Moumou revealed that all intersections and highways have been blocked, and there is no escape even by waterway.\"**7**)\n**{record_delimiter}**\n(\"relationship\"**\"Chen Moumou\"**\"sea fishing\"**\"Chen Moumou hopes to ease the relationship with Gao Moumou and Cheng Moumou through sea fishing.\"**6**)\n**{completion_delimiter}**.
[0051] Among them, entity represents an entity, record_delimiter represents a record delimiter, relationship represents a relationship, and completion_delimiter represents an end delimiter.
[0052] Furthermore, the server 12 parses the output of the large model. For example, during the process of traversing the output of the large model, according to the record_delimiter, the output of the large model is divided into multiple records. For example, the first record is (\"entity\"**\"Gao Moumou\"**\"person\"**\"Gao Moumou is wearing a dark suit, having dinner with An Moumou, and trying to paralyze and test An Moumou during the conversation, appearing nonchalant.\"**). Furthermore, each record is divided into multiple strings according to **, and these multiple strings form a string list. If the first string in the string list is \"entity\", then the subsequent elements are the entity name, entity type, and entity description in sequence. If the first string in the string list is \"relationship\", then the subsequent elements are the first entity name, the second entity name, the relationship description, and the relationship metric in sequence.
[0053] For example, the multiple strings in the first record are \"entity\", \"Gao Moumou\", \"person\", \"Gao Moumou is wearing a dark suit, having dinner with An Moumou, and trying to paralyze and test An Moumou during the conversation, appearing nonchalant.\" \"Gao Moumou\" is the entity name, \"person\" is the entity type, and \"Gao Moumou is wearing a dark suit, having dinner with An Moumou, and trying to paralyze and test An Moumou during the conversation, appearing nonchalant.\" is the entity description.
[0054] For another example, ("relationship" **"Gao Moumou"** "An Moumou" **"Gao Moumou tried to paralyze and probe An Moumou through conversation, while An Moumou remained vigilant and thought about their complex relationship."** 8) is another record, where "Gao Moumou" is the first entity name, "An Moumou" is the second entity name, "Gao Moumou tried to paralyze and probe An Moumou through conversation, while An Moumou remained vigilant and thought about their complex relationship" is the relationship description, and 8 is the relationship measure. This relationship measure is used to characterize the closeness of the relationship between the first entity and the second entity. The larger the value of this relationship measure, the greater the closeness.
[0055] For example, the server 12 forms an entity information by combining the entity name, entity type, and entity description following each "entity", thus obtaining multiple entity information. Additionally, an entity relationship information is formed by combining the first entity name, second entity name, relationship description, and relationship measure following each "relationship", thus obtaining multiple relationship information.
[0056] Furthermore, since the leading actors or important characters in a TV drama will appear in multiple video clips, and even in every video clip, the same entity names will appear in the content description information corresponding to different video clips. In addition, as the plot develops, the fates of the same characters will change, and the relationships between characters will also change. Therefore, the entity descriptions of the same entity generated based on the content description information corresponding to different video clips may be different, and the relationship descriptions and relationship measures of the same relationship generated based on the content description information corresponding to different video clips may also be different.
[0057] Taking "Gao Moumou" as an example, "Gao Moumou" appears in the content description information corresponding to multiple different video clips. For example, it appears in the content description information corresponding to the xth, yth, and zth video clips. According to the content description information corresponding to the xth video clip, the first entity information containing "Gao Moumou" is obtained. The entity description in this first entity information is that "Gao Moumou is wearing a dark suit, having dinner with An Moumou, trying to paralyze and test An Moumou during the conversation, and appearing nonchalant". According to the content description information corresponding to the yth video clip, the second entity information containing "Gao Moumou" is obtained. The entity description in this second entity information is that "Gao Moumou is wearing a black suit and working in the office". According to the content description information corresponding to the zth video clip, the third entity information containing "Gao Moumou" is obtained. The entity description in this third entity information is that "Gao Moumou is sentenced to prison". By statistically analyzing this first entity information, second entity information, and third entity information, the attribute information of the node corresponding to "Gao Moumou" can be obtained. The attribute information of this node includes the node name, node type, node description set, and the source information of each description text in this node description set. The node name is "Gao Moumou", and the node type is "person". The entity descriptions in the first entity information, the second entity information, and the third entity information constitute the node description set. Since each entity description is text information, each entity description is recorded as a description text. That is, this node description set includes 3 description texts. Among them, the 1st description text is generated according to the content description information corresponding to the xth video clip. This content description information comes from the xth video clip. Therefore, the source information of the 1st description text is the identifier of the xth video clip. If the video clips after splitting the whole TV drama are numbered starting from 1, then the identifier of the xth video clip is x. In addition, in other embodiments, it is not limited to starting from 1 as long as each video clip corresponds to a unique identifier. Similarly, the source information of the 2nd description text is the identifier of the yth video clip, such as y. The source information of the 3rd description text is the identifier of the zth video clip, such as z. In addition, since each video clip corresponds to a content description information, each content description information can be recorded as an original document, and each original document can correspond to a unique number. Therefore, the source information of each description text can also be the number of the original document that generates this description text.
[0058] S204. Generate the graphic data structure according to the attribute information of the multiple nodes and the attribute information of the multiple edges.
[0059] Specifically, the server 12 generates a graph data structure based on the attribute information of the multiple nodes and the attribute information of the multiple edges as described above. For example, it generates a Graph Retrieval-augmented Generation (RAG) data structure. Specifically, the attribute information of the multiple nodes and the attribute information of the multiple edges as described above can be saved in a graphml file. Graphml is a graph data format based on the Extensible Markup Language (XML) and is used to represent graph structure data. Taking a certain node as an example, the storage form of the attribute information of this node in the graphml file is as Figure 3 shown. Among them, the node identifier can be the node name, d0 represents the node type, d1 represents the set of node descriptions, and d2 represents the source information of each description text in this set of node descriptions. In addition, the storage form of the attribute information of each edge in the graphml file can be referred to Figure 3 shown. For example, the id of the edge is the name of the first node and the name of the second node, d3 represents the set of relationship metrics, and d4 represents the set of relationship descriptions.
[0060] Optionally, the method further includes: converting each description text in the set of node descriptions and each description text in the set of relationship descriptions into low-dimensional dense vectors respectively, and storing the low-dimensional dense vectors. For example, the graph data structure includes multiple nodes and multiple edges. According to the id of each node, the attribute information of this node can be found, and according to the id of each edge, the attribute information of this edge can be found. Each node corresponds to a set of node descriptions, and this set of node descriptions includes one or more description texts. Each edge corresponds to a set of relationship descriptions, and this set of relationship descriptions includes one or more description texts. In this embodiment, each description text in the set of node descriptions and each description text in the set of relationship descriptions can also be respectively converted into a feature vector. For example, the server 12 uses the embedding technology to convert each description text into a feature vector respectively, and this feature vector is a low-dimensional dense vector. Thus, similar texts are closer in the feature space. For example, "cat" and "dog" are closer in the feature space. Further, the server 12 can store the feature vectors of each description text in a database and assign an index to each feature vector. Taking any description text as an example, the index of the feature vector of this description text is associated with the source information of this description text, for example, associated with the number of the original document corresponding to this description text. In some other embodiments, the index of the feature vector of this description text can also be the video segment identifier corresponding to the content description information that generates this description text.
[0061] In the embodiments of the present disclosure, "low-dimensional" and "dense" are usually used to describe the characteristics of vectors or matrices. Here, "low-dimensional" and "dense" vectors are relative to "high-dimensional" and "sparse". "Low-dimensional" means that the number of elements in the vector is relatively small. For example, a 3-dimensional vector is a low-dimensional vector because it has only 3 elements. In contrast, a vector with thousands or millions of dimensions is usually called a "high-dimensional" vector. In practical applications, low-dimensional vectors are usually used as feature vectors, which can effectively describe and capture the main characteristics of data. A "dense" vector means that most (or all) of the elements of the vector are non-zero values. This is the opposite of a "sparse" vector. A sparse vector is a vector in which most of the elements are zero. Taking a simple example, the vector [1, 2, 3] is a dense vector because all its elements are non-zero. In contrast, the vector [1, 0, 0, 0, 2] is a sparse vector because most of its elements are zero. Low-dimensional dense vectors are often used for the representation of embeddings. In machine learning and deep learning, high-dimensional sparse features are converted into low-dimensional dense forms through learning, aiming to improve the computational efficiency and performance of the model. In the embodiments of the present disclosure, by converting each description text into a low-dimensional dense vector respectively. Thus, similar texts are closer in the feature space. For example, "cat" and "dog" are closer in the feature space. This is beneficial for recall processing of the input retrieval term, facilitating retrieval and matching, and thus improving the retrieval accuracy.
[0062] The video search method provided by the embodiments of the present disclosure will be introduced below in combination with the above-mentioned graphic data structure and database. Figure 4 The flowchart of the video search method provided by the embodiments of the present disclosure. The video search method provided by the embodiments of the present disclosure will be introduced below in combination with the legend. As Figure 4 shown, the specific steps of the method are as follows:
[0063] S401. Obtain a video search request.
[0064] As Figure 2 shown, the user can input a video search request in the form of natural language on the user interface provided by the user terminal 11, and the user terminal 11 sends the video search request to the server 12, so that the server 12 obtains the video search request. For example, the video search request is "An Moumou went to find Gao Moumou to eat pig's feet noodles and was angry and left."
[0065] S402. Identify at least one target entity name in the video search request.
[0066] For example, when the server 12 receives the video search request from the user terminal 11, it identifies at least one target entity name in the video search request, and the identification result is as follows:
[0067] {"person": ["An Moumou", "Gao Moumou"],
[0068] "object": ["Pig's Trotter Noodles"],
[0069] "action": ["look for", "eat"],
[0070] "emotion": ["being angered and leaving"]
[0071] }
[0072] Among them, "person", "object", "action", and "emotion" are entity types respectively, and "An Moumou", "Gao Moumou", "Pig's Trotter Noodles", "look for", "eat", and "being angered and leaving" are target entity names respectively.
[0073] S403. Determine at least one target node corresponding to the at least one target entity name according to the at least one target entity name and the names of each node in the graphic data structure. The at least one target node respectively forms a target set with the corresponding at least one description text. The nodes in the graphic data structure represent entities in the content description information corresponding to video segments, and the edges in the graphic data structure represent relationships between entities.
[0074] In the graphic data structure, each node corresponds to a unique id, and the id of each node can be the node name. Server 12 can find the node name that matches the target entity name according to each target entity name and the names of each node in the graphic data structure, and use the node corresponding to the node name as the target node corresponding to the target entity name. For example, "An Moumou", "Gao Moumou", "Pig's Trotter Noodles", "look for", "eat", and "being angered and leaving" respectively correspond to a target node. Since each target node corresponds to a node description set, and the node description set includes one or more description texts, therefore, one or more description texts corresponding to each target node form a target set.
[0075] In addition, if the video search request includes relationships between different entities, then Server 12 can also parse and process these relationships to obtain the two entities associated with each relationship. Further, according to the names of the two entities, search for the corresponding edges in the graphic data structure, and add the description texts in the relationship description set corresponding to the edges to the target set as described above, so as to enrich the description texts in the target set.
[0076] Optionally, determining at least one target node corresponding to the at least one target entity name according to the at least one target entity name and the names of each node in the graphic data structure includes: determining a target node corresponding to each target entity name in the at least one target entity name and the names of each node in the graphic data structure, where the name of the target node matches the target entity name.
[0077] For example, after the server 12 parses at least one target entity name from the video search request, according to each target entity name, it performs a breadth-first search (BFS) and a depth-first search (DFS) in the graphic data structure to search for a node name that matches the target entity name from the graphic data structure, and takes the node corresponding to the node name as the target node corresponding to the target entity name.
[0078] S404. Determine one or more description texts that meet the preset conditions from the target set according to the target vector of the video search request and the feature vectors of each description text in the target set.
[0079] For example, the server 12 calculates the feature vector of the video search request, which is denoted as the target vector. Further, the server 12 determines one or more description texts that meet the preset conditions from the target set according to the target vector and the feature vectors of each description text in the target set as described above. The one or more description texts that meet the preset conditions are the description texts that are closer to the video search request in the vector space.
[0080] S405. Determine one or more target video segments that match the video search request according to the one or more description texts that meet the preset conditions, where the description text is generated according to the content description information of the target video segment.
[0081] For example, the server 12 determines one or more target video segments that match the video search request according to the one or more description texts that meet the preset conditions, and the description text is generated according to the content description information of the target video segment, that is, according to each description text that meets the preset conditions, a target video segment can be determined.
[0082] Optionally, determining one or more target video segments that match the video search request according to the one or more description texts that meet the preset conditions includes: determining one or more target video segments that match the video search request according to the video segment identifiers corresponding to the one or more description texts that meet the preset conditions.
[0083] As Figure 3 shown, since each description text corresponds to a video segment identifier and each description text is generated based on the content description information of a video segment, when the server 12 determines one or more description texts that meet the preset conditions from the target set, it extracts the video segment identifiers corresponding to each description text that meets the preset conditions from the graphical data structure, that is, extracts one or more video segment identifiers. Further, one or more target video segments are extracted according to the one or more video segment identifiers, each video segment identifier corresponds to a target video segment, and the one or more target video segments are determined as one or more target video segments that match the video search request.
[0084] In an embodiment of the present disclosure, a video search request in natural language form is obtained, and at least one target entity name in the video search request is recognized. Further, according to each target entity name, a node search is performed on the graphical data structure to obtain at least one target node that matches the at least one target entity name. Since each target node corresponds to one or more description texts, each description text corresponds to a feature vector, and each description text is generated based on the content description information of a video segment. Therefore, according to the target vector of the video search request and the feature vectors of each description text, one or more description texts that are closer to the video search request in the vector space can be determined. Since the one or more description texts can accurately express the requirements of the video search request, one or more target video segments that match the video search request can be accurately searched according to the one or more description texts, thereby improving the accuracy of the video search results.
[0085] Optionally, determining one or more description texts that meet the preset conditions from the target set according to the target vector of the video search request and the feature vectors of each description text in the target set includes the following steps as Figure 5 shown:
[0086] S501. Calculate the similarity between the target vector and the feature vectors according to the target vector of the video search request and the feature vectors of each description text in the target set.
[0087] For example, the server 12 calculates the feature vector of the video search request, and this feature vector is denoted as the target vector. Further, the server 12 calculates the similarity between the target vector and the feature vectors of each description text in the target set as described above. For example, the target set includes m description texts, and the server 12 calculates m similarities.
[0088] S502. Sort the multiple description texts in the target set in descending order of the similarity to obtain a first sorting result. The first n description texts in the first sorting result are the description texts that meet the preset conditions, where n is greater than or equal to 1.
[0089] For example, the server 12 sorts the m description texts in descending order of the m similarities, such that the description text with a greater similarity is ranked further forward, to obtain a first sorting result. The first n description texts in the first sorting result are the description texts that meet the preset conditions, where n is greater than or equal to 1 and n is less than m.
[0090] Optionally, after determining one or more target video segments that match the video search request according to the one or more description texts that meet the preset conditions, the method further includes the following steps as Figure 6 shown:
[0091] S601. Calculate the matching degree between the content description information corresponding to each of the multiple target video segments and the video search request.
[0092] For example, when the server 12 determines that there are multiple target video segments that match the video search request, it can also obtain the content description information corresponding to each of the multiple target video segments. Further, calculate the matching degree between the content description information corresponding to each target video segment and the video search request.
[0093] S602. Sort the multiple target video segments according to the matching degree to obtain a second sorting result.
[0094] For example, the server 12 uses rerank to calculate the matching score between the content description information corresponding to each target video segment and the video search request, and sorts the multiple target video segments according to the matching score to obtain a second sorting result. Specifically, the target video segment with a higher matching score is ranked further forward.
[0095] Optionally, after sorting the multiple target video segments to obtain a second sorting result, the method further includes the following several steps as Figure 7 shown:
[0096] S701. For each target video segment in the multiple target video segments, determine the number of description texts in the target set that correspond to the target video segment. The description texts corresponding to the target video segment are generated according to the content description information corresponding to the target video segment.
[0097] For example, the target set described above is a set composed of the description texts corresponding to the target nodes hit after performing a graph search based on the target entity name in the video search request. Each target video segment corresponds to a content description information, and multiple description texts can be generated according to this content description information, that is, each target video segment corresponds to multiple description texts. Further, the server 12 can count how many of the multiple description texts corresponding to each target video segment are in the target set, that is, determine the number of description texts corresponding to each target video segment in the target set. The larger the number, the more the target video segment matches the video search request.
[0098] S702. Adjust the second sorting result according to the number of description texts corresponding to each of the multiple target video segments in the target set, the matching degree, and the distribution of the multiple target video segments in the original video, to obtain a third sorting result.
[0099] For example, the original video can be segmented into 100 video segments through the segmentation process described above, and the multiple target video segments searched by the server 12 that match the video search request are a subset of these 100 video segments. Therefore, some of the target video segments in the multiple target video segments may originate from the same set of videos, and some target video segments may originate from different sets of videos. By counting the distribution of the multiple target video segments in the original video, that is, counting which target video segments in the multiple target video segments are distributed in the same set of videos in the original video, and which target video segments are separately distributed in a certain set of videos. Further, the server 12 can adjust the second sorting result described above according to the number of description texts corresponding to each target video segment in the target set, the matching score between the content description information corresponding to each target video segment and the video search request, and the distribution of the multiple target video segments in the original video, to obtain a third sorting result. That is to say, the second sorting result is the result of sorting the multiple target video segments according to one factor, that is, the matching score between the content description information corresponding to each target video segment and the video search request. The third sorting result is the result of comprehensively sorting the multiple target video segments by combining multiple factors, such as the number of description texts corresponding to each target video segment in the target set, the matching score between the content description information corresponding to each target video segment and the video search request, and the distribution of the multiple target video segments in the original video, etc. That is, the third sorting result can correct the second sorting result.
[0100] Optionally, after adjusting the second sorting result to obtain a third sorting result, the method further includes: sending the third sorting result to the user terminal, and the multiple target video segments are displayed on the interface of the user terminal according to the third sorting result.
[0101] For example, after the server 12 obtains the third sorting result, it can send the third sorting result to the user terminal 11, so that the user terminal 11 displays the multiple target video segments on the user interface according to the third sorting result.
[0102] The embodiments of the present disclosure propose a plot search framework based on GraphRAG. By extracting key elements (such as characters, locations, events, etc.) from the plot segments, that is, the content description information corresponding to the above-mentioned video segments, as nodes. According to the relationships between the nodes, edges in the graph structure are constructed to identify the associations between the nodes. In addition, it supports users to query through video search requests in natural language form. The system automatically analyzes the user's needs and performs searches on the graph structure. It provides a general search function, flexibly adjusts the search depth and breadth according to the user's query needs, and improves the flexibility and efficiency of the search. In addition, using the graph structure to represent key elements and the relationships between key elements, it realizes flexible query processing based on the structure and adapts to diverse and complex user needs. By combining breadth-first search and depth-first search, the search strategy is adaptively adjusted, improving the search efficiency. Through the accurate extraction of nodes and edges, it ensures that the search results have high relevance and accuracy. And it can handle various relationships and interactions between different characters in the video segments, providing users with a complete plot context and key information. In addition, the embodiments of the present disclosure adopt the GraphRAG data structure, overcoming the limitations of traditional RAG methods and very fitting the development context of movie and TV drama plots. Extract entities and relationships in the movie and TV drama from the content description information of the video segments generated by the large model, construct an efficient vector database, support users to query through natural language descriptions, reduce the cost of the large model, and improve the search effect.
[0103] Figure 8 It is a flowchart of the video search method provided by another embodiment of the present disclosure. As Figure 8 shown, it includes the following steps:
[0104] S801. Generate the content description information corresponding to the video segment.
[0105] For example, this step includes preprocessing the original video, that is, splitting the original video into multiple video segments. In addition, this step also includes generating the content description information corresponding to the video segment according to the large model. For example, M frames of images are extracted from the video segment, and the M frames of images and the prompt content are input into the large model. The prompt content is used to instruct the large model to understand the M frames of images and describe the content of the M frames of images, so as to obtain the content description information corresponding to the video segment.
[0106] S802. Generate a graph data structure.
[0107] For example, this step includes the extraction of nodes and edges. For example, entities are extracted from the content description information as nodes, and the relationships between entities are used as edges. Additionally, this step also includes parsing the output of the large model and constructing a GraphRAG data structure. The specific parsing process and construction process are as described above.
[0108] S803. Graph search and general search.
[0109] For example, this step includes the calculation and storage of feature vectors, assigning an index to each feature vector, and querying for target video segments according to a video search request.
[0110] S804. Search result processing and output.
[0111] For example, this step includes sorting multiple target video segments and sending the sorted multiple target video segments to a user terminal.
[0112] It can be understood that the implementation manners and principles of S801 - S804 are as specifically described in the above embodiments and will not be elaborated here.
[0113] The embodiments of the present disclosure use the structure of GraphRAG to represent entities and relationships, which can accurately capture multi - dimensional relationships in video segments and improve the accuracy and relevance of search results. Additionally, the embedding technology is used to convert the description texts corresponding to nodes and edges into feature vectors, and semantic vector matching is performed based on the target vector of the video search request and the feature vectors of the description texts, improving the relevance of search results. Furthermore, the precision of search results is further improved by performing rerank sorting on target video segments.
[0114] Figure 9 It is a schematic structural diagram of the video search device provided by the embodiments of the present disclosure. This video search device is configured in a server or a server cluster. The video search device provided by the embodiments of the present disclosure can execute the processing flow provided by the embodiments of the video search method, as Figure 9 shown, the video search device 90 includes:
[0115] An acquisition module 91, configured to acquire a video search request;
[0116] An identification module 92, configured to identify at least one target entity name in the video search request;
[0117] The first determination module 93 is configured to determine at least one target node corresponding to the at least one target entity name according to the at least one target entity name and the names of each node in the graphic data structure. The at least one target node respectively forms a target set with a corresponding at least one description text. The nodes in the graphic data structure represent entities in the content description information corresponding to video segments, and the edges in the graphic data structure represent the relationships between entities.
[0118] The second determination module 94 is configured to determine one or more description texts that meet preset conditions from the target set according to the target vector of the video search request and the feature vectors of each description text in the target set.
[0119] The third determination module 95 is configured to determine one or more target video segments that match the video search request according to the one or more description texts that meet the preset conditions. The description text is generated according to the content description information of the target video segment.
[0120] Optionally, the video search device 90 further includes:
[0121] The segmentation module 96 is configured to segment the original video into multiple video segments.
[0122] The generation module 97 is configured to generate the content description information corresponding to each of the multiple video segments; generate the attribute information of multiple nodes and the attribute information of multiple edges according to the content description information corresponding to each of the multiple video segments. The attribute information of the node includes the node name, node type, node description set, and the source information of each description text in the node description set. The edge is used to associate the first node and the second node, and the attribute information of the edge includes the first node name, the second node name, the relationship metric set between the first node and the second node, the relationship description set, and the source information of each description text in the relationship description set; generate the graphic data structure according to the attribute information of the multiple nodes and the attribute information of the multiple edges.
[0123] Optionally, the source information of the description text is the video segment identifier corresponding to the content description information used to generate the description text.
[0124] Optionally, when the first determination module 93 determines at least one target node corresponding to the at least one target entity name according to the at least one target entity name and the names of each node in the graphic data structure, it is specifically configured to: determine a target node corresponding to the target entity name according to each target entity name in the at least one target entity name and the names of each node in the graphic data structure, and the name of the target node matches the target entity name.
[0125] Optionally, when the second determination module 94 determines one or more description texts that meet the preset conditions from the target set according to the target vector of the video search request and the feature vectors of each description text in the target set, it is specifically configured to: calculate the similarity between the target vector and the feature vectors according to the target vector of the video search request and the feature vectors of each description text in the target set; sort the multiple description texts in the target set in descending order of the similarity to obtain a first sorting result, and the first n description texts in the first sorting result are the description texts that meet the preset conditions, where n is greater than or equal to 1.
[0126] Optionally, when the third determination module 95 determines one or more target video segments that match the video search request according to the one or more description texts that meet the preset conditions, it is specifically configured to: determine one or more target video segments that match the video search request according to the video segment identifiers corresponding to the one or more description texts that meet the preset conditions.
[0127] Optionally, the video search device 90 further includes: a calculation module 98 and a sorting module 99, where the calculation module 98 is configured to calculate the matching degree between the content description information corresponding to each of the multiple target video segments and the video search request; the sorting module 99 is configured to sort the multiple target video segments according to the matching degree to obtain a second sorting result.
[0128] Optionally, the video search device 90 further includes: a fourth determination module 910 and an adjustment module 911, where the fourth determination module 910 is configured to determine the number of description texts in the target set corresponding to each target video segment among the multiple target video segments, and the description texts corresponding to the target video segments are generated according to the content description information corresponding to the target video segments; the adjustment module 911 is configured to adjust the second sorting result according to the number of description texts in the target set corresponding to the multiple target video segments respectively, the matching degree, and the distribution of the multiple target video segments in the original video to obtain a third sorting result.
[0129] Optionally, the video search device 90 further includes: a sending module 912, configured to send the third sorting result to the user terminal, and the multiple target video segments are displayed on the interface of the user terminal according to the third sorting result.
[0130] Optionally, the video search device 90 further includes: a conversion module 913 and a storage module 914. The conversion module 913 is configured to convert each description text in the node description set and each description text in the relationship description set into low-dimensional dense vectors respectively, and the storage module 914 is configured to store the low-dimensional dense vectors.
[0131] Figure 9 The video search device in the illustrated embodiment can be used to execute the technical solutions of the above method embodiments. The implementation principles and technical effects are similar and will not be elaborated here.
[0132] The internal functions and structures of the video search device have been described above. The device can be implemented as an electronic device. Figure 10 This is a schematic structural diagram of an electronic device embodiment provided by an embodiment of the present disclosure. As Figure 10 shown, the electronic device includes a memory 1001 and a processor 1002.
[0133] The memory 1001 is used to store programs. In addition to the above programs, the memory 1001 can also be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application program or method for operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.
[0134] The memory 1001 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0135] The processor 1002 is coupled to the memory 1001 and executes the programs stored in the memory 1001 to execute the technical solutions of the above method embodiments.
[0136] Further, as Figure 10 shown, the electronic device may further include: other components such as a communication component 1003, a power supply component 1004, an audio component 1005, a display 1006, etc. Figure 10 Only some components are schematically shown and do not mean that the electronic device only includes Figure 10 the components shown.
[0137] The communication component 1003 is configured to facilitate communication between the electronic device and other devices in a wired or wireless manner. The electronic device can access a communication standard-based wireless network, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1003 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1003 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0138] The power supply component 1004 provides power for various components of the electronic device. The power supply component 1004 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device.
[0139] The audio component 1005 is configured to output and / or input audio signals. For example, the audio component 1005 includes a microphone (MIC). When the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory 1001 or transmitted via the communication component 1003. In some embodiments, the audio component 1005 further includes a speaker for outputting audio signals.
[0140] The display 1006 includes a screen, and the screen may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.
[0141] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the method described in the above embodiments.
[0142] An exemplary embodiment of the present disclosure further provides a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute to implement the method described in the above embodiments.
[0143] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0144] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but rather will conform to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A video search method, wherein: The method comprises: Get video search request; Identifying at least one target entity name in the video search request; According to the at least one target entity name and the names of each node in the graph data structure, at least one target node corresponding to the at least one target entity name is determined, wherein the at least one target node and the corresponding at least one description text constitute a target set, the nodes in the graph data structure represent entities in the content description information corresponding to the video clip, and the edges in the graph data structure represent the relationship between the entities; Determine one or more description texts satisfying a preset condition from the target set according to the target vector of the video search request and the feature vector of each description text in the target set; One or more target video segments matching the video search request are determined according to the one or more description texts satisfying the preset condition, wherein the description texts are generated according to content description information of the target video segments.
2. The method according to claim 1, characterized in that The graph data structure is generated according to the following steps: Split the original video into multiple video clips; Generating content description information corresponding to the multiple video clips respectively; Generate attribute information of multiple nodes and attribute information of multiple edges according to the content description information respectively corresponding to the multiple video clips, wherein the attribute information of the nodes includes a node name, a node type, a node description set, and source information of each description text in the node description set; the edge is used to associate a first node and a second node, and the attribute information of the edge includes a first node name, a second node name, a relationship metric set between the first node and the second node, a relationship description set, and source information of each description text in the relationship description set; The graph data structure is generated according to the attribute information of the plurality of nodes and the attribute information of the plurality of edges.
3. The method according to claim 2, wherein: The source information of the description text is a video segment identifier corresponding to the content description information used to generate the description text.
4. The method according to claim 1, wherein: Determining at least one target node corresponding to the at least one target entity name according to the at least one target entity name and the names of the nodes in the graph data structure includes: According to each target entity name in the at least one target entity name and the names of each node in the graph data structure, a target node corresponding to the target entity name is determined, and the name of the target node matches the target entity name.
5. The method according to claim 1, wherein: According to the target vector of the video search request and the feature vector of each description text in the target set, one or more description texts satisfying a preset condition are determined from the target set, including: Calculate the similarity between the target vector and the feature vector according to the target vector of the video search request and the feature vector of each description text in the target set; According to the order of the similarity from large to small, the multiple description texts in the target set are sorted to obtain a first sorting result, and the first n description texts in the first sorting result are description texts that meet the preset conditions, and n is greater than or equal to 1.
6. The method according to claim 3, wherein: Determining one or more target video segments matching the video search request according to the one or more description texts satisfying the preset condition includes: One or more target video segments matching the video search request are determined according to the video segment identifiers corresponding to the one or more description texts that meet the preset conditions.
7. The method according to claim 1, wherein: After determining one or more target video segments matching the video search request according to the one or more description texts satisfying the preset condition, the method further includes: Calculating the matching degree between the content description information respectively corresponding to the multiple target video clips and the video search request; The multiple target video clips are sorted according to the matching degree to obtain a second sorting result.
8. The method according to claim 7, wherein: After sorting the multiple target video clips to obtain a second sorting result, the method further includes: According to each target video segment in the multiple target video segments, determining the number of description texts corresponding to the target video segment in the target set, wherein the description text corresponding to the target video segment is generated according to content description information corresponding to the target video segment; The second sorting result is adjusted according to the number of description texts in the target set respectively corresponding to the multiple target video clips, the matching degree, and the distribution of the multiple target video clips in the original video to obtain a third sorting result.
9. The method according to claim 8, wherein: After adjusting the second sorting result to obtain a third sorting result, the method further includes: The third sorting result is sent to a user terminal, and the multiple target video clips are displayed on an interface of the user terminal according to the third sorting result.
10. The method according to claim 2, characterized in that The method further comprises: Each description text in the node description set and each description text in the relationship description set are respectively converted into low-dimensional dense vectors, and the low-dimensional dense vectors are stored.
11. A video search device, wherein: The device comprises: An acquisition module, used to acquire video search requests; An identification module, used to identify at least one target entity name in the video search request; A first determination module is used to determine at least one target node corresponding to the at least one target entity name according to the at least one target entity name and the names of each node in the graph data structure, wherein the at least one target node and the corresponding at least one description text constitute a target set, the nodes in the graph data structure represent entities in the content description information corresponding to the video clip, and the edges in the graph data structure represent the relationship between the entities; A second determination module, configured to determine one or more description texts satisfying a preset condition from the target set according to the target vector of the video search request and the feature vector of each description text in the target set; The third determination module is used to determine one or more target video segments matching the video search request according to the one or more description texts satisfying the preset conditions, wherein the description texts are generated according to the content description information of the target video segments.
12. An electronic device, wherein: include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.