Object evaluation method and device, electronic equipment and storage medium

The hierarchical structure data of audio and video data is generated through large models, and segmented outlines and mind maps are generated, which solves the problem that users find it difficult to know video knowledge points, realizes the efficiency and experience of users to intuitively obtain information, and ensures content accuracy through automated evaluation.

CN120105013APending Publication Date: 2025-06-06BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510220622.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When watching long-term teaching videos, it is difficult for users to intuitively understand the knowledge points contained in the video, especially when viewing incompletely, it is difficult to judge whether the video meets their learning needs.

Method used

Provide an object evaluation method. By obtaining the objects to be evaluated and reference objects, using a large model to generate hierarchical structure data based on audio and video data, and generating segmented outlines, summary and mind maps based on the data, users can intuitively understand the distribution of knowledge points of video through these contents. At the same time, the accuracy of generated content is automatically evaluated through comparison with real hierarchical structure data.

Benefits of technology

It realizes that users can intuitively understand the distribution of knowledge points in the video, improves the efficiency and user experience of users to acquire information, and ensures the accuracy of generated content through automated evaluation, which is conducive to the optimization of subsequent large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105013A_ABST
    Figure CN120105013A_ABST
Patent Text Reader

Abstract

The invention provides an object evaluation method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to the fields of large models, generative models and the like. According to the specific implementation scheme, a to-be-evaluated object and a reference object are obtained; the to-be-evaluated object is hierarchical structure data determined by a large model based on the content of the audio and video data, and the reference object is real hierarchical structure data of the audio and video data; determining a first node set, a second node set and a third node set; wherein the first node set comprises to-be-evaluated nodes meeting a first predetermined condition and a second predetermined condition, the second node set comprises to-be-evaluated nodes meeting the second predetermined condition, and the third node set comprises reference nodes meeting the second predetermined condition; and performing evaluation according to the first node set, the second node set and the third node set to obtain an evaluation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the fields of large models, generative models, etc. More specifically, the present disclosure provides an object evaluation method, device, electronic device, storage medium, and computer program product. Background Art

[0002] Some applications for playing videos can generate segment outlines, outline summaries, mind maps, and other content for teaching videos based on artificial intelligence models, so that users can intuitively understand the knowledge points contained in the video through the above content. Summary of the invention

[0003] The present disclosure provides an object processing method, an apparatus, an electronic device, a storage medium, and a computer program product.

[0004] According to one aspect of the present disclosure, an object evaluation method is provided, including: obtaining an object to be evaluated and a reference object; the object to be evaluated is hierarchical structure data determined by a large model based on the content of audio and video data, and the object to be evaluated includes multiple nodes to be evaluated; the reference object is real hierarchical structure data of audio and video data, and the reference object includes multiple reference nodes; for each node to be evaluated, determining whether the node to be evaluated satisfies a first predetermined condition based on the similarity between the node to be evaluated and at least one reference node; determining a first node set, a second node set, and a third node set; wherein the first node set includes nodes to be evaluated in the object to be evaluated that meet the first predetermined condition and a second predetermined condition, the second node set includes nodes to be evaluated in the object to be evaluated that meet the second predetermined condition, and the third node set includes reference nodes in the reference object that meet the second predetermined condition; and evaluating the object to be evaluated according to the first node set, the second node set, and the third node set to obtain an evaluation result.

[0005] According to another aspect of the present disclosure, an object evaluation device is provided, comprising: an acquisition module, a determination module, a node set determination module and an evaluation module. The acquisition module is used to acquire an object to be evaluated and a reference object; the object to be evaluated is hierarchical structure data determined by a large model based on the content of audio and video data, and the object to be evaluated includes multiple nodes to be evaluated; the reference object is the real hierarchical structure data of audio and video data, and the reference object includes multiple reference nodes. The determination module is used to determine, for each node to be evaluated, whether the node to be evaluated meets the first predetermined condition according to the similarity between the node to be evaluated and at least one reference node. The node set determination module is used to determine the first node set, the second node set and the third node set; wherein the first node set includes the nodes to be evaluated in the object to be evaluated that meet the first predetermined condition and the second predetermined condition, the second node set includes the nodes to be evaluated in the object to be evaluated that meet the second predetermined condition, and the third node set includes the reference nodes in the reference object that meet the second predetermined condition. The evaluation module is used to evaluate the object to be evaluated according to the first node set, the second node set and the third node set to obtain an evaluation result.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided by the present disclosure.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the method provided by the present disclosure.

[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the method provided by the present disclosure is implemented.

[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0011] Figure 1 is a schematic diagram of an application scenario of an object processing method, an evaluation method, and an apparatus according to an embodiment of the present disclosure;

[0012] Figure 2 is a schematic flow chart of an object processing method according to an embodiment of the present disclosure;

[0013] Figure 3 is a schematic diagram of a structure tree according to an embodiment of the present disclosure;

[0014] Figure 4A is a schematic diagram of labeling an object according to an embodiment of the present disclosure;

[0015] Figure 4B is an example diagram of a structure tree of an object according to an embodiment of the present disclosure;

[0016] Figure 4C is a schematic diagram of labeling another object according to an embodiment of the present disclosure;

[0017] Figure 4D is an example diagram of a structure tree of another object according to an embodiment of the present disclosure;

[0018] Figure 4E is an example diagram of a structure tree of another object according to an embodiment of the present disclosure;

[0019] Figure 5A is a schematic diagram of a segmented outline according to an embodiment of the present disclosure;

[0020] Figure 5B is a schematic diagram of another segment outline according to an embodiment of the present disclosure;

[0021] Figure 6 is a schematic diagram of a prompt information template according to an embodiment of the present disclosure;

[0022] Figure 7 is a schematic diagram of a segmented summary according to an embodiment of the present disclosure;

[0023] Figure 8 is a schematic diagram of a mind map according to an embodiment of the present disclosure;

[0024] Fig. 9 is a schematic flow chart of an evaluation method according to an embodiment of the present disclosure;

[0025] Fig.10 is a schematic structural block diagram of an object processing device according to an embodiment of the present disclosure;

[0026] Fig.11 is a schematic structural block diagram of an evaluation device according to an embodiment of the present disclosure; and

[0027] Fig.12 It is a structural block diagram of an electronic device used to implement the object processing method and the object evaluation method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0029] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0030] In the technical solution of the present disclosure, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0031] Users sometimes watch audio and video resources such as lectures on the Internet to learn knowledge points in related fields. However, some videos are long and explain many knowledge points. It is difficult for users to know whether the knowledge points in the video are what they need without watching the video in its entirety. Some embodiments of the present disclosure provide an object processing method, which can generate a segmented outline for audio and video data such as lectures. Users can intuitively know the distribution of knowledge points through the segmented outline, so that they can easily determine whether to learn the audio and video according to their own needs, or choose to learn some clips in the audio and video according to their own needs, thereby improving the efficiency of users in obtaining information and improving user experience.

[0032] In practical applications, key contents such as structure trees, segment outlines, outline summaries, mind maps, etc. can be generated for lecture videos and displayed to users, thereby explaining the knowledge points contained in the audio and video. In practical applications, the above key contents are obtained by processing audio and video data based on a large model, but it is uncertain whether the content generated by the large model is accurate, so it is necessary to evaluate the generated key contents. Some embodiments of the present disclosure provide an object evaluation method, which can automatically and accurately evaluate key contents such as structure trees, segment outlines, outline summaries, etc. generated based on audio and video data of types such as lectures. The evaluation cost is low, and the accuracy of the evaluation results can be ensured, which is conducive to the subsequent iterative optimization of the large model.

[0033] The technical solution provided by the present disclosure will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0034] Figure 1 It is a schematic diagram of an application scenario of the object processing method, evaluation method and device according to an embodiment of the present disclosure.

[0035] It should be noted that Figure 1What is shown is merely an example of a system architecture to which the embodiments of the present disclosure can be applied, in order to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0036] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0037] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Terminal devices 101, 102, 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptops, desktop computers, etc.

[0038] The server 105 may be a server that provides various services, such as a background management server (only an example) that provides support for websites browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as a structure tree, segment outline, outline summary, mind map, etc. generated based on the audio and video data selected by the user, or an evaluation result of the object to be evaluated generated based on the user request) to the terminal device.

[0039] It should be noted that the object processing method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the object processing device provided in the embodiment of the present disclosure can generally be set in the server 105. The object processing method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the object processing device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.

[0040] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0041] Figure 2 is a schematic flowchart of an object processing method according to an embodiment of the present disclosure.

[0042] like Figure 2 As shown, the object processing method 200 may include operations S210 to S240.

[0043] In operation S210, an initial text is determined according to at least one of an audio and an image included in an object to be processed, the initial text including a plurality of subtexts, each of the subtexts having a first timestamp.

[0044] For example, the object to be processed can be data of types such as audio data and video data. In practical applications, the object to be processed is audio and video data of lectures, such as audio and video data of explaining knowledge in the fields of advanced mathematics, photography, etc. The object to be processed can be processed by OCR (Optical Character Recognition), ASR (Automatic Speech Recognition), etc. to obtain the initial text. The initial text includes multiple sub-texts. For example, the text converted from each sentence in the audio and video is a sub-text, or the text in the image frame in the video is a sub-text. In addition, each sub-text has a first timestamp, which can represent the start time or end time of the sub-text in the audio and video data.

[0045] In operation S220, based on the large model, a structure tree is generated according to the multiple sub-texts and the first timestamps of the multiple sub-texts; the structure tree includes multiple nodes, the attributes of each node include the node name and the second timestamp, each node represents a fragment in the object, and the dependency relationship between the multiple nodes represents: the hierarchical relationship between the multiple contents described by the multiple fragments in the object.

[0046] For example, a large model for generating a structure tree can be pre-trained, and the large model can be a generative model. This embodiment does not limit the structure of the large model. After training the large model, multiple sub-texts carrying the first timestamp are input into the large model, and prompt information (prompt) can be used to guide the large model to determine multiple knowledge points from the multiple sub-texts, and use the knowledge points as nodes, and organize each node into a structure tree according to the hierarchical relationship of each knowledge point in the multiple sub-texts. In this way, the large model can output a structure tree, which is a tree structure used to represent the overall context of audio and video data.

[0047] For example, the object to be processed is a video, in which the video clip of the first chapter describes the study method for the postgraduate entrance examination, the video clip of the second chapter describes the trend of the score line, and the video clip of the third chapter describes the postgraduate entrance examination strategy 1, postgraduate entrance examination strategy 2, and postgraduate entrance examination strategy 3. The structure tree generated in this way may include node 1 corresponding to the video clip of the first chapter, node 2 corresponding to the video clip of the second chapter, and node 3 corresponding to the video clip of the third chapter, and node 3 includes three child nodes corresponding to postgraduate entrance examination strategy 1, postgraduate entrance examination strategy 2, and postgraduate entrance examination strategy 3. In addition, the second timestamp of each node may represent at least one of the start time and the end time of the video clip corresponding to the node.

[0048] In operation S230, a target node is determined from the structure tree according to the dependency relationship of each node in the structure tree and the attribute of each node in the structure tree.

[0049] For example, for a node in a structure tree, if there is no child node of the node in the structure tree, the node is a leaf node. For example, in the above structure tree, node 1, node 2 and three child nodes are leaf nodes. Since node 3 has child nodes, node 3 is not a leaf node. Leaf nodes can be used as target nodes.

[0050] For another example, nodes whose node names have a number of characters less than a predetermined number may be filtered out from the structure tree, and the filtered out nodes may be used as target nodes.

[0051] For another example, nodes whose duration of corresponding video segments is greater than a predetermined duration may be screened from the structure tree, and the screened nodes may be used as target nodes.

[0052] It is understandable that in the actual screening process, the above-mentioned multiple screening strategies can be used in combination.

[0053] In operation S240 , a segment outline is determined according to the node name of the target node and the second timestamp so as to display the segment outline, where the segment outline includes outline titles, and each outline title has an outline timestamp.

[0054] For example, the target node corresponds to the outline title in the segment outline one by one, the node name of the target node can be used as the outline title of the segment outline, and the second timestamp of the target node can be used as the outline timestamp of the outline title.

[0055] For example, the outline titles may be displayed in the order of the target nodes, and the outline timestamp may be displayed near the outline titles.

[0056] For another example, according to the outline timestamp, the progress bar of the audio and video data is divided into multiple intervals, each of which corresponds to an outline title. Then, for each of the multiple intervals, the outline title associated with the interval is displayed in the interval. This allows the user to more intuitively know the specific location of the explanation content corresponding to the outline title in the video, and also to more intuitively know the duration of the explanation content corresponding to the outline title.

[0057] The technical solution provided in this embodiment can generate a segmented outline for audio and video data of the lecture type. The user can intuitively know the distribution of knowledge points through the segmented outline, so that the user can conveniently determine whether to learn the audio and video according to their own needs, or determine to learn part of the audio and video according to their own needs, thereby improving the efficiency of users in obtaining information and improving user experience.

[0058] Next, the process of determining the initial text based on the object to be processed is described.

[0059] In one example, the object to be processed is audio data, and speech recognition can be performed on the audio to obtain speech text and the timestamp of the speech text, and the speech text is used as a subtext in the initial text, and the timestamp of the speech text is used as the first timestamp.

[0060] In another example, the object to be processed is video data. Since the video data includes audio data and image data, for the audio data, the voice text and the timestamp of the voice text can be obtained in a similar manner as described above. For the image data, optical character recognition can be performed on the image frame to obtain the image text and the timestamp of the image text. The image text can be used as a subtext in the initial text, and the timestamp of the image text can be used as the first timestamp.

[0061] After obtaining the initial text, the initial text can be processed by a large model to obtain a structure tree. The large model can be a general large model or a large model dedicated to generating a structure tree. In this case, the large model needs to be trained. Next, the process of training the large model is described.

[0062] In practical applications, video samples or audio samples may be annotated in advance to obtain training samples. The training samples are then used to train the large model. This embodiment does not limit the training process of the large model. The training samples may be constructed first, and then the large model may be trained using the training samples.

[0063] In the process of annotating video samples or audio samples, taking video as an example, the video samples can be divided into intervals first, and the video samples can be cut into segments, thereby dividing them into several video units with no semantic overlap. In the actual division process, the interval division can be carried out according to the explanation logic of the lecturer in the video, and each segment is a complete semantic unit. The interval division can also be carried out layer by layer from large to small until no clear main information can be extracted from the courseware information. Among them, the inability to extract clear main information refers to the summary information without obvious hierarchical structure from the courseware. Next, information can be extracted from each interval of the video. The extraction logic can be to extract semantic information that can represent the overall meaning of the entire interval from the structural information of the courseware.

[0064] In practical applications, the labeled structure tree 301 can be used Figure 3 The form shown.

[0065] like Figure 4A As shown, an object can include Figure 4A The information 401 shown, the structure tree 402 generated based on the object is as follows Figure 4B shown.

[0066] like Figure 4C As shown, an object can include Figure 4C The information 403 shown, the structure tree 404 generated based on the object is as follows Figure 4D shown. Figure 4D It mainly reflects the hierarchical relationship of the knowledge points involved in the object, omitting the first timestamp of each node. Figure 4D In the structure tree 404 shown, the first timestamp corresponding to each node is added.

[0067] It should be noted that the structure tree is a tree structure, and the structure tree can be divided into at least one level, and the path lengths from nodes at the same level to the root node are the same. For example, Figure 4D The safety protection requirements for general operations, the safety production requirements for construction operations, and green construction management are at the same level, and the nodes such as lifting and hoisting operations, diving operations, the classification of production safety accidents in port and waterway engineering, and green construction management content are at the same level. In addition, the structure tree can also be divided into multiple subtrees. For example, a subtree includes the following nodes: construction safety management, safety protection requirements for general operations, lifting and hoisting operations, and diving operations.

[0068] Please refer to Figure 4E , another example of structure tree 405 is as follows Figure 4E shown.

[0069] After obtaining the structure tree of the object, the target node can be determined from the structure tree according to the dependency relationship of each node in the structure tree and the attributes of each node in the structure tree, and then the segment outline can be determined using the target node.

[0070] In this embodiment, the nodes in the structure tree can be screened according to at least one of the node repeatability, node abnormality and node quantity of the structure tree to obtain a screened structure tree. For any node in the screened structure tree, if the screened structure tree lacks other nodes with the node as a parent node, the node is a leaf node. The target node is determined according to the leaf node in the screened structure tree.

[0071] In one example, the process of adding marks to the nodes in the structure tree can be pre-configured with a first condition, and the first condition includes at least one of the following: the number of repeated nodes is greater than or equal to the threshold of the number of repeated nodes, and the node repetition rate is greater than or equal to the repetition rate threshold. According to the hierarchy of the structure tree, each level can be determined as the target level in order from low to high, and each level is processed in turn, with the root node at the highest level. Taking the processing of a target level as an example, if the nodes at the target level in the structure tree meet the first condition, the nodes at the target level in the structure tree are marked respectively, and the target level of the structure tree can be the lowest level. After a number of rounds of processing of the structure tree, the lowest level in the remaining structure tree can continue to be used as the target level. In this embodiment, if the number of repeated nodes in the target layer is large or the repetition rate is high, for example, a layer of nodes includes 5 "diving operations", during the front-end display process, on the one hand, the front-end display area is limited, and on the other hand, multiple repeated segmented outlines cause confusion in the overall knowledge structure, affecting the user experience. Therefore, for the target layer with serious repetition, you can consider tracing back one layer. The implementation method of tracing back one layer includes: after adding labels to the nodes, deleting the nodes with labels from the structure tree, and then selecting the target node from the retained nodes, thereby alleviating the duplication of the segmented outline.

[0072] In another example, the nodes in the structure tree can be marked by pre-configuring a first condition, which can include: the number of nodes is greater than or equal to the single-layer node number threshold, and if the nodes in the target level in the structure tree meet the first condition, the nodes in the target level in the structure tree are marked respectively. In this embodiment, if the number of nodes in the target level is greater than or equal to the single-layer node number threshold, it represents the node angle of the target level. For example, the target level has 20 nodes, but in the front-end display process, the front-end display area is limited, and too many outline titles will cause most of the outline titles to be unable to display normally, lose the function of the segmented outline, and affect the user experience. Therefore, for the target level with too many nodes, it can be considered to trace back one layer upward, thereby reducing the number of outline titles in the segmented outline. In addition, it can also be traced back in subtree units. For example, the repetition rate of leaf nodes in different subtrees is different, and it can be traced back one layer from the target level of the subtree with the highest repetition rate, and the remaining subtrees can be temporarily not traced back.

[0073] In another example, the nodes in the structure tree can be marked by pre-configuring a second condition, the second condition including at least one of the following: the number of abnormal nodes is greater than or equal to a predetermined abnormal number, and the ratio of the number of abnormal nodes to the number of nodes at the target level of the subtree is greater than or equal to a predetermined abnormal ratio. For the subtree in the structure tree, in response to detecting that the nodes at the target level in the subtree meet the second condition, the nodes at the target level in the subtree can be marked, and the target level of the subtree can be the lowest level, or the lowest level in a subtree remaining after filtering the structure tree for some rounds. In this embodiment, if there are many abnormal nodes at the target level, this will affect the normal display of the front end and the user experience. Therefore, it can be considered to trace back one level upward, thereby reducing the number of outline titles in the segment outline.

[0074] For example, the playing time of the clip corresponding to the abnormal node is less than or equal to the predetermined time. Since the clip corresponding to the abnormal node is short, the content of the knowledge point is usually less, so the importance is relatively low. In addition, when displayed on the front end, the interval that can be used to display the node is positively correlated with the duration. A short duration will result in a small display interval, which is not conducive to fully displaying the node name of the node. Therefore, this embodiment traces back one level when the number of abnormal nodes at the target level is large or the proportion is large.

[0075] For example, the number of characters in the node name of the abnormal node is greater than the first number of characters. Since the number of characters in the node name of the abnormal node is large, the node name may not be fully displayed during the front-end display process, which is not conducive to the user knowing the content and structure of the knowledge point. Therefore, this embodiment traces back one level when the number of abnormal nodes at the target level is large or the proportion is large.

[0076] After adding the mark, the nodes in the structure tree can be filtered according to the mark to obtain the filtered structure tree.

[0077] In one example, filtering may be performed in the following manner: nodes with a mark may be deleted from the structure tree.

[0078] In another example, the screening can be performed in the following manner: in response to determining that a node is marked, the marked node is deleted from the structure tree to obtain an intermediate structure tree. Then, according to the node collapse parameter or the level shuffling parameter of the intermediate structure tree, it is determined whether to cancel the mark for the marked node, and then the nodes that are not canceled are deleted to obtain a filtered structure tree.

[0079] For example, the node folding parameter may include the folding rate of each node in the intermediate structure tree, or the node folding parameter may include the folding rate of the intermediate outline composed of the leaf nodes of the intermediate structure tree. For example, a node has a corresponding child node in the structure tree, but the child node is deleted due to reasons such as tracing back one layer, then the node is in a folded state. The folding rate refers to the ratio of the nodes in the folded state to the total nodes. For example, if the intermediate outline includes 10 outline headings, of which 5 outline headings are folded, the folding rate is 50%.

[0080] It should be noted that if the node collapse parameter is greater than the corresponding threshold, it means that more leaf nodes in the structure tree cannot be displayed in the segment outline of the front end. In this way, the granularity of the segment outline is large and cannot represent the knowledge points contained in the object in detail. Therefore, the node mark is unmarked to avoid the node being deleted.

[0081] For example, the level shuffling parameter may include the level shuffling rate of the intermediate structure tree, or the node folding parameter may include the level shuffling rate of the intermediate outline consisting of the leaf nodes of the intermediate structure tree. For example, the intermediate structure tree has 10 nodes in total, 6 of which are in the third layer, 2 nodes in the second layer, and 2 nodes in the first layer. Then, the ratio of the maximum number of nodes in the same level to the total number of nodes can be used as the level shuffling rate, and the level shuffling rate in the present example is 40%.

[0082] It should be noted that if the level shuffling parameter is greater than the corresponding threshold, it can mean that the outline titles of the segment outline are scattered at different levels in the structure tree, so the level of the segment outline determined in this way is relatively chaotic, which will affect the user's understanding of the hierarchical relationship of the knowledge points in the object. Therefore, the node mark is cancelled to avoid deleting the node.

[0083] Through the above method, the nodes in the structure tree can be filtered according to the tags to obtain the filtered structure tree.

[0084] In other embodiments, in the process of selecting target nodes from the structure tree, other constraints may be imposed on the nodes. For example, the total number of characters of each node in the structure tree is less than the maximum number of characters required by the object, the number of all knowledge points is less than the maximum number of predetermined knowledge points, and the number of words in a single segment outline is less than the word limit required by a single time interval. In practical applications, based on obtaining user authorization, the audio and video of lectures stored on a computer, mobile phone, network disk, etc. may be read, and the number of segment outlines of these audio and video of lectures may be counted, and the average number of segment outlines of each audio and video of lectures may be used as the maximum number of predetermined knowledge points.

[0085] In addition, when tracing back one layer, tracing back can be performed based on priority. For example, when the number of characters in the segment outline is greater than the number of characters required for a single time interval, tracing back can be prioritized. For example, if the number of layers of two subtrees is different, tracing back the subtree with more layers is prioritized. In addition, if the display rate of a single node at the front end is less than the predetermined display rate, tracing back one layer can be performed.

[0086] Please refer to Figure 5A and Figure 5B In the case where the word count of the segment outline does not exceed the word count limit, an example of a segment outline 501 is as follows: Figure 5A In the case where the word count of the segment outline exceeds the word count limit, an example of a segment outline 502 is as follows Figure 5B shown.

[0087] According to another embodiment of the present disclosure, when the total duration of the object is greater than a predetermined duration, the number of characters in the initial text is greater than or equal to the second number of characters, and the number of leaf nodes in the structure tree is greater than a predetermined number, an operation of determining the target node from the structure tree is performed, otherwise the operation of determining the target node from the structure tree is not performed. In this embodiment, for videos with a shorter duration or less explanation or fewer knowledge points explained, the time required for users to watch the video is shorter, and such videos can be filtered without being regenerated into a segment outline. This embodiment generates a segment outline only for videos with a longer duration, more explanations, and more knowledge points explained. This embodiment determines whether to generate a segment outline based on the length of the video and the amount of content, which can avoid processing the entire video, reduce the number of large model calls, and reduce the processing cost of audio and video data.

[0088] According to another embodiment of the present disclosure, after generating a segment outline, a reference hierarchical structure can be determined according to the segment outline, and then a segment summary can be generated based on the large model, according to the initial text and the reference hierarchical structure, wherein the segment summary includes a summary title and a summary content text, and then the summary title and the summary content text are displayed.

[0089] For example, the section outline can be used as a reference hierarchy structure, or the section outline can be used as a part of the reference hierarchy structure, and the child nodes of the section outline in the structure tree can be used as another part of the reference hierarchy structure.

[0090] For example, a prompt information template prompt can be pre-configured, which can include a subtitle name, a reference hierarchical structure and an initial text. The subtitle name can be omitted, and the reference hierarchical structure and the initial text can be combined with the prompt information template and then input into the big model, which generates a segmented summary.

[0091] For example, a general large model can be used to generate segment summaries. It is also possible to use a general large model to generate segment summary samples, and then use these segment summary samples to train a large model dedicated to generating segment summaries.

[0092] For example, a prompt information template 601 is as follows: Figure 6 As shown, Figure 6 The subtitle structure tree in is the reference hierarchy structure mentioned above. The segment summary 701 output by the large model is as follows Figure 7 As shown, Figure 7 The title is the summary title, and sumary is the summary content text.

[0093] According to another embodiment of the present disclosure, in this embodiment, the node omission parameter of the structure tree can be determined based on the difference between the key points included in the initial text and the nodes in the structure tree. In addition, a predetermined omission condition is also pre-configured, for example, the node omission parameter includes an omission rate, and the predetermined omission condition includes that the omission rate is greater than or equal to a predetermined ratio. For another example, the node omission parameter includes an omission number, and the predetermined omission condition includes that the omission number is greater than or equal to a predetermined number.

[0094] If the node missing parameter meets the predetermined missing condition, the segment outline can be determined as the reference hierarchical structure. In this embodiment, if the missing rate is high, the segment summary is directly generated based on the segment outline.

[0095] If the node missing parameter does not meet the predetermined missing condition, the reference hierarchical structure can be determined based on the segment outline and the reference child node with the segment outline as the parent node. For example, if the structure tree includes a child node with the segment outline as the parent node, the child node with the segment outline as the parent node in the structure tree is determined as the reference child node. If the structure tree does not include a child node with the segment outline as the parent node, the reference child node is generated based on the large model, and the timestamp of the reference child node can also be generated. Then, the nodes of the segment outline are expanded based on the reference child node, so that the segment outline covers more knowledge points, so that a more comprehensive segment summary of the hierarchical structure and key points can be generated.

[0096] According to another embodiment of the present disclosure, in the process of determining the node missing parameters of the structure tree, for a subtree whose level in the structure tree meets a predetermined level condition, an expanded node for expanding the node of the subtree can be generated based on at least a portion of the initial text associated with the subtree in the initial text, and then the node missing parameters can be determined based on the difference between the expanded node and the node in the subtree.

[0097] For example, the timestamp corresponding to a certain subtree is the 3rd minute to the 23rd minute, then can be associated with this subtree according to the subtext that is between the 3rd minute to the 23rd minute in the initial text, these subtexts are input to the large model, determine the key point corresponding to this fragment by the large model, and the key point that lacks in the subtree is used as the expansion node. Next, can calculate the node omission rate of subtree according to the number of nodes in the subtree before the quantity of expansion nodes and the expansion.If the structure tree comprises the subtree that a plurality of levels satisfy the predetermined level condition, the mean value of the node omission rate of each subtree can be used as the node omission rate of the structure tree.

[0098] This embodiment uses the initial text to determine the key points of the segment, and then determines the node omission situation of the subtree, so that the node omission rate can be accurately determined.

[0099] According to another embodiment of the present disclosure, the predetermined hierarchical condition includes: the number of layers of the subtree is less than or equal to the predetermined number of layers, and the predetermined number of layers may be 1 layer. For example, a subtree has only one layer of nodes. This means that in the process of dividing the audio and video data, the content extracted from the video clip corresponding to the node lacks levels and there is a possibility of omission. Therefore, the subtree with fewer layers is expanded.

[0100] According to another embodiment of the present disclosure, after the structure tree is generated, a mind map may be generated according to the nodes in the structure tree and the dependency relationships between the nodes, and then the mind map may be displayed through the front end.

[0101] For example, a mind map can be generated based on a large model or based on a pre-configured processing logic. The dependency relationship between nodes in the mind map is consistent with the dependency relationship between nodes in the structure tree. In one example, the nodes in the mind map correspond to the nodes in the structure tree one by one. In another example, there are differences between the nodes in the mind map and the nodes in the structure tree. For example, the repeated nodes in the structure tree can be merged and then the mind map can be generated. For example, a schematic diagram of a mind map 801 is shown in FIG. Figure 8 shown.

[0102] Fig. 9 is a schematic flow chart of an evaluation method according to an embodiment of the present disclosure.

[0103] like Fig. 9As shown, the object evaluation method 900 may include operations S910 to S940.

[0104] In operation S910, an object to be evaluated and a reference object are obtained; the object to be evaluated is hierarchical structure data determined by the large model based on the content of the audio and video data, and the object to be evaluated includes multiple nodes to be evaluated; the reference object is real hierarchical structure data of the audio and video data, and the reference object includes multiple reference nodes.

[0105] For example, the audio and video data may be audio and video data of a lecture, such as audio and video data explaining knowledge in the fields of advanced data, photography, etc. For example, the node to be evaluated may represent a segment in the audio and video data, and the reference node may represent a real segment in the audio and video data.

[0106] For example, the object to be evaluated may include a structure tree, which is a tree structure used to represent the overall context of audio and video data, and the node to be evaluated may be a node in the structure tree. In the actual processing process, the initial text may be determined based on the image, voice, etc. of the audio and video data, and then the initial text of the audio and video data may be processed based on the large model to obtain the structure tree, and the initial text may be a text describing the content of the audio and video data.

[0107] For example, the object to be evaluated may include a segment outline, which may be the same as the structure tree, or some or all nodes in the structure tree may be used as target nodes, and then the target nodes may be used as the outline title of the segment outline, so that the segment outline includes multiple outline titles. The node to be evaluated may be the outline title of the segment outline.

[0108] For example, the object to be evaluated may include a segment summary, and the segment summary includes multiple summary nodes, each of which may include a summary title and summary content text. The node to be evaluated may be a summary node. In the actual processing process, a reference hierarchical structure may be obtained according to a structure tree or a segment outline, and then the reference hierarchical structure may be processed based on the large model to obtain a segment summary.

[0109] For example, the reference object may be obtained by pre-marking the audio and video data or by other means, and the type of the reference object is consistent with the type of the object to be evaluated. For example, if the object to be evaluated is a structure tree, the reference structure is a pre-marked reference structure tree, and the reference node is a node in the reference structure tree. If the object to be evaluated is a segment outline, the reference structure is a reference segment outline, and the reference node is an outline title in the reference segment outline. If the object to be evaluated is a segment summary, the reference structure is a reference segment summary, and the reference node is a summary node in the reference segment summary.

[0110] In operation S920, for each node to be evaluated, it is determined whether the node to be evaluated satisfies a first predetermined condition according to a similarity between the node to be evaluated and at least one reference node.

[0111] In one example, the nodes to be evaluated and the reference nodes whose levels and orders match can be determined according to the level and order of the nodes to be evaluated in the object to be evaluated, and the level and order of the reference nodes in the reference object. Taking the object to be evaluated as a structure tree and the reference object as a reference structure tree as an example, the nodes to be evaluated and the reference nodes that are at the same level in the structure tree and the reference structure tree and are in the same order in the level can be used as the matched nodes to be evaluated and the reference nodes. Then the similarity between the matched nodes to be evaluated and the reference nodes can be calculated.

[0112] For example, when the similarity is greater than or equal to the similarity threshold, it can be determined that the node to be evaluated meets the first predetermined condition, otherwise it is determined that the node to be evaluated does not meet the first predetermined condition. It can be seen that satisfying the first predetermined condition can mean that: the node to be evaluated obtained based on the large model is relatively consistent with the reference node, and it can be considered that the reasoning of the large model on the node to be evaluated is credible, otherwise it is considered that the reasoning of the large model on the node to be evaluated is uncredible.

[0113] In another example, the similarity between the node to be evaluated and each reference node in the reference structure tree can be determined, and the reference node with the greatest similarity is matched with the node to be evaluated. If the greatest similarity is greater than or equal to a similarity threshold, it can be determined that the node to be evaluated meets the first predetermined condition; otherwise, it is determined that the node to be evaluated does not meet the first predetermined condition.

[0114] In operation S930, a first node set, a second node set, and a third node set are determined; wherein the first node set includes nodes to be evaluated in the object to be evaluated that satisfy a first predetermined condition and a second predetermined condition, the second node set includes nodes to be evaluated in the object to be evaluated that satisfy the second predetermined condition, and the third node set includes reference nodes in the reference object that satisfy the second predetermined condition.

[0115] For example, the object to be evaluated may have at least one hierarchical structure, for example, the object to be evaluated is a tree structure, and the second predetermined condition indicates a condition that the position of the node to be evaluated in the object to be evaluated needs to satisfy.

[0116] For example, the object to be evaluated is screened for nodes according to the first condition and the second condition, and then the screened nodes are added to the first node set. The object to be evaluated and the reference object can also be screened for nodes according to the second predetermined condition to obtain the second node set and the third node set respectively.

[0117] In operation S940, the object to be evaluated is evaluated according to the first node set, the second node set, and the third node set to obtain an evaluation result.

[0118] For example, the evaluation result may be determined according to the number of nodes in the first node set, the number of nodes in the second node set, and the number of nodes in the third node set.

[0119] This embodiment automatically and accurately evaluates the structure tree, segment outline, outline summary, and other evaluation objects generated based on audio and video data such as lectures. The evaluation cost is low and the accuracy of the evaluation results can be ensured, which is conducive to the subsequent optimization of the large model.

[0120] According to another embodiment of the present disclosure, the object to be evaluated includes a structure tree or a segment outline for audio and video data, and whether the node to be evaluated meets the first predetermined condition can be determined in the following manner: whether the node to be evaluated meets the first predetermined condition can be determined based on the similarity between the node name of the matched node to be evaluated and the node name of the reference node. For example, the node to be evaluated in the structure tree has a node name. For example, the node name of a node to be evaluated is "Study Methods for Postgraduate Entrance Examination in Physics", and the reference node that matches the node to be evaluated is "Study Methods for Postgraduate Entrance Examination in Physics". If the similarity between the two is high, then the node to be evaluated meets the first predetermined condition. This embodiment calculates the similarity through the text of the node name to determine whether the node to be evaluated meets the first predetermined condition, so that it can accurately determine whether there is an error in the node to be evaluated obtained based on large model reasoning.

[0121] According to another embodiment of the present disclosure, the object to be evaluated includes a segmented summary of audio and video data, and whether the node to be evaluated meets the first predetermined condition can be determined in the following manner: whether the node to be evaluated meets the first predetermined condition can be determined based on the similarity between the matched summary content text and the partial sub-text.

[0122] In this embodiment, some subtexts can be screened by timestamps. For example, the initial text can be determined according to the image, voice, etc. of the audio and video data. The initial text includes multiple subtexts, and each subtext has a first timestamp. In addition, the segmented summary generated based on the initial text also has a timestamp. For example, a certain summary node in the segmented summary corresponds to the 3rd minute to the 23rd minute in the audio and video data. The 3rd minute to the 23rd minute part of the subtext corresponding to the audio and video data in the initial text can be screened out, and then the similarity between the part of the subtext and the summary content text of the summary node is calculated. If the similarity is high, it is considered that the summary content text is a summary summary made for the part of the subtext, so the node to be evaluated meets the first predetermined condition. Otherwise, it is considered that the content summarized by the summary content text is greatly different from the part of the subtext, so the node to be evaluated does not meet the first predetermined condition. In this embodiment, the similarity is calculated by the summary content text of the summary node, so as to determine whether the segmented summary meets the first predetermined condition, so that it can accurately determine whether there is an error in the node to be evaluated obtained based on the large model reasoning.

[0123] According to another embodiment of the present disclosure, the second predetermined condition may be at least one of a point dimension sub-condition, a line dimension sub-condition, and a subtree dimension sub-condition. The point dimension sub-condition refers to that the node is a leaf node. For the structure tree, segment outline, and segment summary waiting to be evaluated, for any node to be evaluated, if there are no other nodes with the node as the parent node in the object to be evaluated, then the node to be evaluated is a leaf node. The line dimension sub-condition refers to that the node is a node in the target node line. The node line includes a path from the root node to the leaf node. The target node line can be any node line. The subtree dimension sub-condition refers to that the node is a node in the target subtree. The target subtree includes multiple paths from the root node to the leaf node. The target subtree can be any subtree.

[0124] It should be noted that in the process of determining the first node set, the second node set and the third node set, the sub-conditions in the second condition used are consistent. For example, if the point dimension sub-condition is used to filter the nodes to obtain the first node set, then the point dimension sub-condition needs to be used to filter the nodes to obtain the second node set and the third node set, instead of using the line dimension sub-condition or the subtree dimension sub-condition to filter the nodes to obtain the second node set and the third node set. Similarly, the first node set, the second node set and the third node set are all obtained by filtering based on the line dimension sub-condition, or the first node set, the second node set and the third node set are all obtained by filtering based on the subtree dimension sub-condition.

[0125] According to another embodiment of the present disclosure, in the actual evaluation process, the accuracy sub-result of the object to be evaluated can be determined according to the number of nodes in the first node set and the number of nodes in the second node set, and the comprehensiveness sub-result of the object to be evaluated can also be determined according to the number of nodes in the first node set and the number of nodes in the third node set. Then, the evaluation result is determined according to the accuracy sub-result and the comprehensiveness sub-result.

[0126] Next, the evaluation process of the structure tree is explained. The structure tree can be evaluated from the point dimension, line dimension, subtree dimension, and time dimension respectively.

[0127] In one example, for the point dimension evaluation process of the structure tree, the first node set of the point dimension, the second node set of the point dimension, and the third node set of the point dimension can be determined first, wherein the first node set of the point dimension includes the leaf nodes that meet the first predetermined condition in the structure tree to be evaluated, the second node set of the point dimension includes all the leaf nodes of the structure tree to be evaluated, and the third node set of the point dimension includes all the leaf nodes of the reference structure tree. The accuracy sub-result of the structure tree in the point dimension can be determined based on the ratio between the summary points of the first node set of the point dimension and the summary points of the second node set of the point dimension. The comprehensiveness sub-result of the structure tree in the point dimension can be determined based on the ratio between the summary points of the first node set of the point dimension and the summary points of the third node set of the point dimension.

[0128] In another example, for the line dimension evaluation process of the structure tree, a path from the root node to the leaf node is a node line, and the node line is the line dimension. It can be seen that the structure tree can include multiple node lines. The evaluation process is described by taking a node line as an example. The first node set of the line dimension, the second node set of the line dimension, and the third node set of the line dimension can be determined first, wherein the first node set of the line dimension includes the leaf nodes that meet the first predetermined condition in a node line in the structure tree to be evaluated, the second node set of the line dimension includes all the leaf nodes in a node line in the structure tree to be evaluated, and the third node set of the line dimension includes all the leaf nodes in the matching node line in the reference structure tree. The accuracy sub-result of the node line can be determined according to the ratio between the total number of nodes of the first node set of the line dimension and the total number of nodes of the second node set of the line dimension. The comprehensiveness sub-result of the node line can be determined according to the ratio between the total number of nodes of the first node set of the line dimension and the total number of nodes of the third node set of the line dimension. When the structure tree includes multiple node lines, average values ​​of the accuracy sub-results and comprehensiveness sub-results of the multiple node lines may be determined as the accuracy sub-result and comprehensiveness sub-result of the line dimension of the structure tree.

[0129] In another example, for the subtree dimension evaluation process of the structure tree, multiple paths from the root node to multiple leaf nodes are a subtree. It can be seen that the structure tree can include multiple subtrees. The evaluation process is described by taking a subtree as an example. The first node set of the subtree dimension, the second node subtree set of the subtree dimension, and the third node subtree set of the subtree dimension can be determined first, wherein the first node set of the subtree dimension includes nodes that meet the first predetermined condition in a subtree in the structure tree to be evaluated, the second node subtree set of the subtree dimension includes all nodes in a subtree in the structure tree to be evaluated, and the third node subtree set of the subtree dimension includes all nodes in the matching subtree in the reference structure tree. The accuracy sub-result of the subtree can be determined based on the ratio between the summary points of the first node set of the subtree dimension and the summary points of the second node subtree set of the subtree dimension. The comprehensiveness sub-result of the subtree can be determined based on the ratio between the summary points of the first node set of the subtree dimension and the summary points of the third node subtree set of the subtree dimension. When the structure tree includes multiple subtrees, the average values ​​of the accuracy subresults and comprehensiveness subresults of the multiple subtrees may be determined to determine the accuracy subresults and comprehensiveness subresults of the structure tree in the subtree dimension.

[0130] In another example, for the evaluation process of the time dimension, the time evaluation result of the structure tree can be determined according to the difference between the timestamp of each node in the structure tree and the timestamp of each node in the reference structure tree. For example, the matched node to be evaluated and the reference node constitute a node pair, and the time difference between the timestamp of the node to be evaluated and the timestamp of the reference node can be calculated to obtain the time difference of the node pair. The structure tree to be evaluated and the reference structure tree can form multiple node pairs, and the average value of the time differences of the multiple node pairs can be calculated to evaluate the timestamp accuracy sub-result of the structure tree.

[0131] The evaluation process of the structure tree is described above from the point dimension, line dimension, subtree dimension and time dimension. In practical applications, at least one of the above dimensions can be selected to implement the evaluation process of the structure tree.

[0132] Next, the evaluation process of the segmented outline is explained. The segmented outline can be evaluated from the point dimension and the time dimension respectively.

[0133] In one example, for the point dimension evaluation process of the segmented outline, the first node set of the point dimension, the second node set of the point dimension, and the third node set of the point dimension can be first determined, wherein the first node set of the point dimension includes the leaf nodes of the segmented outline to be evaluated that meet the first predetermined condition, the second node set of the point dimension includes all the leaf nodes of the segmented outline to be evaluated, and the third node set of the point dimension includes all the leaf nodes of the reference segmented outline. The accuracy sub-result of the segmented outline in the point dimension can be determined based on the ratio between the summary points of the first node set of the point dimension and the summary points of the second node set of the point dimension. The comprehensiveness sub-result of the segmented outline in the point dimension can be determined based on the ratio between the summary points of the first node set of the point dimension and the summary points of the third node set of the point dimension.

[0134] It should be noted that, in some embodiments, the outline titles in the segmented outline are leaf nodes of the filtered structure tree, so that the segmented outline only includes leaf nodes and no other nodes. Therefore, the number of leaf nodes in the segmented outline is the same as the number of nodes (i.e., outline titles) in the segmented outline.

[0135] In other embodiments, the segment outline may also be evaluated from the line dimension and the subtree dimension. The evaluation process of the segment outline is similar to the evaluation process of the structure tree, and will not be described in detail in this embodiment.

[0136] In another example, for the evaluation process of the time dimension, the time evaluation result of the segment outline can be determined based on the difference between the timestamps of each node in the segment outline and the timestamps of the nodes of the reference segment outline. For example, the matching node to be evaluated and the reference node constitute a node pair, and the time difference between the timestamp of the node to be evaluated and the timestamp of the reference node can be calculated to obtain the time difference of the node pair. The structure tree to be evaluated and the reference structure tree can form multiple node pairs, and the average value of the time differences of multiple node pairs can be calculated to evaluate the time stamp accuracy of the outline.

[0137] In another example, for the evaluation process of the folding rate, the folding rate of the segmented outline can be determined based on the number of outline titles in the segmented outline that are in a folded state and the total number of outline titles in the segmented outline. For example, if a node has a corresponding child node in the structure tree, but there is no corresponding child node in the segmented outline, then the node (the node in the segmented outline is the outline title) is in a folded state. The folding rate refers to the ratio of the nodes in the folded state to the total nodes. For example, the segmented outline includes 10 outline titles, of which 5 outline titles are in a folded state, and the folding rate is 50%. It should be noted that if the folding rate of the segmented outline is high, it means that more leaf nodes in the structure tree cannot be displayed in the segmented outline of the front end, so the granularity of the segmented outline is large and it is impossible to represent the knowledge points contained in the audio and video data in detail.

[0138] In another example, for the character display rate evaluation process, the character display rate of the segmented outline can be determined based on the total number of characters in the outline title in the segmented outline and the number of characters in the outline title in the segmented outline displayed on the front end. For example, the front end needs to display the segmented outline. For a single outline title in the segmented outline, if the outline title includes 10 characters, the front end displays the first 5 characters, and the remaining 5 characters cannot be displayed normally due to display space limitations. Therefore, the character display rate of the outline title is 50%. If the segmented outline includes multiple outline titles, the average value of the character display rates of the multiple outline titles can be used as the overall character display rate of the segmented outline.

[0139] In another example, for the evaluation process of the level disruption rate, the level disruption rate of the segmentation outline can be determined according to the maximum number of the outline titles at the same level in the segmentation outline, and the total number of the outline titles in the segmentation outline according to the maximum number of the segmentation outlines at the same level, and the total number of the segmentation outlines. For example, the segmentation outline has a total of 10 nodes, wherein 6 nodes are in the 3rd layer of the structure tree, 2 nodes are in the 2nd layer of the structure tree, and 2 nodes are in the 1st layer of the structure tree, then the ratio of the maximum number of nodes at the same level to the total number of nodes can be used as the level disruption rate, such as the level disruption rate in this example is 40%. It should be noted that if the level disruption rate is higher, it can be represented that the nodes in the segmentation outline are more dispersedly distributed in the different levels of the structure tree, and the level of the segmentation outline determined in this way is more chaotic, which will affect the user's understanding of the hierarchical relationship of the knowledge point in the audio and video data.

[0140] The above describes the evaluation process of the segmented outline from the perspectives of point dimension, time dimension, folding rate, character display rate, and level shuffle rate. In practical applications, at least one of the above dimensions can be selected to implement the evaluation process of the segmented outline.

[0141] Next, the evaluation process of the segment summary is explained.

[0142] In one example, for the point dimension evaluation process of the segmented summary, the first node set, the second node set and the third node set can be determined first, wherein the first node set includes the summary nodes that meet the first predetermined condition in the segmented summary to be evaluated, the second node set includes all the summary nodes of the segmented summary to be evaluated, and the third node set includes all the summary nodes of the reference segmented summary. The accuracy sub-result of the segmented summary can be determined based on the ratio between the summary points of the first node set and the summary points of the second node set. The comprehensiveness sub-result of the segmented summary can be determined based on the ratio between the summary points of the first node set and the summary points of the third node set.

[0143] The above explains the evaluation process of the segment summary.

[0144] Fig.10 is a schematic structural block diagram of an object processing device according to an embodiment of the present disclosure.

[0145] like Fig.10 As shown, the object processing device 1000 may include an initial text determination module 1010 , a structure tree generation module 1020 , a target node determination module 1030 , and an outline determination module 1040 .

[0146] The initial text determination module 1010 is used to determine the initial text according to at least one of the audio and the image contained in the object to be processed, where the initial text includes a plurality of subtexts, each of which has a first timestamp.

[0147] The structure tree generation module 1020 is used to generate a structure tree based on the large model according to multiple sub-texts and the first timestamps of each of the multiple sub-texts; the structure tree includes multiple nodes, the attributes of each node include a node name and a second timestamp, each node represents a fragment in the object, and the dependency relationship between the multiple nodes represents: the hierarchical relationship between the multiple contents described by the multiple fragments in the object.

[0148] The target node determination module 1030 is used to determine the target node from the structure tree according to the dependency relationship of each node in the structure tree and the attribute of each node in the structure tree.

[0149] The outline determination module 1040 is used to determine the segment outline according to the node name of the target node and the second timestamp so as to display the segment outline. The segment outline includes an outline title, and each outline title has an outline timestamp.

[0150] According to another embodiment of the present disclosure, the target node determination module includes: a processing submodule and a node determination submodule. The processing submodule is used to filter the nodes in the structure tree according to at least one of the node repeatability, node abnormality and node quantity of the structure tree to obtain a filtered structure tree. The node determination submodule is used to determine the target node according to the leaf nodes in the filtered structure tree.

[0151] According to another embodiment of the present disclosure, the structure tree is divided into at least one level, and the path lengths from nodes at the same level to the root node are the same; the processing submodule includes: a first adding submodule and a filtering submodule. The first adding submodule is used to add marks to the nodes at the target level in the structure tree in response to detecting that the nodes at the target level in the structure tree meet the first condition; the first condition includes at least one of the following: the number of repeated nodes is greater than or equal to the repeated number threshold, the node repetition rate is greater than or equal to the repetition rate threshold, and the number of nodes is greater than or equal to the single-layer node number threshold. The filtering submodule is used to filter the nodes in the structure tree according to the marks to obtain a filtered structure tree.

[0152] According to another embodiment of the present disclosure, the structure tree includes at least one layer of nodes, and the path lengths from the nodes at the same level to the root node are the same. The processing submodule includes: a second adding submodule and a screening submodule. The second adding submodule is used for adding marks to the nodes at the target level in the subtree in the structure tree in response to detecting that the nodes at the target level in the subtree meet the second condition, and the second condition includes at least one of the following: the number of abnormal nodes is greater than or equal to a predetermined abnormal number, and the ratio of the number of abnormal nodes to the number of nodes at the target level of the subtree is greater than or equal to a predetermined abnormal ratio. The screening submodule is used to screen the nodes in the structure tree according to the marks to obtain a screened structure tree.

[0153] According to another embodiment of the present disclosure, it also includes: an abnormal node determination module, which is used to determine a node as an abnormal node in response to determining that the playback duration of the segment corresponding to the node is less than or equal to the predetermined duration, or the number of characters in the node name of the node is greater than the first number of characters.

[0154] According to another embodiment of the present disclosure, the screening submodule further includes: a first deletion submodule, a determination submodule and a second deletion submodule. The first deletion submodule is used to delete the marked nodes from the structure tree in response to determining that the nodes are marked, so as to obtain an intermediate structure tree. The determination submodule is used to determine whether to cancel the mark for the marked nodes according to the node collapse parameter or the level shuffling parameter of the intermediate structure tree. The second deletion submodule is used to delete the nodes that have not been cancelled, so as to obtain a structure tree after screening.

[0155] According to another embodiment of the present disclosure, it also includes: a trigger module, which is used to trigger the operation of determining the target node from the structure tree in response to detecting that the total duration of the object is greater than a predetermined duration, the number of characters in the initial text is greater than or equal to the second number of characters, and the number of leaf nodes in the structure tree is greater than a predetermined number.

[0156] According to another embodiment of the present disclosure, it further includes: a reference level determination module and a summary generation module. The reference level determination module is used to determine the reference level structure according to the segment outline. The summary generation module is used to generate a segment summary based on the large model, the initial text and the reference level structure so as to display the segment summary, wherein the segment summary includes a summary title and a summary content text.

[0157] According to another embodiment of the present disclosure, the reference hierarchy determination module includes: an omission rate determination submodule, a first determination submodule, and a second determination submodule. The omission rate determination submodule is used to determine the node omission parameters of the structure tree according to the difference between the key points included in the initial text and the nodes in the structure tree. The first determination submodule is used to determine the segment outline as the reference hierarchy structure in response to detecting that the node omission parameters meet the predetermined omission condition. The second determination submodule is used to determine the reference hierarchy structure according to the segment outline and the reference child node with the segment outline as the parent node in response to detecting that the node omission parameters do not meet the predetermined omission condition.

[0158] According to another embodiment of the present disclosure, the reference child node is determined by the following modules: a first module and a second module. The first module is used to determine the child node with the segment outline as the parent node in the structure tree as the reference child node when the structure tree includes the child node with the segment outline as the parent node. The second module is used to generate the reference child node based on the large model when the structure tree does not include the child node with the segment outline as the parent node.

[0159] According to another embodiment of the present disclosure, the predetermined omission condition includes at least one of the following: the node omission parameter includes a omission rate, and the omission rate is greater than or equal to a predetermined ratio; the node omission parameter includes a omission number, and the omission number is greater than or equal to a predetermined number.

[0160] According to another embodiment of the present disclosure, the omission rate determination submodule includes: an expansion unit and a parameter determination unit. The expansion unit is used to generate an expansion node for expanding the node of the subtree according to at least part of the initial text associated with the subtree in the initial text for the subtree whose level meets the predetermined level condition in the structure tree. The parameter determination unit is used to determine the node omission parameter according to the difference between the expansion node and the node in the subtree.

[0161] According to another embodiment of the present disclosure, the predetermined level condition includes: the number of levels of the subtree is less than or equal to the predetermined number of levels.

[0162] According to another embodiment of the present disclosure, the object is a video, and the initial text determination module includes: a first recognition submodule and a second recognition submodule. The first recognition submodule is used to recognize image frames in the video to obtain image text and a timestamp of the image text. The second recognition submodule is used to recognize speech in the video to obtain speech text and a timestamp of the speech text. Among them, the multiple subtexts in the initial text include image text and speech text, and the first timestamp includes image text, a timestamp of the image text, and a timestamp of the speech text.

[0163] According to another embodiment of the present disclosure, an outline display module is further included, and the outline display module includes: a division submodule and an outline display submodule. The division submodule is used to divide the progress bar of the object into multiple intervals according to the outline timestamp. The outline display submodule is used to display the outline title associated with the interval in each of the multiple intervals.

[0164] According to another embodiment of the present disclosure, it also includes: a mind map generation module is used to generate a mind map according to the nodes in the structure tree and the dependency relationships between the nodes, so as to display the mind map.

[0165] Fig.11 is a schematic structural block diagram of an evaluation device according to an embodiment of the present disclosure.

[0166] like Fig.11 As shown, the object evaluation device 1100 may include an acquisition module 1110 , a determination module 1120 , a node set determination module 1130 and an evaluation module 1140 .

[0167] The acquisition module 1110 is used to acquire the object to be evaluated and the reference object; the object to be evaluated is the hierarchical structure data determined by the large model based on the content of the audio and video data, and the object to be evaluated includes multiple nodes to be evaluated; the reference object is the real hierarchical structure data of the audio and video data, and the reference object includes multiple reference nodes.

[0168] The determination module 1120 is used to determine, for each node to be evaluated, whether the node to be evaluated meets a first predetermined condition according to the similarity between the node to be evaluated and at least one reference node.

[0169] The node set determination module 1130 is used to determine a first node set, a second node set and a third node set; wherein the first node set includes nodes to be evaluated in the object to be evaluated that meet the first predetermined condition and the second predetermined condition, the second node set includes nodes to be evaluated in the object to be evaluated that meet the second predetermined condition, and the third node set includes reference nodes in the reference object that meet the second predetermined condition.

[0170] The evaluation module 1140 is used to evaluate the object to be evaluated according to the first node set, the second node set and the third node set to obtain an evaluation result.

[0171] According to another embodiment of the present disclosure, the object to be evaluated includes a structure tree or a segment outline for audio and video data; the determination module includes: a first determination submodule, used to determine whether the node to be evaluated meets the first predetermined condition based on the similarity between the node name of the matched node to be evaluated and the node name of the reference node.

[0172] According to another embodiment of the present disclosure, the object to be evaluated includes a segmented summary of audio and video data; the node to be evaluated includes a summary content text, and the reference node includes a partial sub-text in the content text of the audio and video data; the determination module includes: a second determination sub-module, used to determine whether the node to be evaluated meets the first predetermined condition based on the similarity between the matched summary content text and the partial sub-text.

[0173] According to another embodiment of the present disclosure, the object to be evaluated includes one of a structure tree, a segment outline and a segment summary; the second predetermined condition includes at least one of the following: the node is a leaf node; the node is a node in a target node line, and the target node line includes a path from a root node to a leaf node; and the node is a node in a target subtree, and the target subtree includes multiple paths from the root node to the leaf nodes.

[0174] According to another embodiment of the present disclosure, the evaluation module includes: an accuracy sub-result sub-module, a comprehensiveness sub-result sub-module and a result determination sub-module. The accuracy sub-result sub-module is used to determine the accuracy sub-result of the object to be evaluated based on the number of nodes in the first node set and the number of nodes in the second node set. The comprehensiveness sub-result sub-module is used to determine the comprehensiveness sub-result of the object to be evaluated based on the number of nodes in the first node set and the number of nodes in the third node set. The result determination sub-module is used to determine the evaluation result based on the accuracy sub-result and the comprehensiveness sub-result.

[0175] According to another embodiment of the present disclosure, it further includes: a time evaluation module, which is used to determine the time evaluation result of the object to be evaluated according to the difference between the time stamp of the node to be evaluated and the time stamp of the reference node.

[0176] According to another embodiment of the present disclosure, the object to be evaluated includes a segmented outline, and the method further includes at least one of the following: a folding rate determination module, a character display rate determination module, and a hierarchical shuffle rate determination module. The folding rate determination module is used to determine the folding rate of the segmented outline according to the number of outline titles in a folded state in the segmented outline and the total number of outline titles in the segmented outline. The character display rate determination module is used to determine the character display rate of the segmented outline based on the total number of characters in the segmented outline and the number of characters displayed at the front end of the segmented outline. The hierarchical shuffle rate determination module is used to determine the hierarchical shuffle rate of the segmented outline according to the maximum number of outline titles at the same level in the segmented outline and the total number of outline titles in the segmented outline.

[0177] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, comprising at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.

[0178] According to an embodiment of the present disclosure, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the above method.

[0179] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, including a computer program, and the computer program implements the above method when executed by a processor.

[0180] Fig.12 It is a block diagram of an electronic device for implementing the object processing method and the object evaluation method of the embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0181] like Fig.12 As shown, the device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the device 1200 can also be stored. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0182] A number of components in the device 1200 are connected to the I / O interface 1205, including: an input unit 1206, such as a keyboard, a mouse, etc.; an output unit 1207, such as various types of displays, speakers, etc.; a storage unit 1208, such as a disk, an optical disk, etc.; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1209 allows the device 1200 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0183] The computing unit 1201 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1201 performs the various methods and processes described above, such as any of the above methods. For example, in some embodiments, any of the above methods may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 1208. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of any of the above methods described above may be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to perform any of the above methods in any other appropriate manner (e.g., by means of firmware).

[0184] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0185] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0186] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0187] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0188] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0189] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.

[0190] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0191] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for evaluating an object, comprising: Acquire an object to be evaluated and a reference object; the object to be evaluated is hierarchical structure data determined by the large model based on the content of the audio and video data, and the object to be evaluated includes multiple nodes to be evaluated; The reference object is real hierarchical structure data of the audio and video data, and the reference object includes a plurality of reference nodes; For each of the nodes to be evaluated, determining whether the node to be evaluated satisfies a first predetermined condition according to a similarity between the node to be evaluated and at least one of the reference nodes; Determine a first node set, a second node set, and a third node set; wherein the first node set includes the nodes to be evaluated in the object to be evaluated that meet the first predetermined condition and the second predetermined condition, the second node set includes the nodes to be evaluated in the object to be evaluated that meet the second predetermined condition, and the third node set includes the reference nodes in the reference object that meet the second predetermined condition; as well as The object to be evaluated is evaluated according to the first node set, the second node set and the third node set to obtain an evaluation result.

2. The method according to claim 1, wherein: The object to be evaluated includes a structure tree or a segment outline for the audio and video data; and determining whether the node to be evaluated satisfies a first predetermined condition according to the similarity between the node to be evaluated and at least one of the reference nodes includes: According to the similarity between the node name of the node to be evaluated and the node name of at least one of the reference nodes, it is determined whether the node to be evaluated meets the first predetermined condition.

3. The method according to claim 1, wherein: The object to be evaluated includes a segmented summary of audio and video data; the node to be evaluated includes a summary content text, the reference node includes a partial subtext in the content text of the audio and video data, and the summary content text and the partial subtext correspond to the same segment in the audio and video data; the determining whether the node to be evaluated satisfies the first predetermined condition based on the similarity between the node to be evaluated and at least one of the reference nodes includes: According to the similarity between the summary content text and the partial subtext, it is determined whether the node to be evaluated meets a first predetermined condition.

4. The method according to claim 1, wherein: The object to be evaluated has at least one hierarchical structure, and the second predetermined condition represents a condition that needs to be satisfied by the position of the node to be evaluated in the at least one hierarchical structure.

5. The method according to claim 4, wherein: The object to be evaluated includes one of a structure tree, a segment outline, and a segment summary; and the second predetermined condition includes at least one of the following: The node is a leaf node; The node is a node in a target node line, wherein the target node line includes a path from a root node to a leaf node; and The node is a node in a target subtree, and the target subtree includes multiple paths from a root node to a leaf node.

6. The method according to claim 1, wherein: The step of evaluating the object to be evaluated according to the first node set, the second node set, and the third node set to obtain an evaluation result includes: Determining the accuracy sub-result of the object to be evaluated according to the number of nodes in the first node set and the number of nodes in the second node set; Determining a comprehensiveness sub-result of the object to be evaluated according to the number of nodes in the first node set and the number of nodes in the third node set; and The evaluation result is determined according to the accuracy sub-result and the comprehensiveness sub-result.

7. The method according to claim 1, further comprising: The time evaluation result of the object to be evaluated is determined according to the difference between the time stamp of the node to be evaluated and the time stamp of the reference node.

8. The method according to claim 1, wherein: The object to be evaluated includes a segment outline, and the segment outline includes an outline title; the method further includes at least one of the following: Determining the folding rate of the segment outline according to the number of outline titles in the folded state in the segment outline and the total number of outline titles in the segment outline; Determine the character display rate of the segmented outline according to the total number of characters of the outline title in the segmented outline and the number of characters of the outline title displayed at the front end; as well as The level shuffling rate of the segment outline is determined according to the maximum number of outline titles at the same level in the segment outline and the total number of outline titles in the segment outline.

9. An object evaluation device, comprising: An acquisition module is used to acquire an object to be evaluated and a reference object; the object to be evaluated is hierarchical structure data determined by the large model based on the content of the audio and video data, and the object to be evaluated includes multiple nodes to be evaluated; The reference object is real hierarchical structure data of the audio and video data, and the reference object includes a plurality of reference nodes; A determination module, configured to determine, for each of the nodes to be evaluated, whether the node to be evaluated satisfies a first predetermined condition according to a similarity between the node to be evaluated and at least one of the reference nodes; A node set determination module, configured to determine a first node set, a second node set, and a third node set; wherein the first node set includes the nodes to be evaluated in the object to be evaluated that meet the first predetermined condition and the second predetermined condition, the second node set includes the nodes to be evaluated in the object to be evaluated that meet the second predetermined condition, and the third node set includes the reference nodes in the reference object that meet the second predetermined condition; as well as An evaluation module is used to evaluate the object to be evaluated according to the first node set, the second node set and the third node set to obtain an evaluation result.

10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.

12. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.