Text processing method and device, model training method and device, electronic equipment and medium

By automatically identifying the matching relationship between chapter titles and content in story text, and combining semantics, sentiment, theme matching degree and the importance of main characters, the problem of time-consuming manual annotation and susceptibility to subjective factors is solved, and key chapter identification is achieved efficiently and accurately.

CN120874808APending Publication Date: 2025-10-31BAIDU COM TIMES TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510884524.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In existing technologies, the identification of key chapters mainly relies on manual annotation, which is time-consuming and easily affected by subjective factors, and cannot efficiently process story texts of different types and styles.

Method used

By acquiring the chapter titles and content of the text to be processed, key chapters are automatically identified using matching relationships. Combining semantics, sentiment, topic matching degree, and the importance of main characters, a model training method is used to improve recognition efficiency and accuracy.

Benefits of technology

It enables automatic identification of key chapters in story texts of different types and styles, improving identification efficiency, reducing subjective influence, and ensuring the accuracy and efficiency of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120874808A_ABST
    Figure CN120874808A_ABST
Patent Text Reader

Abstract

The invention provides a text processing method and device, a model training method and device, electronic equipment and a medium, and relates to the field of artificial intelligence, in particular to the technical fields of NLP, intelligent retrieval and the like. According to the specific implementation scheme, a chapter title and chapter content corresponding to any chapter in a to-be-processed text are obtained, and the to-be-processed text corresponds to a main role; according to the matching relationship between the chapter content corresponding to any chapter and the chapter title and / or the main role, determining chapter information corresponding to any chapter; and determining key chapters in the to-be-processed text based on the chapter information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of AI (Artificial Intelligence) technology, specifically to the fields of NLP (Natural Language Processing) and intelligent retrieval, and particularly to a text processing method, model training method, device, electronic device and medium. Background Technology

[0002] Key chapters are important parts of novels and other story texts that serve to connect the preceding and following parts, drive the plot forward, showcase the core conflict, or deeply portray key characters. Accurately identifying key chapters helps users quickly grasp the story's structure and understand its core direction. Summary of the Invention

[0003] This disclosure provides a text processing method, a model training method, an apparatus, an electronic device, and a medium.

[0004] According to one aspect of this disclosure, a text processing method is provided, comprising:

[0005] Obtain the chapter title and chapter content corresponding to any chapter in the text to be processed, wherein the text to be processed corresponds to a main character;

[0006] Based on the matching relationship between the chapter content corresponding to any chapter and the chapter title and / or the main character, determine the chapter information corresponding to any chapter;

[0007] Based on the chapter information, the key chapters in the text to be processed are identified.

[0008] According to another aspect of this disclosure, a model training method is provided, comprising:

[0009] Obtain the model to be trained;

[0010] Obtain sample chapters and corresponding chapter tags for the sample chapters, wherein the chapter tags are used to indicate the matching relationship between the chapter content of the corresponding sample chapter and the chapter title and / or the main role of the text to which the sample chapter belongs;

[0011] Based on the sample chapters and their corresponding chapter labels, the model to be trained is trained to obtain the target model;

[0012] The target model is used to determine the chapter information corresponding to any chapter based on the matching relationship between the chapter content corresponding to any chapter in the text to be processed and the chapter title and / or the main characters in the text to be processed, so as to determine the key chapters in the text to be processed based on the chapter information.

[0013] According to another aspect of this disclosure, a text processing apparatus is provided, comprising:

[0014] The acquisition module is used to acquire the chapter title and chapter content corresponding to any chapter in the text to be processed, wherein the text to be processed corresponds to a main character;

[0015] The first determining module is used to determine the chapter information corresponding to any chapter based on the matching relationship between the chapter content corresponding to any chapter and the chapter title and / or the main character;

[0016] The second determining module is used to determine the key chapters in the text to be processed based on the chapter information.

[0017] According to another aspect of this disclosure, a model training apparatus is provided, comprising:

[0018] The first acquisition module is used to acquire the model to be trained;

[0019] The second acquisition module is used to acquire sample chapters and chapter tags corresponding to the sample chapters, wherein the chapter tags are used to indicate the matching relationship between the chapter content of the corresponding sample chapter and the chapter title and / or the main role of the text to which the sample chapter belongs;

[0020] The training module is used to train the model to be trained based on the sample chapters and the corresponding chapter labels to obtain the target model;

[0021] The target model is used to determine the chapter information corresponding to any chapter based on the matching relationship between the chapter content corresponding to any chapter in the text to be processed and the chapter title and / or the main characters in the text to be processed, so as to determine the key chapters in the text to be processed based on the chapter information.

[0022] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0023] At least one processor; and

[0024] A memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the text processing method proposed in one aspect of this disclosure or the model training method proposed in another aspect.

[0026] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing the computer to perform the text processing method proposed in one aspect of this disclosure or the model training method proposed in another aspect.

[0027] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the text processing method proposed in one aspect of this disclosure or the model training method proposed in another aspect.

[0028] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0029] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0030] Figure 1 This is a flowchart illustrating a text processing method provided according to an embodiment of the present disclosure;

[0031] Figure 2 This is a flowchart illustrating a text processing method according to another embodiment of the present disclosure;

[0032] Figure 3 This is a flowchart illustrating a text processing method according to another embodiment of the present disclosure;

[0033] Figure 4 This is a schematic flowchart of a model training method provided according to another embodiment of the present disclosure;

[0034] Figure 5 This is a schematic flowchart of a model training method provided according to another embodiment of the present disclosure;

[0035] Figure 6 A schematic diagram of the structure of a text processing apparatus according to an embodiment of the present disclosure;

[0036] Figure 7 A schematic diagram of the structure of a model training device according to an embodiment of the present disclosure;

[0037] Figure 8 This is a block diagram of an electronic device used to implement the text processing method or model training method of the embodiments of this disclosure. Detailed Implementation

[0038] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0039] In related technologies, the identification of key chapters mainly relies on manual annotation. However, many story texts are complex and have many chapters. The work of manually identifying key chapters is time-consuming and easily affected by subjective factors, making it impossible to efficiently process story texts of different types and styles.

[0040] Therefore, to address at least one of the aforementioned problems, this disclosure proposes a text processing method, a model training method, an apparatus, an electronic device, and a medium. This disclosure can be applied to digital reading platforms, intelligent retrieval systems, and literary research tools, helping users quickly grasp the story's plot and understand its core development.

[0041] The following description, with reference to the accompanying drawings, describes text processing methods, model training methods, apparatuses, electronic devices, and media according to embodiments of the present disclosure.

[0042] Figure 1 This is a flowchart illustrating a text processing method according to an embodiment of the present disclosure.

[0043] This disclosure illustrates the example of a text processing method configured in a text processing device, which can be applied to any electronic device to enable the electronic device to perform text processing functions.

[0044] Among them, electronic devices can be any device with computing capabilities, such as personal computers, mobile terminals, servers, etc. Mobile terminals can be hardware devices with various operating systems, touch screens and / or displays, such as mobile phones, tablets, personal digital assistants, wearable devices, etc.

[0045] like Figure 1 As shown, the text processing method may include the following steps S101 to S103:

[0046] Step S101: Obtain the chapter title and chapter content corresponding to any chapter in the text to be processed, wherein the text to be processed corresponds to a main character.

[0047] The texts to be processed include, but are not limited to, story texts such as novel texts, script texts, and biographical texts.

[0048] The text to be processed can be in any e-book format, including but not limited to: EPUB (Electronic Publication), MOBI (Mobipocket eBook Format), TXT (Plain Text File), and PDF (Portable Document Format).

[0049] In order to improve the efficiency of identifying key chapters, word segmentation and stop word removal can be performed on the chapter titles and chapter content in the text to be processed.

[0050] The main characters in the text to be processed can refer to the protagonist of the story, or to characters whose importance is greater than a set importance threshold.

[0051] As an example, character names can be extracted from the text to be processed; the importance of character names can be determined based on at least one of the following: the frequency of occurrence of character names in the text to be processed, positional distribution score, social centrality, and the degree of correlation between character names and emotional fluctuations in the text to be processed; and the main characters in the text to be processed can be determined based on the importance of character names.

[0052] To improve the accuracy and efficiency of character name extraction, the text to be processed can be formatted before extraction to avoid garbled characters. For example, an e-book parsing tool can be used to convert the text to be processed into UTF-8 (8-bit Unicode Transformation Format).

[0053] Step S102: Determine the chapter information corresponding to any chapter based on the matching relationship between the chapter content and the chapter title and / or the main character of any chapter.

[0054] The matching relationship between chapter content and chapter title can be used to indicate whether the chapter title and chapter content match; the matching relationship between chapter content and main character can be used to indicate whether the chapter content includes the main character, and / or to indicate the importance of the main character in the chapter content; chapter information can be used to determine the importance of a chapter or its recommendation priority.

[0055] As an example, the chapter information for any given chapter can be determined based on the matching relationship between the chapter content and the chapter title; or, the chapter information for any given chapter can be determined based on the matching relationship between the chapter content and the main character; or, the chapter information for any given chapter can be determined based on the matching relationship between the chapter content and the chapter title, as well as the matching relationship between the chapter content and the main character.

[0056] Step S103: Based on the chapter information, identify the key chapters in the text to be processed.

[0057] Among these features, the importance or recommendation priority of any chapter in the text to be processed can be determined based on chapter information; key chapters can be identified based on the importance or recommendation priority of the corresponding chapters.

[0058] The text processing method of this disclosure obtains the chapter title and chapter content corresponding to any chapter in the text to be processed, wherein the text to be processed corresponds to a main character; determines the chapter information corresponding to any chapter based on the matching relationship between the chapter content and the chapter title and / or the main character; and determines the key chapters in the text to be processed based on the chapter information. The chapter title is used to indicate the theme direction, and the main character is the core bearer of the plot. Therefore, the chapter information obtained based on the matching relationship between the chapter content and the chapter title and / or the main character is a deep condensation of the core features of the chapter. Thus, based on the chapter information, key chapters whose content closely follows the theme and / or revolves around the main character can be accurately selected. Furthermore, compared to manual selection of key chapters, this disclosure can automatically identify key chapters in different types and styles of text to be processed based on the chapter information, which is more efficient and less susceptible to subjective factors.

[0059] It should be noted that the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solution disclosed herein are all carried out with the consent of the user, and all comply with the provisions of relevant laws and regulations, and do not violate public order and good morals.

[0060] To clearly illustrate how chapter information is determined in any embodiment of this disclosure, this disclosure also proposes a text processing method.

[0061] Figure 2 This is a flowchart illustrating a text processing method provided according to another embodiment of the present disclosure.

[0062] like Figure 2 As shown, the text processing method may include the following steps S201 to S203:

[0063] Step S201: Obtain the chapter title and chapter content corresponding to any chapter in the text to be processed, wherein the text to be processed corresponds to a main character.

[0064] The explanation of step S201 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.

[0065] Step S202: Determine the chapter information corresponding to any chapter based on the matching relationship between the chapter content and the chapter title and / or the main character of any chapter; wherein, the chapter information includes at least one of the following: semantic matching degree, sentiment matching degree, theme matching degree between the chapter title and the chapter content, and the importance of the main character in the chapter content.

[0066] The importance of the main character in the chapter content can be used to indicate whether the main character has driven significant developments in the plot, such as triggering key conflicts or driving plot twists; the semantic matching degree between the chapter title and the chapter content can be used to indicate the degree of semantic similarity between the chapter title and the chapter content; the emotional matching degree between the chapter title and the chapter content can be used to indicate the consistency or similarity of the chapter title and the chapter content in terms of emotional expression; and the thematic matching degree between the chapter title and the chapter content can be used to indicate the consistency or similarity of the chapter title and the chapter content in terms of content theme.

[0067] In any embodiment of this disclosure, obtaining the importance of a main character in the chapter content includes: obtaining the number of times the main character appears in the chapter content; determining the behavioral importance score of the main character in the chapter content based on the behavioral events in which the main character participates; and determining the importance of the main character in the chapter content based on the number of appearances and the behavioral importance score.

[0068] Among them, the main characters are the core of the story's progression. The number of times the main characters appear in the chapter content can indicate whether the main characters are closely intertwined with the plot. The importance score of the main characters' actions in the chapter content can indicate whether the main characters' actions have a significant impact on the story's direction. Therefore, combining the two can accurately measure the value of the main characters in the chapter content, thereby helping to improve the accuracy of identifying key chapters.

[0069] Among them, the behavioral events involving the main characters can be identified through an event extraction model.

[0070] As an example, the importance score of a main character's actions in a chapter can be determined based on the number of action events the main character participates in within the chapter content.

[0071] As another example, syntactic analysis can be performed on the behavioral event to obtain the target behavioral event with the main character as the subject of the event; the behavioral actions of the main character in the target behavioral event and the number of times the main character's behavioral actions appear in the target behavioral event can be obtained; and the behavioral importance score can be determined based on the behavioral actions and the number of times the actions appear.

[0072] This process involves determining the actions of the main characters in each target event, resulting in a set of actions. For any given action, the number of times that action appears in the set is counted, yielding the frequency of that action. For example, if the main characters' actions in the three target events are "discover," "decide," and "discover," then the action "discover" appears 2 times, and the action "decide" appears 1 time.

[0073] Therefore, focusing on the target behavioral events of the main characters as the subjects of the events ensures that the objects of analysis are the key behaviors that truly drive the plot forward in the story; combining behavioral actions and the frequency of their occurrence to measure the importance of the main characters' behavior in the plot progression makes the determination of behavioral importance scores more objective and accurate.

[0074] As an example, the action score of a corresponding behavior can be determined based on the number of times the action occurs; and the importance score of a behavior can be determined based on the action score of each behavior.

[0075] As an example, the weights corresponding to the actions can be obtained; based on the weights corresponding to the actions, the number of times the actions occur is weighted and summed to obtain the importance score of the actions.

[0076] Different actions have different weights. For example, the weights of key actions such as "discovery", "decision", and "sacrifice" are greater than the weights of "walking" and "speaking".

[0077] Different actions have different effects on plot progression. Therefore, combining the weight of actions and the frequency of their occurrence to determine the importance score of actions can more accurately measure the importance of the main characters' actions in advancing the plot.

[0078] In any embodiment of this disclosure, obtaining the semantic matching degree between the chapter title and the chapter content includes at least one of the following: determining the semantic matching degree based on the similarity between the feature vector corresponding to the chapter title and the feature vector corresponding to the chapter content; determining the semantic matching degree based on the similarity between the feature vector corresponding to the chapter title and the feature vector corresponding to the content summary of the chapter content.

[0079] This involves calculating the cosine similarity or Euclidean distance between feature vectors, and determining the semantic matching degree based on the cosine similarity or Euclidean distance.

[0080] As an example, semantic matching can include the similarity between the feature vector corresponding to the chapter title and the feature vector corresponding to the chapter content, as well as the similarity between the feature vector corresponding to the chapter title and the feature vector corresponding to the content summary of the chapter content.

[0081] As another example, in order to balance the efficiency of obtaining semantic matching degree with the accuracy of key chapter identification, for instance, when the content length of the chapter content is less than a set length threshold, the semantic matching degree is determined based on the similarity between the feature vector corresponding to the chapter title and the feature vector corresponding to the chapter content; when the content length of the chapter content is greater than or equal to the set length threshold, the semantic matching degree is determined based on the similarity between the feature vector corresponding to the chapter title and the feature vector corresponding to the content summary of the chapter content.

[0082] Therefore, determining the semantic matching degree based on the feature vectors corresponding to the chapter titles and the chapter content can provide an overall understanding of the fit between the chapter titles and the chapter content; determining the semantic matching degree based on the feature vectors corresponding to the chapter titles and the feature vectors corresponding to the content summaries of the chapter content can compensate for semantic matching errors caused by excessively long chapter content and improve accuracy.

[0083] In any embodiment of this disclosure, obtaining the topic matching degree between chapter titles and chapter content includes: obtaining a first topic distribution corresponding to the chapter titles and a second topic distribution corresponding to the chapter content through a topic model; wherein, the first topic distribution is used to indicate the degree of matching between the corresponding chapter titles and each topic in the topic set, and the second topic distribution is used to indicate the degree of matching between the corresponding chapter content and each topic in the topic set; and determining the topic matching degree based on the similarity between the first topic distribution and the second topic distribution.

[0084] The topic model can refer to the LDA (Latent Dirichlet Allocation) model; the topic set includes multiple defined topics; the first topic distribution and the second topic distribution are probability distributions, where the first topic distribution reflects the probability distribution of chapter titles on each topic in the topic set, and the second topic distribution reflects the probability distribution of chapter content on each topic in the topic set.

[0085] Therefore, by obtaining the topic distribution of chapter titles and chapter content through topic modeling, it is possible to clearly present the degree of matching between them and each topic in the topic set. Furthermore, the topic matching degree can be determined based on the similarity of the topic distribution of the two, which can accurately quantify the consistency of chapter titles and chapter content in terms of topics.

[0086] In any embodiment of this disclosure, obtaining the sentiment matching degree between the chapter title and the chapter content includes: performing sentiment analysis on the chapter title to obtain a first sentiment distribution; performing sentiment analysis on the chapter content to obtain a second sentiment distribution; and determining the sentiment matching degree based on the similarity between the first sentiment distribution and the second sentiment distribution.

[0087] The first sentiment distribution indicates the degree of match between the corresponding chapter title and each sentiment in the sentiment set, while the second sentiment distribution indicates the degree of match between the corresponding chapter content and each sentiment in the sentiment set. For example, the sentiment set may include emotions such as happiness, sadness, anger, and fear.

[0088] Step S203: Based on the chapter information, identify the key chapters in the text to be processed.

[0089] In any embodiment of this disclosure, at least one of the semantic matching degree, sentiment matching degree, topic matching degree, and the importance of the main character in the chapter content of each chapter in the text to be processed can be input into the trained key chapter recognition model to obtain the key chapter output by the key chapter recognition model.

[0090] The text processing method of this disclosure determines key chapters in the text to be processed based on at least one of the following: semantic matching degree, sentiment matching degree, thematic matching degree between chapter titles and chapter content, and the importance of main characters in the chapter content. Semantic matching degree reflects the semantic consistency between chapter titles and chapter content; sentiment matching degree reflects the emotional consistency between chapter titles and chapter content; thematic matching degree reflects the thematic consistency between chapter titles and chapter content; and the importance of main characters reflects their role in driving the plot. Therefore, determining key chapters by comprehensively considering the semantic matching degree, sentiment matching degree, thematic matching degree between chapter titles and chapter content, and the importance of main characters in the chapter content can effectively filter out irrelevant chapters and improve the accuracy of key chapter identification.

[0091] To clearly illustrate how key chapters are determined based on chapter information in any embodiment of this disclosure, this disclosure also proposes a text processing method.

[0092] Figure 3 This is a flowchart illustrating a text processing method provided according to another embodiment of the present disclosure.

[0093] like Figure 3 As shown, the text processing method may include the following steps S301 to S303:

[0094] Step S301: Obtain the chapter title and chapter content corresponding to any chapter in the text to be processed, wherein the text to be processed corresponds to a main character.

[0095] Step S302: Determine the chapter information corresponding to any chapter based on the matching relationship between the chapter content and the chapter title and / or the main character of any chapter; wherein, the chapter information includes at least one of the following: semantic matching degree, sentiment matching degree, theme matching degree between the chapter title and the chapter content, and the importance of the main character in the chapter content.

[0096] The explanation of steps S301 and S302 can be found in the relevant descriptions in any embodiment of this disclosure, and will not be repeated here.

[0097] Step S303: Determine the score of the first chapter based on at least one of semantic matching degree, sentiment matching degree, topic matching degree, and importance degree; determine the key chapters based on the score of the first chapter.

[0098] For example, the weights corresponding to semantic matching degree, sentiment matching degree, topic matching degree, and importance can be obtained; based on the obtained weights, the semantic matching degree, sentiment matching degree, topic matching degree, and importance are weighted and summed to obtain the score of the first chapter.

[0099] As an example, the score of the first chapter can be compared with a first score threshold, and the chapters whose scores of the first chapter are greater than the first score threshold can be identified as key chapters.

[0100] Therefore, the importance of the main character in the chapter content can be used to indicate whether the main character has driven a significant development in the plot; the semantic matching degree between the chapter title and the chapter content can be used to indicate the degree of semantic similarity between the chapter title and the chapter content; the emotional matching degree between the chapter title and the chapter content can be used to indicate the consistency or similarity of the chapter title and the chapter content in terms of emotional expression; the thematic matching degree between the chapter title and the chapter content can be used to indicate the consistency or similarity of the chapter title and the chapter content in terms of content theme. Determining the first chapter score based on at least one of semantic matching degree, emotional matching degree, thematic matching degree, and importance degree can objectively quantify the chapter quality. Selecting key chapters based on the first chapter score can avoid the subjectivity of manual selection and improve the efficiency and accuracy of key chapter selection.

[0101] In any embodiment of this disclosure, the chapter information also includes at least one of plot density, information content, and plot twist score.

[0102] In any embodiment of this disclosure, chapters are filtered based on the first chapter score to obtain candidate chapters; the second chapter score of the candidate chapters is determined based on at least one of the plot density, information content, and plot twist score corresponding to the chapter content; and the key chapters are determined from the candidate chapters based on the second chapter score.

[0103] Among them, plot density can be used to indicate the tightness and richness of the plot in the chapter content; information content can be used to indicate the amount of effective information in the chapter content; and plot turning point score can be used to indicate whether there are key plot turning points in the chapter content.

[0104] As an example, the score of the second chapter can be compared with a second score threshold, and candidate chapters whose scores of the second chapter are greater than the second score threshold can be identified as key chapters.

[0105] Therefore, by using multiple parameters such as semantic matching, sentiment matching, topic matching, and importance to determine the score of the first chapter, and then using this score to initially filter out a large number of irrelevant chapters, the selection scope can be quickly narrowed down. Next, the score of the second chapter is determined based on factors such as plot density, information content, and plot twists, which can further focus on chapters with key plot points and rich content. Thus, this method effectively filters out irrelevant chapters, significantly improving the efficiency and accuracy of key chapter identification.

[0106] In any embodiment of this disclosure, the method further includes: sorting key chapters based on the first chapter score; obtaining adjacent chapter pairs whose difference in the first chapter score is less than a set threshold based on the sorted key chapters; determining the third chapter score corresponding to any chapter in the adjacent chapter pair based on at least one of the plot density, information content, and plot twist score corresponding to the chapter content; reordering the chapters in the adjacent chapter pair based on the third chapter score to obtain reordered key chapters; and generating chapter recommendation information based on the reordered key chapters.

[0107] The chapter recommendation information may include a chapter recommendation list, which includes chapter number, chapter title, chapter information, chapter summary, and other information.

[0108] Therefore, by first sorting based on the score of the first chapter, then querying adjacent chapter pairs with similar scores, and finally sorting based on the score of the third chapter, key chapters with plot twists can be recommended first, thus helping users quickly grasp the story's outline.

[0109] In any embodiment of this disclosure, the method further includes: identifying plot elements in the chapter content based on a plot recognition model and / or set plot keywords; and determining plot density based on the ratio of the number of plot elements to the number of characters in the chapter content.

[0110] The plot recognition model can refer to a trained model used to identify plot elements; plot keywords can include: discovering the truth, betrayal, death, marriage proposal, explosion, escape, etc.

[0111] For example, the ratio of the number of plot points to the number of characters in a chapter can be used as the plot density. For instance, if a chapter has 4000 characters and 15 plot points are identified, the corresponding plot density is 0.00375; if a chapter has 3000 characters and 3 plot points are identified, the corresponding plot density is 0.001.

[0112] Therefore, determining the plot density by the ratio of the number of plot points to the number of characters in a chapter can objectively quantify the compactness of the plot in a chapter, which is helpful for accurately selecting key chapters with dense plots and rich content.

[0113] In any embodiment of this disclosure, the method further includes: obtaining target information that first appears in the chapter content; classifying and statistically analyzing the target information according to the information type to obtain the number of information under any information type, and weighting and summing the number of information based on the weight corresponding to the information type to obtain the information amount; and / or determining the information amount based on the term frequency (TF) of the target information in the text to be processed and the inverse document frequency (IDF) of the target information in a set text library.

[0114] Each chapter has a corresponding chapter number. For any given chapter, the content of that chapter can be compared with the content of the chapter whose chapter number precedes it to obtain the target information that appears for the first time in that chapter.

[0115] The information types can include people, places, revelations of important secrets / truths, background information, explanations of rules, and changes in relationships between characters.

[0116] Therefore, by classifying and statistically summing information by type to obtain the information content, or by combining TF-IDF to determine the information content, we can accurately quantify the information richness of chapters, thus providing a reliable basis for the identification and recommendation of key chapters.

[0117] In any embodiment of this disclosure, the method further includes: performing emotion change detection on the chapter content, or performing emotion change detection on the chapter content and adjacent chapter content; detecting transition cue words in the chapter content; identifying whether there is a plot transition in the chapter content based on the emotion change detection results and / or transition cue word detection results; and determining a plot transition score based on the plot transition identification results.

[0118] Specifically, the system performs emotion change detection on the chapter content to determine whether there are any emotional changes within the chapter content; it also performs emotion change detection on the chapter content and the content of adjacent chapters to determine whether the chapter content exhibits emotional changes compared to the adjacent chapter content.

[0119] For example, if there is an emotional change within the chapter content, and / or if there is an emotional change in the chapter content compared to adjacent chapter content, then it is determined that there is a plot twist in the chapter content; if there is no emotional change within the chapter content, and there is no emotional change in the chapter content compared to adjacent chapter content, then it is determined that there is no plot twist in the chapter content.

[0120] This includes detecting transitional words in the chapter content to determine whether the chapter content contains such words. For example, transitional words may include: suddenly, unexpectedly, at that moment, etc.

[0121] For example, if the chapter content includes a transitional word, it is determined that there is a plot twist in the chapter content; if the chapter content does not include a transitional word, it is determined that there is no plot twist in the chapter content.

[0122] Therefore, emotional changes can reflect the emotional fluctuations of the text content, while turning point prompts directly indicate changes in the plot direction. Combining the two to determine the plot turning point score can accurately capture the climax, trough and turning point of the plot in the chapter, and thus accurately locate the plot's twists and turns.

[0123] The text processing method of this disclosure determines a first chapter score based on at least one of semantic matching degree, sentiment matching degree, topic matching degree, and importance degree; and determines key chapters based on the first chapter score. The first chapter score can objectively quantify the quality of chapters, avoid the subjectivity of manual selection, and improve the efficiency and accuracy of key chapter selection.

[0124] This disclosure also proposes a model training method. Figure 4 This is a schematic flowchart of a model training method provided according to another embodiment of the present disclosure.

[0125] This disclosure illustrates the example of the model training method being configured in a model training device, which can be applied to any electronic device to enable the electronic device to perform model training functions.

[0126] Among them, electronic devices can be any device with computing capabilities, such as personal computers, mobile terminals, servers, etc. Mobile terminals can be hardware devices with various operating systems, touch screens and / or displays, such as mobile phones, tablets, personal digital assistants, wearable devices, etc.

[0127] like Figure 4 As shown, the model training method may include the following steps S401 to S403:

[0128] Step S401: Obtain the model to be trained.

[0129] Here, the model to be trained can refer to the initial model that has not been trained.

[0130] Step S402: Obtain the sample chapter and the corresponding chapter tag. The chapter tag is used to indicate the matching relationship between the chapter content of the corresponding sample chapter and the chapter title and / or the main role of the text to which the sample chapter belongs.

[0131] The chapter tag can be used to indicate the matching relationship between the chapter content and the chapter title of the corresponding sample chapter; or, the chapter tag can be used to indicate the matching relationship between the chapter content and the main character of the text to which the sample chapter belongs; or, the chapter tag can be used to indicate the matching relationship between the chapter content and the chapter title and the main character of the text to which the sample chapter belongs.

[0132] As an example, chapter tags include at least one of the following: semantic matching tags between the chapter title and the chapter content of the sample chapter; sentiment matching tags between the chapter title and the chapter content of the sample chapter; topic matching tags between the chapter title and the chapter content of the sample chapter; and importance tags of the main characters in the text to which the sample chapter belongs within the chapter content of the sample chapter.

[0133] The semantic matching label indicates the degree of semantic matching between the chapter title and the chapter content of the sample chapter; the sentiment matching label indicates the degree of sentiment matching between the chapter title and the chapter content of the sample chapter; the topic matching label indicates the degree of topic matching between the chapter title and the chapter content of the sample chapter; and the importance label indicates the importance of the main characters in the text to which the sample chapter belongs within the chapter content of the sample chapter.

[0134] The importance of main characters within a chapter can indicate whether they significantly advance the plot; the semantic matching degree between chapter titles and content can indicate the degree of semantic similarity between them; the sentiment matching degree between chapter titles and content can indicate the consistency or similarity in their emotional expression; and the thematic matching degree between chapter titles and content can indicate the consistency or similarity in their thematic content. Therefore, training the model based on these chapter tags helps the target model output more accurate chapter information.

[0135] Step S403: Based on the sample chapters and their corresponding chapter labels, train the model to be trained to obtain the target model.

[0136] The target model is used to determine the chapter information corresponding to any chapter based on the matching relationship between the chapter content corresponding to any chapter in the text to be processed and the chapter title and / or the main characters in the text to be processed, so as to identify the key chapters in the text to be processed based on the chapter information.

[0137] Specifically, the model can be trained based on the difference between the predicted content output by the model to be trained for sample chapters and sample titles and the corresponding chapter labels, thus obtaining the target model.

[0138] The model training method of this disclosure includes: obtaining a model to be trained; obtaining sample chapters and corresponding chapter tags, wherein the chapter tags indicate the matching relationship between the chapter content of the corresponding sample chapter and the chapter title and / or the main character of the text to which the sample chapter belongs; and training the model to be trained based on the sample chapters and corresponding chapter tags to obtain a target model. The target model obtained by training the model to be trained can determine the corresponding chapter information based on the matching relationship between the chapter content of any chapter and the chapter title and / or the main character, thereby identifying key chapters in the text to be processed based on the chapter information. The chapter title indicates the theme direction, and the main character is the core bearer of the plot. Therefore, the chapter information obtained based on the matching relationship between the chapter content and the chapter title and / or the main character is a deep condensation of the core features of the chapter. Thus, based on the chapter information, key chapters whose content closely follows the theme and / or revolves around the main character can be accurately selected. Furthermore, compared to manual selection of key chapters, this disclosure can automatically obtain chapter information through the target model, and then identify key chapters in different types and styles of text to be processed based on the chapter information, which is more efficient and less susceptible to subjective factors.

[0139] The model to be trained is a pre-trained model, such as the BERT model; or, the model to be trained is a model obtained by training an initial model based on sample chapters and corresponding semantic matching labels; in order to clearly illustrate how the model to be trained is trained in any embodiment of this disclosure, this disclosure also proposes a model training method.

[0140] Figure 5 This is a schematic flowchart of a model training method provided according to another embodiment of the present disclosure.

[0141] like Figure 5 As shown, the model training method may include the following steps S501 to S503:

[0142] Step S501: Input the sample chapter title and sample chapter content into the model to be trained to obtain the predicted semantic matching degree, predicted sentiment matching degree and predicted topic matching degree output by the model to be trained.

[0143] The model to be trained can refer to a model with the ability to predict semantic matching degree.

[0144] Based on the fact that the model to be trained already has the ability to predict semantic matching degree, this disclosure combines topic matching features and sentiment matching features to optimize the model to improve the model's generalization and expressive ability.

[0145] Step S502: Determine the target loss based on the difference between the predicted semantic matching degree and the semantic matching degree label, the difference between the predicted sentiment matching degree and the sentiment matching degree label, and the difference between the predicted topic matching degree and the topic matching degree label.

[0146] Specifically, the target loss can be determined based on the difference between the predicted semantic matching degree and the semantic matching degree label, the difference between the predicted sentiment matching degree and the sentiment matching degree label, and the difference between the predicted topic matching degree and the topic matching degree label.

[0147] Step S503: Based on the target loss, train the model to be trained to obtain the target model.

[0148] In the case of training the model based on the target loss, regularization, Dropout and other methods can be used to prevent overfitting, and cross-validation and grid search can be used to tune the parameters.

[0149] It should be noted that, given that the model to be trained already has the ability to predict semantic matching degree, the model to be trained can be fine-tuned based on the target loss.

[0150] The model training method of this disclosure involves inputting sample chapter titles and content into the model to be trained, obtaining the predicted semantic matching degree, predicted sentiment matching degree, and predicted topic matching degree output by the model. Based on the differences between the predicted semantic matching degree and semantic matching degree labels, the differences between the predicted sentiment matching degree and sentiment matching degree labels, and the differences between the predicted topic matching degree and topic matching degree labels, a target loss is determined. Based on the target loss, the model to be trained is trained to obtain the target model. By optimizing the model to be trained by combining topic matching features and sentiment matching features, based on the model's existing semantic matching degree prediction capability, the generalization and expressive ability of the model can be improved, thereby enabling the target model to output a more accurate semantic matching degree.

[0151] With the above Figures 1 to 3 Corresponding to the text processing method provided in the embodiments, this disclosure also provides a text processing apparatus. Because the text processing apparatus provided in the embodiments of this disclosure is similar to the one described above... Figures 1 to 3 The text processing method provided in the embodiments corresponds to the text processing apparatus provided in the embodiments of this disclosure, and will not be described in detail in the embodiments of this disclosure.

[0152] Figure 6 This is a schematic diagram of the structure of a text processing apparatus provided according to an embodiment of the present disclosure.

[0153] like Figure 6 As shown, the text processing device 600 may include: an acquisition module 610, a first determination module 620, and a second determination module 630.

[0154] The acquisition module 610 is used to acquire the chapter title and chapter content corresponding to any chapter in the text to be processed, wherein the text to be processed corresponds to a main character;

[0155] The first determining module 620 is used to determine the chapter information corresponding to any chapter based on the matching relationship between the chapter content corresponding to any chapter and the chapter title and / or the main character;

[0156] The second determining module 630 is used to determine the key chapters in the text to be processed based on the chapter information.

[0157] In one possible implementation of this disclosure, the chapter information includes the importance of the main characters in the chapter content, and the first determining module 620 is specifically used for:

[0158] Get the number of times the main characters appear in the chapter content;

[0159] Based on the behavioral events that the main characters participate in within the chapter content, determine the importance score of the main characters' actions within the chapter content;

[0160] The importance of key characters in the chapter content is determined based on their frequency of appearance and the importance score of their actions.

[0161] In one possible implementation of this disclosure, the first determining module 620 is specifically used for:

[0162] Perform syntactic analysis on the action event to obtain the target action event with the main character as the subject of the event;

[0163] Obtain the actions of the main characters in the target behavior event, and the number of times the actions of the main characters appear in the target behavior event;

[0164] The importance score of a behavior is determined based on the behavior and the frequency of its occurrence.

[0165] In one possible implementation of this disclosure, the first determining module 620 is specifically used for:

[0166] Obtain the weights corresponding to the actions;

[0167] Based on the weights corresponding to the actions, the frequency of each action is weighted and summed to obtain the action importance score.

[0168] In one possible implementation of this disclosure, the chapter information includes the semantic matching degree between the chapter title and the chapter content, and the first determining module 620 is specifically used for at least one of the following:

[0169] The semantic matching degree is determined based on the similarity between the feature vectors corresponding to the chapter titles and the feature vectors corresponding to the chapter content.

[0170] The semantic matching degree is determined based on the similarity between the feature vector corresponding to the chapter title and the feature vector corresponding to the content summary of the chapter content.

[0171] In one possible implementation of this disclosure, the chapter information includes the topic matching degree between the chapter title and the chapter content, and the first determining module 620 is specifically used for:

[0172] Using a topic model, we obtain the first topic distribution corresponding to the chapter titles and the second topic distribution corresponding to the chapter content. The first topic distribution indicates the degree of matching between the corresponding chapter titles and each topic in the topic set, and the second topic distribution indicates the degree of matching between the corresponding chapter content and each topic in the topic set.

[0173] The topic matching degree is determined based on the similarity between the first topic distribution and the second topic distribution.

[0174] In one possible implementation of this disclosure, the chapter information includes at least one of the following: semantic matching degree, sentiment matching degree, theme matching degree between the chapter title and chapter content, and the importance of the main character in the chapter content; the second determining module 630 is specifically used for:

[0175] The score for the first chapter is determined based on at least one of the following: semantic matching degree, sentiment matching degree, topic matching degree, and importance degree.

[0176] Based on the score of the first chapter, the key chapters are identified.

[0177] In one possible implementation of this disclosure, the second determining module 630 is specifically used for:

[0178] Based on the score of the first chapter, chapters are filtered to obtain candidate chapters;

[0179] The second chapter score of a candidate chapter is determined based on at least one of the following: plot density, information content, and plot twist score.

[0180] Based on the score of the second chapter, key chapters are identified from the candidate chapters.

[0181] In one possible implementation of this disclosure, the apparatus further includes a chapter recommendation module, used for:

[0182] Based on the score of the first chapter, the key chapters are sorted.

[0183] Based on the sorted key chapters, obtain adjacent chapter pairs whose difference in scores of the first chapter is less than a set threshold;

[0184] Based on at least one of the plot density, information content and plot twist score of the chapter content, determine the score of the third chapter corresponding to any chapter in the adjacent chapter pair;

[0185] Based on the score of the third chapter, the chapters in adjacent chapter pairs are reordered to obtain the reordered key chapters;

[0186] Based on the reordered key chapters, chapter recommendation information is generated.

[0187] In one possible implementation of this disclosure, the apparatus further includes a third determining module, configured to:

[0188] Based on the plot recognition model and / or the set plot keywords, identify the plot in the chapter content;

[0189] Plot density is determined by the ratio of the number of plot points to the number of characters in the chapter content.

[0190] In one possible implementation of this disclosure, the apparatus further includes a fourth determining module, configured to:

[0191] Retrieve the target information that appears for the first time in the chapter content;

[0192] Based on the information type, the target information is classified and statistically analyzed to obtain the number of information under any information type. Based on the weight corresponding to the information type, the number of information is weighted and summed to obtain the information volume. And / or, the information volume is determined based on the word frequency of the target information in the text to be processed and the inverse document frequency of the target information in the set text library.

[0193] In one possible implementation of this disclosure, the apparatus further includes a fifth determining module, configured to:

[0194] Perform sentiment change detection on the chapter content, or perform sentiment change detection on the chapter content and the adjacent chapter content;

[0195] Detect transition words in the chapter content;

[0196] Based on the results of emotion change detection and / or transition cue word detection, identify whether there are plot twists in the chapter content;

[0197] Based on the plot twist identification results, a plot twist score is determined.

[0198] The text processing apparatus of this disclosure acquires the chapter title and chapter content corresponding to any chapter in the text to be processed, wherein the text to be processed corresponds to a main character; based on the matching relationship between the chapter content corresponding to any chapter and the chapter title and / or the main character, the chapter information corresponding to any chapter is determined; based on the chapter information, key chapters in the text to be processed are determined. The chapter title is used to indicate the theme direction, and the main character is the core bearer of the plot. Therefore, the chapter information obtained based on the matching relationship between the chapter content and the chapter title and / or the main character is a deep condensation of the core features of the chapter. Thus, based on the chapter information, key chapters whose content closely follows the theme and / or revolves around the main character can be accurately selected. In addition, compared with the method of manually selecting key chapters, this disclosure can automatically identify key chapters in different types and styles of text to be processed based on the chapter information, which is more efficient and less affected by subjective factors.

[0199] With the above Figures 4 to 5 Corresponding to the model training method provided in the embodiments, this disclosure also provides a model training apparatus. Since the model training apparatus provided in the embodiments of this disclosure is similar to the one described above... Figures 1 to 5The model training method provided in the embodiments corresponds to the model training device provided in the embodiments of this disclosure, and will not be described in detail in the embodiments of this disclosure.

[0200] Figure 7 This is a schematic diagram of the structure of a model training device provided according to an embodiment of the present disclosure.

[0201] like Figure 7 As shown, the model training device 700 may include: a first acquisition module 710, a second acquisition module 720, and a training module 730.

[0202] The first acquisition module 710 is used to acquire the model to be trained;

[0203] The second acquisition module 720 is used to acquire sample chapters and corresponding chapter tags for the sample chapters. The chapter tags are used to indicate the matching relationship between the chapter content of the corresponding sample chapter and the chapter title and / or the main characters of the text to which the sample chapter belongs.

[0204] Training module 730 is used to train the model to be trained based on the sample chapters and corresponding chapter labels to obtain the target model.

[0205] The target model is used to determine the chapter information corresponding to any chapter based on the matching relationship between the chapter content corresponding to any chapter in the text to be processed and the chapter title and / or the main characters in the text to be processed, so as to identify the key chapters in the text to be processed based on the chapter information.

[0206] In one possible implementation of this disclosure, the section labels include at least one of the following:

[0207] Sentiment matching labels between the chapter titles and chapter content of the sample chapters;

[0208] The topic matching degree tags between the chapter titles and chapter content of the sample chapters;

[0209] The importance of the main characters in the text to which the sample chapter belongs within the chapter content is labeled.

[0210] In one possible implementation of this disclosure, the model to be trained is a pre-trained model; or, the model to be trained is a model obtained by training an initial model based on sample chapters and corresponding semantic matching labels; the training module 730 is specifically used for:

[0211] Input the sample chapter title and sample chapter content into the model to be trained to obtain the predicted semantic matching degree, predicted sentiment matching degree and predicted topic matching degree output by the model to be trained.

[0212] The target loss is determined based on the differences between predicted semantic matching degree and semantic matching degree label, predicted sentiment matching degree and sentiment matching degree label, and predicted topic matching degree and topic matching degree label.

[0213] Based on the target loss, the model to be trained is trained to obtain the target model.

[0214] The model training apparatus of this disclosure acquires a model to be trained; acquires sample chapters and corresponding chapter tags, wherein the chapter tags are used to indicate the matching relationship between the chapter content of the corresponding sample chapter and the chapter title and / or the main character of the text to which the sample chapter belongs; and trains the model to be trained based on the sample chapters and corresponding chapter tags to obtain a target model. The target model obtained by training the model to be trained can determine the corresponding chapter information based on the matching relationship between the chapter content of any chapter and the chapter title and / or the main character, thereby identifying key chapters in the text to be processed based on the chapter information. The chapter title is used to indicate the theme direction, and the main character is the core bearer of the plot. Therefore, the chapter information obtained based on the matching relationship between the chapter content and the chapter title and / or the main character is a deep condensation of the core features of the chapter. Thus, based on the chapter information, key chapters whose content closely follows the theme and / or revolves around the main character can be accurately selected. Furthermore, compared to manual selection of key chapters, this disclosure can automatically acquire chapter information through the target model, and then identify key chapters in different types and styles of text to be processed based on the chapter information, which is more efficient and less susceptible to subjective factors.

[0215] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0216] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device 800 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0217] like Figure 8As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 802 or a computer program loaded from storage unit 808 into RAM (Random Access Memory) 803. The RAM 803 can also store various programs and data required for the operation of the electronic device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An I / O (Input / Output) interface 805 is also connected to the bus 804.

[0218] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows electronic device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0219] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as text processing methods and / or model training methods. For example, in some embodiments, the text processing methods and / or model training methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the text processing methods and / or model training methods described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform text processing methods and / or model training methods by any other suitable means (e.g., by means of firmware).

[0220] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0221] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0222] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0223] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0224] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.

[0225] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0226] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0227] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0228] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A text processing method, comprising: Obtain the chapter title and chapter content corresponding to any chapter in the text to be processed, wherein the text to be processed corresponds to a main character; Based on the matching relationship between the chapter content corresponding to any chapter and the chapter title and / or the main character, determine the chapter information corresponding to any chapter; Based on the chapter information, the key chapters in the text to be processed are identified.

2. The method according to claim 1, wherein, The chapter information includes the importance of the main character in the chapter content. Determining the chapter information corresponding to any chapter based on the matching relationship between the chapter content and the chapter title and / or the main character includes: Obtain the number of times the main character appears in the chapter content; Based on the behavioral events in which the main character participates in the chapter content, determine the behavioral importance score of the main character in the chapter content; The importance of the main character in the chapter content is determined based on the number of times it appears and the importance score of the behavior.

3. The method according to claim 2, wherein, The determination of the importance score of the main character's behavior in the chapter content based on the behavioral events the main character participates in includes: Perform syntactic analysis on the behavioral event to obtain the target behavioral event with the main character as the subject of the event; Obtain the actions of the main character in the target behavior event, and the number of times the actions of the main character appear in the target behavior event; The importance score of the behavior is determined based on the behavior and the number of times the behavior occurs.

4. The method according to claim 3, wherein, The determination of the importance score of the behavior based on the behavior and the frequency of the behavior includes: Obtain the weight corresponding to the behavior action; Based on the weights corresponding to the actions, the frequency of occurrence of the actions is weighted and summed to obtain the importance score of the actions.

5. The method according to claim 1, wherein, The chapter information includes the semantic matching degree between the chapter title and the chapter content. Determining the chapter information corresponding to any chapter based on the matching relationship between the chapter content corresponding to any chapter and the chapter title and / or the main character includes at least one of the following: The semantic matching degree is determined based on the similarity between the feature vector corresponding to the chapter title and the feature vector corresponding to the chapter content; The semantic matching degree is determined based on the similarity between the feature vector corresponding to the chapter title and the feature vector corresponding to the content summary of the chapter content.

6. The method according to claim 1, wherein, The chapter information includes the theme matching degree between the chapter title and the chapter content. Determining the chapter information corresponding to any chapter based on the matching relationship between the chapter content corresponding to any chapter and the chapter title and / or the main character includes: Using a topic model, a first topic distribution corresponding to the chapter title and a second topic distribution corresponding to the chapter content are obtained; wherein, the first topic distribution is used to indicate the degree of matching between the corresponding chapter title and each topic in the topic set, and the second topic distribution is used to indicate the degree of matching between the corresponding chapter content and each topic in the topic set; The topic matching degree is determined based on the similarity between the first topic distribution and the second topic distribution.

7. The method according to claim 1, wherein, The chapter information includes at least one of the following: semantic matching degree, sentiment matching degree, theme matching degree between the chapter title and the chapter content, and the importance of the main character in the chapter content; The step of determining the key chapters in the text to be processed based on the chapter information includes: The score of the first chapter is determined based on at least one of the semantic matching degree, the sentiment matching degree, the topic matching degree, and the importance degree. Based on the score of the first chapter, the key chapters are determined.

8. The method according to claim 7, wherein, The chapter information also includes at least one of plot density, information content, and plot twist score. Determining the key chapters based on the first chapter score includes: Based on the scores of the first chapter, chapters are filtered to obtain candidate chapters; The second chapter score of the candidate chapter is determined based on at least one of the plot density, information content, and plot twist score corresponding to the chapter content. Based on the score of the second chapter, the key chapter is determined from the candidate chapters.

9. The method according to claim 7, wherein, The method further includes: Based on the scores of the first chapter, the key chapters are ranked; Based on the sorted key chapters, obtain adjacent chapter pairs whose difference in scores of the first chapter is less than a set threshold; Based on at least one of the plot density, information content and plot twist score of the chapter content, determine the score of the third chapter corresponding to any chapter in the adjacent chapter pair; Based on the score of the third chapter, the chapters in the adjacent chapter pairs are reordered to obtain the reordered key chapters; Based on the reordered key chapters, chapter recommendation information is generated.

10. The method according to claim 8 or 9, wherein, The method further includes: Based on the plot recognition model and / or the set plot keywords, identify the plot plots in the chapter content; The plot density is determined based on the ratio of the number of plot points to the number of characters in the chapter content.

11. The method according to claim 8 or 9, wherein, The method further includes: Obtain the target information that first appears in the content of the chapter; Based on the information type, the target information is classified and statistically analyzed to obtain the number of information under any information type, and the number of information is weighted and summed based on the weight corresponding to the information type to obtain the information volume; and / or, the information volume is determined based on the word frequency of the target information in the text to be processed and the inverse document frequency of the target information in a set text library.

12. The method according to claim 8 or 9, wherein, The method further includes: Emotional change detection is performed on the content of the chapter, or emotional change detection is performed on the content of the chapter and the content of adjacent chapters. The content of the aforementioned chapters was tested for transitional words; Based on the results of emotion change detection and / or transition cue word detection, identify whether there is a plot twist in the content of the chapter; Based on the plot twist identification results, the plot twist score is determined.

13. A model training method, comprising: Obtain the model to be trained; Obtain sample chapters and corresponding chapter tags for the sample chapters, wherein the chapter tags are used to indicate the matching relationship between the chapter content of the corresponding sample chapter and the chapter title and / or the main role of the text to which the sample chapter belongs; Based on the sample chapters and their corresponding chapter labels, the model to be trained is trained to obtain the target model; The target model is used to determine the chapter information corresponding to any chapter based on the matching relationship between the chapter content corresponding to any chapter in the text to be processed and the chapter title and / or the main characters in the text to be processed, so as to determine the key chapters in the text to be processed based on the chapter information.

14. The method according to claim 13, wherein, The chapter tags include at least one of the following: Sentiment matching labels between the chapter titles and chapter content of the sample chapters; The topic matching degree tags between the chapter titles and chapter content of the sample chapters; The labels indicate the importance of the main characters in the text to which the sample chapter belongs within the chapter content of the sample chapter.

15. The method according to claim 14, wherein, The model to be trained is a pre-trained model; or, the model to be trained is a model obtained by training an initial model based on sample chapters and corresponding semantic matching degree labels. The step of training the model to be trained based on the sample chapters and corresponding chapter labels to obtain the target model includes: Input the sample chapter title and the sample chapter content into the model to be trained to obtain the predicted semantic matching degree, predicted sentiment matching degree and predicted topic matching degree output by the model to be trained; The target loss is determined based on the differences between the predicted semantic matching degree and the semantic matching degree label, the differences between the predicted sentiment matching degree and the sentiment matching degree label, and the differences between the predicted topic matching degree and the topic matching degree label. Based on the target loss, the model to be trained is trained to obtain the target model.

16. A text processing apparatus, comprising: The acquisition module is used to acquire the chapter title and chapter content corresponding to any chapter in the text to be processed, wherein the text to be processed corresponds to a main character; The first determining module is used to determine the chapter information corresponding to any chapter based on the matching relationship between the chapter content corresponding to any chapter and the chapter title and / or the main character; The second determining module is used to determine the key chapters in the text to be processed based on the chapter information.

17. A model training device, comprising: The first acquisition module is used to acquire the model to be trained; The second acquisition module is used to acquire sample chapters and chapter tags corresponding to the sample chapters, wherein the chapter tags are used to indicate the matching relationship between the chapter content of the corresponding sample chapter and the chapter title and / or the main role of the text to which the sample chapter belongs; The training module is used to train the model to be trained based on the sample chapters and the corresponding chapter labels to obtain the target model; The target model is used to determine the chapter information corresponding to any chapter based on the matching relationship between the chapter content corresponding to any chapter in the text to be processed and the chapter title and / or the main characters in the text to be processed, so as to determine the key chapters in the text to be processed based on the chapter information.

18. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-12 or 13-15.

19. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-12 or 13-15.

20. A computer program product comprising a computer program that, when executed by a processor, implements the method of any one of claims 1-12 or 13-15.