A documentary story writing system and a documentary story writing device
Through the documentary story writing system, a large model is used to talk to users to generate documentary stories that conform to their personal life and personality, solving the problem of unable to automatically record and reproduce personal memories in the existing technology, and achieving efficient and personalized story generation.
Patent Information
- Application Number
- CN202411111997.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-08-14
AI Technical Summary
The existing story continuation system cannot record and reproduce personal memories without the user's conception in advance, and the generated stories cannot fully reflect the user's personality and emotions, resulting in low satisfaction.
It provides a documentary story writing system, including theme setting unit, voice dialogue unit and story generation unit. Through a big model, it conducts voice dialogue with users, recognizes intentions, manages dialogue context, explores problems, and generates documentary stories that conform to the user's life and personality.
It can effectively generate structured, emotional and personal characteristics without user conception, improve user satisfaction and is especially suitable for the elderly.
Smart Images

Figure CN119272732B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly to a documentary story writing system and a documentary story writing device. Background Art
[0002] Existing story continuation and article generation systems often require the author to have a certain conception of the story content (plot, characters) in advance. The main purpose of the system is to diverge thinking (forming multiple possible storylines) and generate stylized texts. After receiving the initial story input by the user, the characters in the dialogue are determined. As the starting point of the dialogue generation process, the basis of the story and the participating characters are determined. "Writing Frog", a popular online literature creation assistant, mainly functions to continue the story according to the user's existing desire for expression of the theme and plot.
[0003] However, the expansion or dialogue function in the known systems focuses on diverging imagination, controlling the twists and climaxes of complex plot designs, adjusting the rhythm of the characters' entrance and exit, creating various characters with different personalities to match the plot, and ultimately achieving the purpose of attracting readers. However, these are contrary to the principles of writing personal memoirs (documentary literature). Therefore, in the process of interaction between the system and the user and the collection of user responses for the purpose of confirming interest, it is impossible to aim at excavating the user's personal experiences and emotional states, and to control the allocation among open-ended questions, detailed inquiries, and emotional confirmation strategies.
[0004] Currently, there does not exist such a system that, without any prior writing conception by the user, aims at recording and reproducing personal memories, obtains user information indexed by the chronological order of the user's life, and generates a story based on this.
[0005] Currently, there does not exist such a system that generates articles according to the writing styles of personal biographies and documentary prose. When using a general or other writing style model to write personal stories, the diction and emotional expression cannot fully interpret the story and the user's personality, resulting in low satisfaction. For example, large models tend to generate copy with strong emotional colors, such as "Whenever I recall that spring morning, I am filled with emotion", but not all authors appreciate such an expression when narrating their own experiences. Summary of the Invention
[0006] Embodiments of the present invention provide a documentary story writing system and a documentary story writing device, which can solve the technical problem in the prior art that "without any prior writing conception by the user, it is impossible to aim at recording and reproducing personal memories, obtain user information indexed by the chronological order of the user's life, and generate a story based on this".
[0007] To achieve the above object, in a first aspect, an embodiment of the present invention provides a documentary story writing system, including:
[0008] A theme setting unit, configured to guide a user to determine a theme of a documentary story, and extract content parameters related to the theme from a database according to the theme;
[0009] A voice dialogue unit, configured to carry out a voice dialogue with the user around the theme content based on a large model in combination with the content parameters, and mine and raise questions by means of identifying the intention of the user's reply and managing the dialogue context, dynamically guiding the user to share personal experiences and emotions, and forming a current dialogue record when the voice dialogue ends;
[0010] A story generation unit, configured to generate a documentary story according to the literary style and the current dialogue record after the user selects a literary style, where the text style has corresponding generation rules and algorithms, so that the story conforms to literary artistry and reflects the user's life and personality at the same time.
[0011] In a second aspect, an embodiment of the present invention provides a documentary story writing device, where the documentary story writing device has a visual interface or does not have a visual interface, and the documentary story writing device includes:
[0012] A processor; and a memory arranged to store executable instructions, where the executable instructions, when executed, cause the processor to implement through the foregoing documentary story writing system;
[0013] The documentary story writing device further includes:
[0014] A start and reset key, configured to enter a documentary story writing mode or re-enter another documentary story writing mode;
[0015] A dialogue control key, starting the dialogue control key to enter a voice dialogue mode, and generating a documentary story through the theme setting unit, the voice dialogue unit and the story generation unit.
[0016] The above technical solution has the following beneficial effects: The voice dialogue for the purpose of collecting personal experience information can automatically extract key information from the life events and memories provided by the user, effectively collect, process and utilize the user's personal information, and generate a structured, emotionally rich and personalized memoir article, that is, generate a documentary story. It can provide a simple, efficient and more personalized story generation tool without sacrificing text quality, and can be used by users who hope to write personal memoirs or have a documentary nature, especially suitable for the elderly. It can solve the problem in the prior art that it is impossible to automatically and accurately generate a documentary text that highly meets personalized needs according to personal memories and life experiences. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 is the logical structure diagram of a documentary story writing system according to an embodiment of the present invention;
[0019] Figure 2 is the flow chart of writing a documentary story according to an embodiment of the present invention;
[0020] Figure 3 is the framework diagram of the hardware structure and the software and hardware dual front ends according to an embodiment of the present invention. Specific Embodiments
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0022] As Figure 1 shown, in combination with the embodiments of the present invention, a documentary story writing system is provided, including:
[0023] A theme setting unit 11, configured to guide a user to determine the theme of a documentary story, and extract content parameters related to the theme from a database, where the content parameters are used to construct an initial direction and a dialogue depth of the documentary story;
[0024] A voice dialogue unit 12, configured to carry out a voice dialogue with the user around the theme content based on a large model in combination with the content parameters, identify the intention of the user's reply, manage the dialogue context, mine questions and ask questions, dynamically guide the user to share personal experiences and emotions, and form a current dialogue record when the voice dialogue ends;
[0025] A story generation unit 13, configured to generate a documentary story according to the selected literary style and the current dialogue record after the user selects the literary style.
[0026] The prior art mainly focuses on story continuation based on existing creative ideas and stylized text generation, and has a high dependence on user input. Especially in dealing with personal memories and documentary literature creation, the existing systems cannot effectively capture and reproduce the user's personalized and detailed personal experiences.
[0027] A documentary story writing system according to an embodiment of the present invention, especially a voice conversation for the purpose of collecting personal experience information, can automatically extract key information from the life events and memories provided by users, effectively collect, process and utilize the personal information of users, and generate memoir articles that are structured, rich in emotion and conform to personal characteristics, that is, generate documentary stories. It can provide a simple, efficient and more personalized story generation tool without sacrificing text quality, which can be used by users who hope to write personal memoirs or have a documentary nature, especially suitable for the elderly. It can solve the problem in the prior art that it is impossible to automatically and accurately generate documentary texts that highly meet personalized needs according to personal memories and life experiences.
[0028] Preferably, the theme setting unit 11 includes a first theme setting subunit or a second theme setting subunit:
[0029] The first theme setting subunit is used to guide the user to select a theme from the preset themes as the theme of the documentary story through a conversation on the UI interface or a button.
[0030] The second theme setting subunit is used to guide the user to speak out the content information points that the theme may contain through a large model, and logically concatenate all the content information points to obtain an initial theme. The large model expands, selects and refines the content information points and the theme background of the initial theme to obtain the final theme; wherein, during the process of the large model expanding, selecting and refining the content information points and the theme background of the initial theme, Internet retrieval can be used for connection.
[0031] Preferably, the theme setting unit 11 further includes:
[0032] The content parameter extraction subunit is used to extract the plot guiding factors related to the theme from the database, and extract the inspiring questions corresponding to the plot guiding factors, and use the inspiring questions as the content parameters related to the theme. The content parameters are used as prompt words during the voice conversation to guide the large model to ask questions to the user when the large model conducts a voice conversation with the user around the theme content, and construct the initial direction and conversation depth of the documentary story.
[0033] Among them, the plot guiding factors of various themes and the inspiring questions transformed according to the plot guiding factors are stored in the database. The plot guiding factors are extracted from a large number of documentary literatures, and the extraction of the plot guiding factors at least includes: event type, cause of contradiction, relationship between characters and personal feelings.
[0034] Preferably, this documentary story writing system further includes:
[0035] An information collection unit, which is used to, if the user is not generating a documentary story for the first time, after guiding the user to determine the theme of the documentary story, obtain the user's basic information from the database and provide it to the large model. The obtained user's basic information is used to avoid repeated questioning of the user during the process of the large model having a voice conversation with the user around the theme content. Among them, the basic information includes: gender, age, hometown, educational level, and highest honor.
[0036] Preferably, the voice conversation unit 12 includes a single large model-based agent, and the single large model-based agent is used for:
[0037] Combined with the content parameters, during the process of having a conversation with the user around the theme content, use an emotional feedback strategy to ask questions to the user, and form the current conversation record when the voice conversation ends; the emotional feedback strategy includes: confirming positive emotions and soothing negative emotions;
[0038] Or,
[0039] The single large model-based agent combined with the content parameters, during the process of having a voice conversation with the user around the theme content, use a continued questioning strategy to ask questions to the user, and form the current conversation record when the voice conversation ends; the continued questioning strategy includes: open-ended questions, continuing to dig into a certain character, asking for details of the event, scenario description, contradictory assumptions, asking for feelings, and switching topics;
[0040] Among them, if the user is not generating a documentary story for the first time, when the single large model-based agent has a voice conversation with the user around the theme content, use the memory link as the opening statement of the voice conversation to ask questions to the user; afterwards, for each group of voice conversations, the single large model-based agent simultaneously asks at least one question. When at least two questions are asked, guide the user to choose one to answer.
[0041] Preferably, the voice conversation unit 12 further includes multiple large model-based agents, and the multiple agents include: a decision maker, an opinion collector, and a detail miner, where:
[0042] The multiple large model-based agents are used to, combined with the content parameters, collaborate to ask questions to the user during the process of having a voice conversation with the user around the theme content, and form the current conversation record when the voice conversation ends; where:
[0043] The decision maker is used to control the rhythm of asking for details and changing topics, and according to the rhythm of asking for details and changing topics, decide the working time and working goals of the opinion collector and the detail miner respectively, and form the current conversation record when the voice conversation ends;
[0044] An opinion collector, responsible for determining the background of the times based on the generated voice conversation, and forming corresponding questions in combination with the common sense of the background of the times. The corresponding questions include at least one of the following: emotional identification, contradictory assumptions, and opinion inquiries; among them, the method of determining the background of the times includes: searching for background information of the times on the Internet based on the generated voice conversation to determine the background of the times;
[0045] A detail miner, responsible for mining stories related to the specific events mentioned and / or stories related to specific characters, raising questions based on the stories related to specific events and / or stories related to specific characters, and supplementing the stories related to specific events and / or stories related to specific characters as supplementary content to the current conversation record;
[0046] Among them, if the user generates a documentary story for non-first time, when the multi-agent based on the large model conducts a voice conversation with the user around the theme content, the memory link is used as the opening statement of the voice conversation to ask the user questions; afterwards, for each group of voice conversations, the multi-agent based on the large model simultaneously raises at least one question. When at least two questions are raised, the user is guided to select and answer.
[0047] Preferably, the story generation unit 13 includes a first generation sub-unit and a second generation sub-unit. After each voice conversation unit forms the current conversation record, a documentary story is generated through the first generation sub-unit or the second generation sub-unit, where:
[0048] The first generation sub-unit is used to select the style of writing after the voice conversation. According to at least one group of voice conversations selected from the current conversation record and the selected style of writing, the text style has generation rules and algorithms for generating a documentary story that conforms to artistry and reflects the user's life and character; through the synthesis key, at least one group of voices selected from the current conversation record is integrated into a narrative story draft according to the generation rules and algorithms. If the user modifies it through the editing key, the modified story draft is used as the documentary story. If the user does not modify it, the story draft is used as the documentary story, and the generated documentary story is in text form, voice form or video form;
[0049] The second generation sub-unit is used to select the style of writing after the voice conversation and generate a documentary story through multiple steps according to the style of writing and the current conversation record; generating a documentary story through multiple steps according to the style of writing and the current conversation record includes:
[0050] First, automatically generate a thread running through the entire text based on the current conversation record; second, enhance the artistry of the current conversation record. The methods for enhancing the artistry of the article include: polishing the beginning, adding descriptions of the environment, and increasing dialogues; third, adjust the article structure by setting suspense to increase the interest of the documentary story. The ways of setting suspense include: different time jumps and character perspective limitations; fourth, generate a documentary story according to the corresponding generation rules and algorithms of the described writing style. The generated documentary story is in the form of text, voice, or video.
[0051] Preferably, the documentary story writing system further includes an analysis unit, which is provided in the documentary story writing system or in a remote server. The analysis unit is used for:
[0052] Convert the current conversation into text form;
[0053] Analyze the current conversation record in text form, generate a continuity question based on the current conversation record, use the continuity question as a memory link for the opening remarks of a voice conversation on the same topic next time, and save the memory link to the database;
[0054] Analyze the current conversation record in text form, extract basic information from the current conversation record. The basic information includes: gender, age, family members, hometown, health status, educational level, and highest honor. Compare the extracted basic information with the basic information in the database, and update or replace the same basic information in the database;
[0055] Analyze the current conversation record in text form, extract core elements related to mobilizing memories and writing memoirs from the current conversation record, and save them to the database. The core elements include: relevant dates, relevant people, memories, major life experiences, major events, and personal hobbies;
[0056] Analyze the current conversation record in text form, extract key information other than core elements from the current conversation record, compare the extracted key information with the key information in the database, and update, replace, or leave unchanged the similar key information in the database;
[0057] Among them, the key information includes: core elements, and historical conversation records with a semantic similarity higher than the similarity threshold to the current conversation record;
[0058] User information includes: core elements, and basic information.
[0059] Preferably, the documentary story writing system further includes a music generation unit, a picture generation unit, and a video generation unit, where:
[0060] The music generation unit arranges the documentary story into a poem through a large model, uses the poem as a prompt to input into a music generation large model, and combines the music style selected by the user to generate a song.
[0061] The picture generation unit is used to extract the picture elements in the documentary story (for example: bicycle, plum blossom bricks at the foot of the city gate, catching butterflies under the small mountain, stealing melons in the field) and generate corresponding copywriting for the picture elements. The copywriting is a more concise text description. Input the copywriting as a prompt into an open-source picture generation model (such as StableDiffusion) that has been stylistically fine-tuned to generate pictures.
[0062] The video generation unit is used to receive the photos that the user can upload through the picture import unit, and by calling the video generation API service interface, generate corresponding short videos for the photos and the generated pictures respectively. Use the short videos, the set matching copywriting, and the music as video materials, and splice these video materials into a video.
[0063] Combined with the embodiments of the present invention, a documentary story writing device is provided. The documentary story writing device may or may not have a visual interface. The documentary story writing device includes:
[0064] A processor; and a memory arranged to store executable instructions, which when executed cause the processor to implement through any one of the foregoing documentary story writing systems;
[0065] The documentary story writing device further includes:
[0066] A start and reset key, used to enter a documentary story writing mode, or re-enter another documentary story writing mode;
[0067] A dialogue control key. When the dialogue control key is activated, it enters the voice dialogue mode, and a documentary story is generated through the theme setting unit, the voice dialogue unit, and the story generation unit.
[0068] Preferably, the documentary story writing device further includes: a dialogue history list for the user to review previous conversation records;
[0069] A dialogue record display unit, implemented through a dialogue record display key, for viewing the dialogue content;
[0070] A story list, used to view all generated documentary stories. By clicking, it enters the story display page for editing and sharing;
[0071] A story generation unit, capable of generating a documentary story and also editing the documentary story;
[0072] The Story Square is used to browse documentary stories shared by other users and has like and comment buttons;
[0073] Memory Cards are used to view one's own story memory cards on the page.
[0074] As Figure 2 shown, in combination with Figure 3 , the steps to implement documentary story writing using the documentary story writing system of the embodiments of the present invention are as follows:
[0075] 1. Guide the user to determine the conversation topic, extract content parameters related to a specific topic from the database (topic material database), and the content parameters are used for the initial direction and depth of the conversation to ensure that the conversation is carried out directionally and purposefully. This database is pre-constructed and contains core story guiding factors for various topics, such as the cause of the event, the relationship between characters, etc., ensuring the richness and pertinence of the conversation content.
[0076] (1) Specifically, guide the user to select one from the preset topics such as childhood impressions, childhood playmates, family memories, educational enlightenment... through the physical keys, touch keys in the UI interface, or the physical keys without a UI interface to start the conversation, and automatically match the most relevant topic tags according to the user's short description. A custom topic conversation can also be started. Topic tags include but are not limited to: childhood playmates, family memories, youthful friendships; and also include some other personalized and optional topics: military career, foreign lands, etc. Among them, topic tags, such as character relationships, life events, emotional changes, etc. These tags are classified, sorted, and labeled by a professional team based on historical data, literary works, and psychological research.
[0077] (2) In a custom topic conversation, the system will first guide the user to tell the content that the topic may contain, and the content has sparse and fragmented plots (information points). Use a large model to string these together with a certain logic to define the content of this topic), obtain the initial topic, and then define and expand the topic content to obtain the final topic. Background expansion can also be performed. The steps of expanding the topic may include multiple calls to the large model (even online retrieval) to expand, select, and refine the topic content. For example, ask the user through the large model what kind of story they want to tell. Support the adaptation of multiple large pre-trained language models (collectively referred to as large models below), such as: GPT-4, GLM-4, Dark Side of the Moon, ERNIE Bot, Tongyi Qianwen, etc.
[0078] For users without usage records, regarding a topic, such as "the home in childhood", it will start with an open-ended question first, and then, according to the user's answer, ask questions about the themes that the user may be interested in and that can enrich the personal story content.
[0079] If there is an accumulation of the user's conversation records, questions will be asked based on the user's background. For example: "You once mentioned the story of picking wild vegetables. Would you like to talk about it again?" Next, the system will autonomously adjust to pursue details or switch to a new topic. For example: "Did you find any other pleasures at your grandmother's house?" "Then let's change the topic. Who was your best friend when you were a child?"
[0080] The content parameters may include some colloquial question examples related to the theme and descriptions of the conversation strategies related to the theme. For example: What were the school conditions like when you were a child? How did you go to school? Did you go with anyone? When the user mentions a person and indicates a deep friendship, ask in detail about the person's background in an attempt to understand the person's behaviors and decisions.
[0081] (3) For various themes, plot guiding factors such as event types, causes of contradictions, character relationships, and personal feelings have been extracted from a large number of documentary literatures, and these plot guiding factors have been transformed into inspiring questions and stored in the database to be used as prompt words to guide the large model to ask questions during the conversation.
[0082] 2. Obtain user information related to the theme: After guiding the user to determine the theme of the documentary story through the preset information collection unit, obtain the user's basic information from the database (user information database), and the system can extract key information from the user's narrative, such as dates, important events of relevant characters, preferences, etc., and compare and supplement this information with the historical key information in the database. At the same time, as background knowledge for the large model, the advantage is that the large model will not ask repeated questions when chatting with the user again, such as how many people are in the family.
[0083] The large model itself does not have the "memory ability", so memory management has become an important topic for applications built based on the large model. Storing overly long memories will not only cause high costs and a decline in the user experience brought about by latency, but also increase the complexity of instructions, posing a challenge to the instruction following ability of existing large models. In the embodiments of the present invention, core elements in the key information related to mobilizing memories and writing memoirs are extracted from the interaction records between the user and the system for storage. The core elements include, but are not limited to: memories about people, major life experiences, personal preferences, etc.
[0084] Core elements: Currently include dates, relevant characters, major events, preferences, extracted from past conversation records.
[0085] Key information: Includes core elements and historical conversation records with high semantic similarity to the current conversation.
[0086] User information: This concept is the broadest. It includes core elements, key information, and other basic information: gender, age, hometown, educational level, highest honor, etc.
[0087] 3. The system uses a large language model to dynamically adjust the dialogue strategy and conduct dynamic dialogue management around the subject content (i.e., topic). It combines the set initial direction and depth of the dialogue, and guides users to share their personal experiences and emotions in depth through intent recognition, sentiment analysis tools, and context management technology to form a dialogue record. In particular, the system analyzes the user's emotional response and topic adaptability in real time during the conversation. If the user shows hesitation or discomfort, the system will automatically switch to a more soothing topic or adopt a more delicate questioning method (guided by the large model through prompt words). Questioning strategies include: open-ended questions, focusing on personal resumes, asking about situational details, exploring event contradictions, paying attention to the feelings of the interviewee, guiding the topic, etc.
[0088] In order to improve the user experience, the big model asks questions for non-first conversations, especially in the first sentence of each conversation to attract users by establishing memory links. Specifically, the conversation will first echo the last conversation, and then the questions asked are based on the known content of the last conversation, and cut from a different angle. For example: Last time we talked about your family having parents and a brother, and the experience of visiting your grandma's house was also very interesting, so what is the most unforgettable time you spent with your family? Are there any fixed family activities or traditions?
[0089] After the conversation on each topic is over, a memory link is generated in advance as the opening statement for the next conversation to ensure the memory link between the new conversation and the old conversation.
[0090] (1) Single agent: The dialogue format includes two parts: emotional feedback and continued questioning. Emotional feedback strategies include but are not limited to confirming positive emotions and soothing negative emotions. Continued questioning strategies include but are not limited to open-ended questions, continuing to explore a certain character, asking for event details, scenario descriptions, contradictory assumptions, inquiring about feelings, and switching topics. For each group of voice dialogues, the single agent based on the large model asks at least one question at the same time. When there are at least two questions raised, the user is guided to choose from them to answer. The number of questions can be set according to the size of the user interaction interface and the user's operation difficulty. For example, in the web version, three questions can be issued at the same time, and the user can choose one of the three to answer, thereby providing the user with the freedom to control the direction of the story.
[0091] (2) Multi-agent: A good interview needs to consider many factors, such as: controlling the rhythm of asking details and changing the topic according to the user's response; combining emotional identification, contradictory assumptions, and opinion inquiries with common sense of the times; comparing the information collected with the elements required for a good story, and judging the next question. Elements include, for example, the user only provided scattered information "Once I mixed up my sheep with someone else's sheep, and the old man didn't get angry, but just quietly sorted the sheep again." This lacks story elements: what kind of person is the old man, who was there at the time, why did he mix up the sheep, and what lessons did he learn afterwards.
[0092] With the intelligence capabilities of existing large models, it is difficult to take all aspects into consideration at the same time, so a multi-agent collaborative information collection solution is adopted. Depending on the cost and delay acceptance of the usage scenario, a multi-agent model can be selected. The system can include but is not limited to three agents: decision makers, opinion collectors, and detail miners.
[0093] Decision makers determine when and to what end the idea gatherers and detail miners work.
[0094] The opinion collector is responsible for searching the Internet for information to obtain the historical context. For example, if he determines that the character's experience fits the label "going into business in the 1980s", he will search the Internet for controversial points related to this, and then combine this with common sense of the historical context to conduct emotional identification, contradictory assumptions, and opinion inquiries.
[0095] Detail miners are responsible for unearthing stories about specific events and specific people, which are used to supplement the conversation record and can also be used to elicit questions to enrich the content of the conversation.
[0096] 4. Store the conversation records in the database and extract key information from them as user memory.
[0097] After the conversation is over, the current conversation record is formed, the content of the current conversation record is analyzed, key information such as major events, emotional changes, etc. are extracted, and stored for subsequent use. Multiple key information related to the topic is extracted from the conversation record, and these key information are matched and updated with the user's existing information, such as updating health status or changes in family members. The purpose is to manage the user memory stored in the database. Because if the user memory only increases and does not decrease, it will cause confusion, so the comparison and replacement method is adopted to manage the user memory.
[0098] In order to improve the accuracy of information processing and reduce redundancy, the traditional vector calculation semantic similarity technology is not used, but the memory is maintained through the topic unit. A large-scale language model, i.e. a large model, is used to judge story generation and memory update. By analyzing the relevance and information density of the conversation content, it is decided whether to update or replace the memory unit of a specific topic. This can avoid memory conflicts, improve the relevance of memory, and ensure that only sufficiently important or detailed information is updated.
[0099] At the beginning of each conversation, the system automatically generates an engaging opening statement based on updated topic-related memory links, aiming to increase the fun of the conversation and attract user engagement.
[0100] 5. Generate a story based on the conversation record and the style selected by the user, as shown in 6. The system designs corresponding generation rules and algorithms for different text styles (such as documentary, romantic, tragic, etc.) to ensure that the story not only conforms to literary artistry, but also accurately reflects the user's life and personality. The story generation module uses a hybrid algorithm that combines rule-driven and data-driven methods. The rule-driven part builds the story skeleton based on the basic rules of literary creation (for example: adding a sketch of the environment, filling in the dialogue between characters, using specific and vivid details to explain the emotions of the characters, etc.), while the data-driven part uses a deep learning model to fill in specific details and emotional colors. Users can choose the tone of the story (such as simplicity, strong literary color, sensational, etc.), and the system will adjust the story content according to these preferences to ensure that each story reflects the user's personality and emotions.
[0101] (1) The system allows users to quickly integrate a short story draft into a plain narrative by selecting only a group of dialogues and a style. The user can modify it if he wants. I provide editable functions (such as text boxes) on the application, which can view the dialogue records of a certain year, month, and day at a certain time in the morning on the front end of the software. For example, the first dialogue record generated by Grandma Liu is: OK, when I was a child, I always lived at my grandma's house. There were three or four houses around, so I often ran around with her during the summer vacation...; the second one is: We don't build houses with complex structures, such as very high towers, and now I think...
[0102] (2) At the same time, users can also generate more artistic articles through multiple steps. The first step is to use the dialogue material to determine an obvious clue that runs through the entire article. The second step is to enhance the artistry of the article by polishing the beginning, adding environmental descriptions, and dialogues. The third step is to set up suspense again through different time jumps, character perspective restrictions, and other information arrangements, and adjust the article structure again to add interest.
[0103] For example, part of the description of user Grandma Liu in the conversation record is:
[0104] Grandma Liu: I don't really remember these actually.
[0105] Grandma Liu: He has always been the solution for all the big and small illnesses in our family. For example, when I had a fever, I would go there to get an injection. When I broke my leg and had a fracture and rested for half a year, I also went there to change the dressing every time. Including when my grandfather later got cancer and was on his deathbed, he also drank traditional Chinese medicine there all the time. Although we knew it was useless, they could always give you care there, at least seemingly professional medical support. It's really a good kind of model. In the city now, which patient doesn't have some complaints about the hospital? I don't know if high medical skills are really what our society should pursue.
[0106] Grandma Liu: Yes, you're absolutely right. The connections in the village are very close. Everyone will help you with everything. It's very difficult for you to do something alone. And it seems that whatever you do, it will have an impact on a large group of people.
[0107] Generate a draft of all of Grandma Liu's descriptions this time, then save the draft. The generated draft is:
[0108] In that childhood time shrouded in mist, Hao Yinsu was my inseparable little friend. She was timid. Something as simple as learning to ride a bicycle seemed like a huge adventure to her. And I, who had already been galloping in the wind, was willing to slow down and become her solid support. Whenever it rained and the muddy road became impassable, we would abandon the bike and walk. That was a long but wonderful way to school. Our laughter echoed in the thick fog, and the frost on our eyebrows was like Santa Claus's beard, making our world full of the warmth of a fairy tale.
[0109] Remember once, Yinsu couldn't ride a bike because of illness, and I skipped class to send her home. That time, there was a firm light in her eyes, and she was determined to learn to ride a bike. Finally, she succeeded in junior high school. Although she could only ride but couldn't stop, and needed to rely on the steps or my help every time she got on and off the bike. I always rode a little faster and then waited for her in front. Her shout full of panic and victory still rings in my ears to this day.
[0110] During those days, I became her guardian, and a sense of pride welled up. I accompanied her, semester after semester, rain or shine. The warmth in her heart and the pride in my eyes became the firmest bridge of our friendship.
[0111] However, as time passed by, we each embarked on different life journeys. Working and getting married, we were like sand scattered by the torrent of time, each settling in a corner of life. Our connections gradually diminished, but those shared moments, those paths we walked together through the wind and snow, still shine with an eternal light in my heart. Hao Yinsu, she is the warmest page in my childhood memories and will always be worthy of my affectionate perusal.
[0112] A documentary story writing system or a documentary story writing system device according to an embodiment of the present invention is a story sharing application for elderly users. By conversing with a large language model, it records and generates personal stories. The application supports functions such as multiple theme selections, conversation record preservation, historical conversation viewing, story generation and editing, and story sharing. It aims to provide a convenient platform for elderly users to record their life experiences and share them with others. It can tell life stories by conversing with a large language model, can review and edit the recorded conversations and generated stories, and can share the generated stories with family and friends. The use is made more interesting by customizing themes and Easter egg themes.
[0113] The documentary story writing system device may have a visual interface or, as Figure 3 shown, can combine the specific work of both software and hardware fronts at the same time. The process of using the documentary story writing system device is described as follows:
[0114] 1. The front end is responsible for user registration and login. The user presses the "Start / Reset" button on the hardware terminal to start a conversation. The conversation with the user is completed through the voice conversation unit. The user's voice is collected through the voice collection unit of the hardware terminal, and the voice is recorded and sent to the server for further processing. The user is allowed to choose to hold down the "Conversation" button to speak, or the natural telephone is voice interaction. Among them, in the process corresponding to the user, in response to the user's reply, the large model converts the response into audio through text-to-speech (TTS) technology.
[0115] 2. As needed, pictures can be imported through the picture import unit and displayed on the image display unit; after the conversation ends, a documentary story is generated through the story generation unit; the user can view the "Conversation History List" to review the conversation records at a certain time on a certain date in the morning. For example:
[0116] A: Hello! Nice to chat with you. I heard that your childhood was filled with warm stories of villagers and neighbors. I'm very interested. Can you share some extraordinary stories that happened between you and your neighbors during your childhood?
[0117] B: Sure. When I was a kid, I lived at my grandma's house. My cousin's house was just three or four houses away from mine. So during the summer vacation, I often played with her. Her grandma took care of her, and my grandma took care of me. She was only one year younger than me. Next to my house was my old aunt's house. They ran a grocery store. We often went there to buy snacks. Well, it was more like taking some snacks without really buying them. I remember the bubble gum was what we bought most often. Across from my house was a family of hereditary doctors. When I was a kid, whenever I had a fever or a fracture, I would go there to see a doctor. Just one injection and I would get better.
[0118] A: Hearing you talk like this, it seems your childhood was full of warmth and joy. The bubble gum you mentioned reminds me of the taste of my childhood. What kind of taste was it? Also, when you played with your cousin, were there any particularly unforgettable games or experiences?
[0119] B: I remember at my cousin's house, there was a plastic toy car. There was a hole in the bottom of the car. We looked inside through the hole. There might have been a small plastic button inside. My sister insisted that it was a piece of candy. We tried every way to...
[0120] 3. Users can select a certain conversation record and enter the "Conversation Record Display" page to view the conversation content. They can also choose to generate a story and enter the "Story Generation" page to generate and edit the story.
[0121] 4. Users can view all the generated stories in the "Story List", click to enter the "Story Display" page for further editing and sharing.
[0122] 5. Users can browse the stories shared by other users in the "Story Square", like and comment on them. They can view their own user memories on the "Memory Card" page.
[0123] The beneficial technical effects achieved by the embodiments of the present invention are as follows:
[0124] 1. Provide a friendly and highly interactive platform that enables elderly users to easily tell and record their life stories. It can reduce the initial thinking burden of users during the article creation process without the need to have a detailed preconception of the story in advance.
[0125] 2. Automatically allocate open-ended questions, detail follow-ups, and emotional confirmation strategies to better explore and reflect users' personal experiences and emotional states. The high-quality story generation and editing functions can improve user satisfaction.
[0126] 3. Increase user stickiness and promote user interaction through the story sharing and story square functions.
[0127] 4. The generated articles are closer to users' real feelings and personal styles, improving the personalization of the text and the sense of resonance among readers. It supports the adaptation of multiple large models, i.e., large pre-trained language models, such as GPT-4, GLM-4, Darkside of the Moon, ERNIE Bot, Tongyi Qianwen, etc. It can instantly connect to the interfaces of other large models when one large model fails to work, so it has good flexibility and high robustness.
[0128] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the scope of the present disclosure. The appended method claims present the elements of the various steps in an exemplary order and are not intended to be limited to the specific order or hierarchy described.
[0129] In the above detailed description, various features are combined in a single embodiment to simplify the present disclosure. This method of disclosure should not be construed as reflecting an intention that the embodiments of the claimed subject matter require more features than are clearly recited in each claim. On the contrary, as reflected in the appended claims, the present invention lies in a state with fewer features than all the features of the disclosed single embodiment. Therefore, the appended claims are hereby expressly incorporated into the detailed description, where each claim stands alone as a separate preferred embodiment of the present invention.
[0130] To enable any person skilled in the art to implement or use the present invention, the above-described disclosed embodiments have been described. For those skilled in the art, various modification methods of these embodiments are obvious, and the general principles defined herein can also be applied to other embodiments without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments given herein, but is consistent with the broadest scope of the principles and novel features disclosed in this application.
[0131] The above description includes examples of one or more embodiments. Of course, it is impossible to describe all possible combinations of components or methods for describing the above embodiments, but those of ordinary skill in the art should recognize that the various embodiments can be further combined and arranged. Therefore, the embodiments described herein are intended to cover all such changes, modifications, and variations that fall within the scope of the appended claims. In addition, with respect to the term "comprising" used in the specification or claims, this term is covered in a manner similar to the term "including," as interpreted when "including" is used as a transitional word in the claims. In addition, any term "or" used in the specification of the claims is intended to mean "non-exclusive or."
[0132] Those skilled in the art can also understand that the various illustrative logical blocks, units, and steps listed in the embodiments of the present invention can be implemented by electronic hardware, computer software, or a combination of both. To clearly show the interchangeability of hardware and software, the above-mentioned various illustrative components, units, and steps have been generally described in terms of their functions. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art can use various methods to implement the described functions for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of the present invention.
[0133] In the embodiments of the present invention, the various illustrative logical blocks or units can be implemented or operate the described functions through a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of the above designs. The general-purpose processor can be a microprocessor. Optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors combined with a digital signal processor core, or any other similar configuration.
[0134] The steps of the methods or algorithms described in the embodiments of the present invention can be directly embedded in hardware, software modules executed by a processor, or a combination of the two. The software modules can be stored in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be disposed in an ASIC, and the ASIC can be disposed in a user terminal. Optionally, the processor and the storage medium can also be disposed in different components of the user terminal.
[0135] In one or more exemplary designs, the functions described in embodiments of the present invention may be implemented in hardware, software, firmware, or any combination of the three. If implemented in software, these functions may be stored on a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. A computer-readable medium includes both computer storage media and communication media that facilitate transfer of a computer program from one place to another. The storage media may be any available media that can be accessed by a general or special purpose computer. By way of example, and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store program code in the form of instructions or data structures and that can be accessed by a general or special purpose computer, or a general or special purpose processor. In addition, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, DSL, or wireless means such as infrared, radio, and microwave, it is included in the definition of computer-readable medium. Disk and disc include compact disc, laser disc, optical disc, DVD, floppy disk, and Blu-ray disc, where disks usually reproduce data magnetically, while discs usually reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0136] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A documentary story writing system, characterized in that, Including: A theme setting unit, configured to guide a user to determine the theme of a documentary story, and extract content parameters related to the theme from a database according to the theme; A voice dialogue unit, configured to, based on a large model and in combination with the content parameters, conduct a voice dialogue with the user around the theme content, and by means of identifying the intention of the user's reply and managing the dialogue context, dig out questions and ask questions, dynamically guide the user to share personal experiences and emotions, and form a current dialogue record when the voice dialogue ends; A story generation unit, configured to generate a documentary story according to the literary style selected by the user and the current dialogue record; The voice dialogue unit includes a single large-model-based intelligent agent, or further includes multiple large-model-based intelligent agents; The single large-model-based intelligent agent is used for: In combination with the content parameters, during the process of conducting a dialogue with the user around the theme content, asking questions to the user by adopting an emotional feedback strategy, and forming a current dialogue record when the voice dialogue ends; The emotional feedback strategy includes: confirming positive emotions and soothing negative emotions; Or, The single large-model-based intelligent agent in combination with the content parameters, during the process of conducting a voice dialogue with the user around the theme content, asks questions to the user by adopting a continuous questioning strategy, and forms a current dialogue record when the voice dialogue ends; the continuous questioning strategy includes: open-ended questions, continuously digging out a certain character, pursuing event details, scenario description, contradictory assumptions, asking for feelings, and switching topics; Among them, if the user is not generating a documentary story for the first time, when the single large-model-based intelligent agent conducts a voice dialogue with the user around the theme content, a memory link is used as the opening statement of the voice dialogue to ask questions to the user; afterwards, for each group of voice dialogues, the single large-model-based intelligent agent simultaneously asks at least one question, and when at least two questions are asked, the user is guided to select and answer from them; The multiple large-model-based intelligent agents include: a decision maker, a view collector, and a detail digger, where: The multiple large-model-based intelligent agents are used for, in combination with the content parameters, during the process of conducting a voice dialogue with the user around the theme content, collaborating to ask questions to the user, and forming a current dialogue record when the voice dialogue ends; where: The decision maker is used for controlling the rhythm of pursuing details and changing topics, determining the working time and working objectives of the view collector and the detail digger according to the rhythm of pursuing details and changing topics, and forming a current dialogue record when the voice dialogue ends; The view collector is used for being responsible for determining the background of the times according to the generated voice dialogue, and forming corresponding questions in combination with the common sense of the background of the times, and the corresponding questions include at least one of the following: emotional identification, contradictory assumptions, and view inquiry; among them, the way of determining the background of the times includes: searching for background information of the times on the Internet according to the generated voice dialogue to determine the background of the times; Detail miners are responsible for mining stories related to the specific events mentioned and / or stories related to specific individuals, raising questions based on the stories related to specific events and / or stories related to specific individuals, and supplementing the stories related to specific events and / or specific individuals as supplementary content to the current conversation record; Among them, if the user generates a documentary story for non - first time, when the multi - agent based on the large - model conducts a voice conversation with the user around the theme content, the memory link is used as the opening statement of the voice conversation to ask the user questions; afterwards, for each group of voice conversations, the multi - agent based on the large - model simultaneously raises at least one question. When at least two questions are raised, the user is guided to select and answer from them.
2. The documentary story writing system according to claim 1, characterized in that, The theme setting unit includes a first theme setting sub - unit or a second theme setting sub - unit: The first theme setting sub - unit is used to guide the user to select a theme from the preset themes as the theme of the documentary story through a conversation on the UI interface or buttons; The second theme setting sub - unit is used to guide the user to tell the content information points that the theme may contain through the large - model, and logically connect all the content information points to obtain the initial theme. The large - model expands, selects, and refines the content information points and theme background of the initial theme to obtain the final theme; among them, during the process of the large - model expanding, selecting, and refining the content information points and theme background of the initial theme, retrieving through connecting to the Internet can be adopted.
3. The documentary story writing system according to claim 1, characterized in that, The theme setting unit also includes: The content parameter extraction unit is used to extract the plot guiding factors related to the theme from the database, and extract the inspiring questions corresponding to the plot guiding factors, and use the inspiring questions as the content parameters related to the theme. The content parameters are used as prompt words during the voice conversation to guide the large - model to ask the user questions when the large - model conducts a voice conversation with the user around the theme content; Among them, in the database, the plot guiding factors of various themes and the inspiring questions transformed according to the plot guiding factors are stored. The plot guiding factors are extracted from documentary literature, and the extraction of the plot guiding factors at least includes: event type, cause of contradiction, relationship between characters, and personal feelings.
4. The documentary story writing system according to claim 1, characterized in that, It also includes: The information collection unit is used to, if the user generates a documentary story for non - first time, after guiding the user to determine the theme of the documentary story, obtain the user's basic information from the database and provide it to the large - model. The obtained user's basic information is used to avoid repeated questions to the user during the process of the large - model conducting a voice conversation with the user around the theme content. Among them, the basic information includes: gender, age, hometown, education level, and highest honor.
5. The documentary story writing system according to claim 1, characterized in that, The story generation unit includes a first generation sub - unit and a second generation sub - unit. After each voice conversation unit forms the current conversation record, a documentary story is generated through the first generation sub - unit or the second generation sub - unit, where: A first generation subunit, which is used to select a writing style after a voice conversation ends. The writing style includes at least one of the following: documentary style, romantic style, and tragic style. According to at least one set of voice conversations selected from the current conversation record and the selected writing style, the writing style has generation rules and algorithms for generating a documentary story that conforms to artistry and reflects the user's life and personality; through a synthesis key, according to the generation rules and algorithms, at least one set of voices in the current conversation record is integrated into a narrative story draft. If the user modifies it through an edit key, the modified story draft is used as the documentary story. If the user does not modify it, the story draft is used as the documentary story. The generated documentary story is in text form, voice form, or video form; A second generation subunit, which is used to select a writing style after a voice conversation ends. According to the writing style and the current conversation record, a documentary story is generated through multiple steps; according to the writing style and the current conversation record, a documentary story is generated through multiple steps, including: First, automatically generate a thread that runs through the whole text based on the current conversation record; second, enhance the artistry of the article for the current conversation record. The methods for enhancing the artistry of the article include: polishing the beginning, adding descriptions of the environment, and increasing dialogues; third, adjust the article structure by setting suspense to increase the interest of the documentary story. The ways of setting suspense include: different time jumps and character perspective restrictions; fourth, according to the generation rules and algorithms corresponding to the writing style, generate a documentary story. The generated documentary story is in text form, voice form, or video form.
6. The documentary story writing system according to claim 1, characterized in that, It also includes an analysis unit. The analysis unit is set in the documentary story writing system or in a remote server. The analysis unit is used for: Converting the current conversation into text form; Analyzing the current conversation record in text form, generating a continuity question based on the current conversation record, using the continuity question as a memory link for the opening remarks of the next voice conversation on the same topic, and saving the memory link to the database; Analyzing the current conversation record in text form, extracting basic information from the current conversation record. The basic information includes: gender, age, family members, hometown, health status, educational level, and highest honor. Comparing the extracted basic information with the basic information in the database, updating or replacing the same basic information in the database; Analyzing the current conversation record in text form, extracting core elements related to mobilizing memories and writing memoirs from the current conversation record, and saving them to the database. The core elements include: relevant dates, relevant people, memories, major life experiences, major events, and personal preferences; Analyzing the current conversation record in text form, extracting key information other than the core elements from the current conversation record, comparing the extracted key information with the key information in the database, and updating, replacing, or not processing the similar key information in the database; Among them, the key information includes: core elements, and historical conversation records with a semantic similarity higher than the similarity threshold to the current conversation record; User information includes: core elements and basic information.
7. The documentary story writing system according to claim 1, characterized in that It also includes a music generation unit, a picture generation unit, and a video generation unit, where: The music generation unit arranges the documentary story into a poem through a large model, uses the poem as a prompt to input into a music generation large model, and generates a song in combination with the music style selected by the user; The picture generation unit is used to extract the picture elements in the documentary story and generate corresponding copywriting for the picture elements, and uses the copywriting as a prompt to input into an open-source picture generation model after stylistic fine-tuning to generate pictures; The video generation unit is used to receive the photos that the user can upload through the picture import unit, and by calling the video generation API service interface, generate corresponding short videos for the photos and the generated pictures respectively, and splice them into a video according to the arrangement order of the short videos adjusted by the user, the matching copywriting set, and the music.
8. A documentary story writing device, characterized in that, The documentary story writing device may or may not have a visual interface, and the documentary story writing device includes: A processor; and a memory arranged to store executable instructions that, when executed, cause the processor to implement through the documentary story writing system according to any one of claims 1-7; The documentary story writing device further includes: A start and reset key for entering a documentary story writing mode or re-entering another documentary story writing mode; A dialogue control key. When the dialogue control key is activated, it enters the voice dialogue mode, and a documentary story is generated through the theme setting unit, the voice dialogue unit, and the story generation unit; The documentary story writing device further includes: A dialogue history list for the user to review previous conversation records; A dialogue record display unit implemented through a dialogue record display key for viewing the conversation content; A story list for viewing all generated documentary stories. By clicking, it enters the story display page for editing and sharing; A story generation unit that can generate a documentary story and can also edit the documentary story; A story square for browsing documentary stories shared by other users, with like and comment keys; Memory cards for viewing one's own story memory cards on the page.
Citation Information
Patent Citations
An interview intelligent robot device and an intelligent interview method for automatically generating an interview draft
CN109918650A
Digital human real-time interaction and generative artificial intelligence-based memoir generation system
CN118246428A