Method and device for duplicating real figure based on large model, electronic equipment and computer readable medium

By using a large-scale model to replicate real people, and combining multimodal data to generate training corpora and construct large-scale character models, the problem of accurately replicating identity, knowledge boundaries, and speaking style was solved, achieving a realistic interactive experience.

CN120930771APending Publication Date: 2025-11-11BEIJING ZERO-1000 TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510749542.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately recreate the identity, knowledge boundaries, and speaking style of real people, resulting in a lack of authenticity and reliability in interactive experiences.

Method used

By constructing a large-scale model-based replication method, input information is received and it is determined whether knowledge base retrieval is needed. Enhanced prompts or direct responses are generated. Training corpora for specific characters are generated by combining multimodal data, including identity features, knowledge boundaries, and speaking styles, to construct a large-scale character model.

Benefits of technology

It accurately replicates the identity, knowledge, and speaking style of real people, providing a more realistic and reliable interactive experience and avoiding the Turing test.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930771A_ABST
    Figure CN120930771A_ABST
Patent Text Reader

Abstract

The invention relates to a method and device for duplicating a real character based on a large model, electronic equipment and a computer readable medium. The method for duplicating the real figure based on the large model comprises the following steps: receiving input information; determining whether knowledge base retrieval is needed or not according to the content of the input information; when it is determined that knowledge base retrieval needs to be carried out, an enhancement prompt is formed based on the input information and the content retrieved from the knowledge base, and the enhancement prompt is input into the role large model to generate a response; and when it is determined that knowledge base retrieval does not need to be performed, directly inputting the input information into the role large model to generate a response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of artificial intelligence and digital human technology, and in particular to a method, apparatus, electronic device, and computer-readable medium for replicating real people based on a large model. Background Technology

[0002] In real life, there are various scenarios where there is a potential demand for replicating real people. On the one hand, people often have deep feelings for deceased relatives and friends and yearn to chat and keep them company to alleviate their longing. While traditional chatbots can provide basic conversational functions, they cannot offer an interactive experience with the characteristics of a real person, and thus cannot truly satisfy users' desire for emotional companionship. On the other hand, in the field of commercial live streaming, high-traffic IPs such as influential figures in e-commerce and knowledge bloggers, due to their limited energy, find it difficult to meet the demands of long live streams to fully interact with the audience, promote products, or share knowledge, which to some extent limits the full realization of their commercial value. While existing pre-recorded video solutions can provide some content, they lack interactivity and cannot interact with the audience in real time, making it difficult to meet the diverse needs of viewers.

[0003] In related technologies, dialogue systems based on general large models can be applied to scenarios such as virtual companionship, live-streaming e-commerce, and knowledge services. Although dialogue systems based on general large models have a certain degree of interactivity, they face difficulties in maintaining consistency in human characteristics, making it hard for users to feel like they are interacting with a real person.

[0004] In summary, existing technologies are insufficient in meeting users' needs for character replication. Therefore, it is necessary to propose a method for replicating real-life characters based on large models to provide users with a more realistic and reliable interactive experience. Summary of the Invention

[0005] This disclosure relates to an apparatus and method for digitally replicating real people, applicable to scenarios such as virtual companionship, live-streaming e-commerce, and knowledge services, which can solve one or more of the aforementioned technical problems.

[0006] This disclosure provides a method for replicating real people based on large models, which can provide users with a more realistic and reliable interactive experience.

[0007] This disclosure also provides a device based on a large model that replicates real people, enabling users to have a more realistic and reliable interactive experience.

[0008] The first aspect of this disclosure provides a method for replicating real-life figures based on a large character model, comprising: receiving input information; determining whether a knowledge base retrieval is required based on the content of the input information; when it is determined that a knowledge base retrieval is required, forming an enhanced prompt based on the input information and the content retrieved from the knowledge base, and inputting the enhanced prompt into the character large character model to generate a response; and when it is determined that a knowledge base retrieval is not required, directly inputting the input information into the character large character model to generate a response.

[0009] In an exemplary embodiment, when it is determined that a knowledge base retrieval is required, a corresponding query vector is generated based on the input information, and multiple relevant corpus segments are retrieved from the knowledge base using the query vector; the retrieved corpus segments are concatenated with the input information to form an enhanced prompt; the enhanced prompt is input into the role model; and the role model generates a response based on the retrieved corpus segments.

[0010] In an exemplary embodiment, the method for replicating real-life figures based on a large model further includes: collecting multimodal data including video, audio, and text data; generating training corpus for a specific figure based on the multimodal data; constructing a knowledge base for the specific figure based on the multimodal data; and inputting the training corpus into a base model for fine-tuning to obtain a large model of the figure.

[0011] In an exemplary embodiment, generating training corpus for a specific person based on multimodal data includes: constructing question-answer pairs related to identity features based on text data.

[0012] In an exemplary embodiment, generating training corpus for a specific person based on multimodal data further includes: extracting question-and-answer pairs from text data that can express the viewpoint of the specific person, and pre-setting unknown domains and rejection templates for the specific person.

[0013] In an exemplary embodiment, generating training corpus for a specific person based on multimodal data further includes: constructing question-and-answer pairs that simulate the annoyance of a specific person.

[0014] In an exemplary embodiment, generating training corpus for a specific person based on multimodal data further includes constructing realistic multi-turn dialogue data by labeling language style elements that reflect the speaking style of the specific person.

[0015] A second aspect of this disclosure provides an apparatus for replicating real-life figures based on a large model, comprising: an input information receiving module configured to receive input information; a knowledge base retrieval determination module configured to determine whether a knowledge base retrieval is required based on the content of the input information; a retrieval enhancement generation execution module configured to, when it is determined that a knowledge base retrieval is required, generate an enhancement prompt based on the input information and the content retrieved from the knowledge base, and input the enhancement prompt into the large model of the character to generate a response; and a direct generation module configured to, when it is determined that a knowledge base retrieval is not required, directly input the input information into the large model of the character to generate a response.

[0016] A third aspect of this disclosure provides an electronic device including a processor and a memory, the memory being used to store a program that, when executed by the processor, performs the method described above.

[0017] A fourth aspect of this disclosure provides a computer-readable medium having a program stored thereon that, when executed by a processor, performs the method described above.

[0018] However, the aspects of this disclosure are not limited to those set forth herein. These and other aspects of the disclosure will become apparent to those skilled in the art upon reference to the detailed description of the disclosure given below.

[0019] The embodiments of this disclosure, through technologies such as identity feature construction, knowledge boundary control, and style transfer, aim to solve problems such as style distortion and knowledge overreach in the process of character replication, thereby replicating a real person who possesses that person's identity, knowledge, and speaking style, while also avoiding the Turing test and providing users with a more realistic and reliable interactive experience. The above-mentioned technical solutions of this disclosure only need to achieve one of the aforementioned effects; it is not required that every technical solution achieve all of the aforementioned technical effects.

[0020] Furthermore, the effects of this disclosure include not only those set forth herein, but also other effects that will be apparent to those skilled in the art upon reference to the claims, the specification, and the accompanying drawings. Attached Figure Description

[0021] These and / or other aspects will become apparent and more readily understood from the following description of embodiments taken in conjunction with the accompanying drawings, in which:

[0022] Figure 1 This is a flowchart of a method for replicating real people based on a large model according to an embodiment of the present disclosure;

[0023] Figure 2 This is a block diagram of a system for replicating real people based on a large model, according to an embodiment of the present disclosure;

[0024] Figure 3 This is a block diagram of an apparatus for replicating real people based on a large model, according to an embodiment of the present disclosure;

[0025] Figure 4 A block diagram of an electronic device according to embodiments of the present disclosure; and

[0026] Figure 5 This is a block diagram of a computer-readable medium according to embodiments of the present disclosure. Detailed Implementation

[0027] This disclosure will now be described more fully below with reference to the accompanying drawings, in which embodiments of the disclosure are illustrated. However, this disclosure may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be more detailed and thorough, and will fully convey the scope of this disclosure to those skilled in the art.

[0028] It will also be understood that throughout the specification, the same reference numerals denote the same parts. Each of the features of the various embodiments of this disclosure can be combined partially or entirely with each other, and various technical associations and drives are possible. Each embodiment can be implemented independently of each other or can be implemented together in association.

[0029] Unless otherwise defined or implied herein, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It will be further understood that terms such as those defined in common dictionaries shall be interpreted as having a meaning consistent with their meaning in the context of the relevant field and in this disclosure, and shall not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0030] As described in the background section, there are various scenarios in real life where there is a potential demand for replicating real people. However, existing technologies have shortcomings in achieving this.

[0031] The inventors of this application have discovered that, from the perspective of personal characteristics, real-life individuals possess multi-dimensional and unique attributes. First, real-life individuals have a unique identity, encompassing many aspects from birth to adulthood, such as family background, educational experience, and job search process. These elements collectively constitute a person's growth trajectory and social network. Second, real-life individuals possess their own knowledge reserves, including both general knowledge and professional field knowledge. When constructing a replica model of a real-life individual, it should be ensured that their knowledge content is as consistent with reality as possible to avoid misinformation. Furthermore, real-life individuals have clear knowledge boundaries; regarding areas they are unfamiliar with, they should demonstrate ignorance rather than fabricating answers. In addition, real-life individuals have unique speaking styles, reflected in multiple dimensions such as tone of voice, pauses, and emotional expression.

[0032] The existing technology falls short in meeting users' needs for character replication for several reasons: First, the identity features are incompletely constructed, making it difficult to accurately recreate core identity elements such as a person's upbringing and social relationships, resulting in a lack of authenticity and completeness in the replicated character's background. Second, the knowledge system lacks boundary control, often resulting in "illusory answers" when answering user questions—responding to answers beyond the character's knowledge or inconsistent with their identity, affecting the reliability of the replicated character. Third, the imitation of language style is insufficient, failing to reproduce the tone, habits, and emotional expression of a real person, making the replicated character lack authenticity in communication.

[0033] Therefore, a method for replicating real people based on a large model is needed to address one or more of the aforementioned shortcomings, so that the replicated real person possesses the characteristics of the original person, such as identity, knowledge, and speaking style, while also avoiding the Turing test and providing users with a more realistic and reliable interactive experience.

[0034] It should be noted that the data collection methods involved in the following embodiments all comply with relevant laws and regulations and data privacy protection requirements.

[0035] In the following description, embodiments will be illustrated with reference to the accompanying drawings.

[0036] The first aspect of this disclosure provides a method for replicating real people based on large models. Figure 1 This is a flowchart of a method for replicating real people based on a large model according to embodiments of this disclosure. Figure 1As shown, the method for replicating a real person based on a large model according to an embodiment of the present disclosure includes: receiving input information (S101); determining whether a knowledge base retrieval is needed based on the content of the input information (S102); when it is determined that a knowledge base retrieval is needed, forming an enhanced prompt based on the input information and the content retrieved from the knowledge base, and inputting the enhanced prompt into the large model of the character to generate a response (S103); and when it is determined that a knowledge base retrieval is not needed, directly inputting the input information into the large model of the character to generate a response (S104).

[0037] In step S101, user input is received. The input can be a natural language query from the user. The input can be in the form of text and / or voice. For example, the user can voice input: "How much did you sell during last year's Singles' Day?"

[0038] In step S102, it is determined whether a knowledge base retrieval is needed based on the content of the input information. For example, a fine-tuned role model or awareness model can be used to predict the user's intent to determine whether a knowledge base retrieval is needed. These models can be trained using historical data and user behavior to improve prediction accuracy. When determining whether a knowledge base retrieval or access to Retrieval Augmentation (RAG) is needed, the fine-tuned role model or awareness model can utilize the general capabilities of the base model to determine whether to trigger a retrieval. The retrieval results can use the top five most relevant fragments recalled by FAISS. FAISS (Facebook AI Similarity Search) is an efficient similarity search and dense vector clustering library specifically designed for fast retrieval of large-scale vector data.

[0039] For example, when a user enters "What do you think of AI?", the user's input is a query for opinion rather than fact, so it does not need to be searched; however, when a user enters "What is your mother's name?", since accurate identity facts are required, it is necessary to search for this information (RAG calls data stored in the knowledge base from sources such as Wikipedia, biographies, and family interview records).

[0040] However, this disclosure is not limited to this. For example, a combination of multiple technologies, such as rule engines and entity recognition, can be used to determine whether a knowledge base retrieval is necessary. For instance, when user input contains certain keywords or matches a certain pattern, the rule engine can determine whether a knowledge base retrieval is required.

[0041] Entity recognition technology is used to identify specific entities from user input, such as names of people, places, and products. This entity information helps the system more accurately understand the user's needs, thereby determining whether a knowledge base search is necessary.

[0042] By defining the knowledge domains that a model can answer and the domains it cannot answer, we can avoid the model guessing or providing incorrect information in areas it doesn't understand. For example, when a user's question exceeds the model's knowledge boundaries, we can determine not to trigger a retrieval and instead have the model output a rejection template. For instance, when a user enters, "What medicine should I take for a cold?", we can avoid a retrieval and instead have the large model output a rejection template based on its knowledge boundaries.

[0043] In step S103, when it is determined that a knowledge base retrieval is needed, an enhanced prompt is generated based on the input information and the content retrieved from the knowledge base. This enhanced prompt is then input into the role-based large model to generate a response. Once a knowledge base retrieval is determined, the system searches the knowledge base based on the key information input by the user (such as keywords, topics, etc.). The knowledge base can be a structured database storing a large amount of professional knowledge and information. The system uses appropriate retrieval algorithms and techniques to retrieve content related to the user's input from the knowledge base. The system integrates (e.g., concatenates) the content retrieved from the knowledge base with the user's original input information to generate an enhanced prompt. The enhanced prompt includes the user's original question, contextual information, and relevant knowledge retrieved from the knowledge base. By generating enhanced prompts, the system can provide richer and more accurate input information for the subsequent role-based large model, thereby improving the accuracy and quality of response generation. Finally, the system inputs the enhanced prompt into the role-based large model. The role-based large model is a trained natural language processing model that can understand the content of the input information and generate corresponding responses based on this information. When generating a response, the big model comprehensively considers the user's original question, contextual information, and relevant knowledge retrieved from the knowledge base to produce an answer that is both user-relevant and professionally accurate. This approach reduces the illusion problem inherent in current big models.

[0044] In step S104, when it is determined that a knowledge base retrieval is not required, the input information is directly input into the role model to generate a response. Conversely, in step S103, when processing user input, if the system determines that the user's input question does not require retrieving specific information from the knowledge base, then the system will decide that a knowledge base retrieval is unnecessary. Once it is determined that a knowledge base retrieval is not required, the system will directly input the user's original input information into the role model. The role model will then directly generate a response based on this information, without requiring an additional knowledge base retrieval process. This approach simplifies the system process and improves response speed, especially when handling simple, general questions.

[0045] According to an exemplary embodiment, when it is determined that a knowledge base retrieval is required, a corresponding query vector is generated based on the input information, and multiple relevant corpus segments are retrieved from the knowledge base using the query vector; the retrieved corpus segments are concatenated with the input information to form an enhanced prompt; the enhanced prompt is input into the role model; and the role model generates a response based on the retrieved corpus segments.

[0046] The bge-large-zh model is used to vectorize user queries. The bge-large-zh model is a large pre-trained model specifically trained for Chinese natural language processing tasks. When a user inputs a natural language query, the bge-large-zh model analyzes and processes the input information, converting it into a fixed-dimensional vector. Each value in this vector represents a feature of the input information in a specific semantic dimension, and together they constitute the representation of the input information in vector space.

[0047] Numerous document fragments can be contained in a vector database (e.g., the FAISS vector database), and these fragments can contain information relevant to the user's query. During similarity retrieval, the system inputs the user's query vector into the FAISS vector database and uses the algorithms and indexing structure provided by FAISS to perform a nearest neighbor search. The goal of the nearest neighbor search is to find the vectors that are closest to the query vector in the vector space; the document fragments corresponding to these vectors are the most relevant documents to the user's query. Ultimately, the system returns the Top-5 most relevant document fragments, sorted from highest to lowest similarity to the query vector, returning the five most similar document fragments.

[0048] Directly inputting user queries into the role model to generate answers may result in inaccurate responses due to limited query information or the illusion problem of the large model. Context enhancement aims to provide the role model with more contextual information by concatenating relevant document fragments retrieved from the FAISS vector database with the user query. This allows the model to better understand the context and needs of the user query, thereby generating higher-quality answers based on the retrieved document fragments. The system concatenates the top-5 most relevant document fragments retrieved from the FAISS vector database with the user's original query text (i.e., input information). The concatenation method can be designed according to specific needs; for example, query text, system instructions, and document fragments can be combined in a certain format to form a longer text sequence. This concatenated text sequence constitutes the enhanced Prompt, which contains the user query and related contextual information, providing the role model with richer context and helping the model generate answers that better meet user needs. The enhanced Prompt is then input into the role model, where the model understands and analyzes the input text, generating corresponding answers based on learned language knowledge and semantic relationships. During the answer generation process, the model comprehensively considers information from the user's query and relevant document fragments retrieved, ensuring that the generated answer is closely related to the user's needs and possesses high accuracy and professionalism. Furthermore, post-processing techniques (such as removing duplicate content) can be used to further optimize the generated answer and improve its quality.

[0049] Figure 2 This is a block diagram of a system for replicating real people based on a large model, according to an embodiment of the present disclosure.

[0050] The system for replicating real people based on a large model according to embodiments of the present disclosure may include a data collection unit, a data storage unit, a data processing unit, a training corpus generation unit, a knowledge base construction unit, and a large character model. However, the system for replicating real people based on a large model according to embodiments of the present disclosure is not limited to this and may include more or fewer components than those described above.

[0051] The Data Collection Department is dedicated to collecting data from a wide range of sources to build rich and diverse datasets. These sources can include Wikipedia, public speeches, interviews, live videos, relevant books, and websites. Data collection methods can include web scraping, scanning printed books, capturing live videos using recording tools, and manually organizing data from specific sources. For example, web scraping technology is used to automatically scrape data from various websites. For publicly available online data (such as encyclopedia websites), the targeted crawling module of the Scrapy framework is configured, setting a delay strategy with a request interval of ≥2 seconds to automatically extract the main text of web pages and filter out advertising scripts. Scrapy is a powerful Python web scraping framework; by configuring its targeted crawling module, the web pages to be scraped and the data fields to be extracted can be precisely specified, improving the efficiency and accuracy of the crawling process.

[0052] For printed documents, they are converted to electronic format and stored using professional scanners. For example, PDF files are generated using a professional scanner, and ISBN metadata tags are attached. ISBN (International Standard Book Number) is an important identifier for books. Attaching ISBN metadata tags to scanned PDF files facilitates book classification, retrieval, and management, improving data organization and accessibility. For data from special sources, such as internal documents obtained through specific channels or exclusive interview transcripts, which may have issues like non-standard formatting and fragmented information, manual sorting and classification are necessary to ensure data quality and usability.

[0053] Recording tools can be used to capture live video streams in real time and save them as video files. For video / live data, the FFmpeg toolkit is used to capture RTMP streaming media. FFmpeg is an open-source multimedia processing toolkit that supports encoding, decoding, and transcoding of various audio and video formats. When capturing video / live data, appropriate video frame rate (video_fps) and audio sample rate (audio_sample_rate) parameters are set to ensure that the acquired data has high quality and meets the needs of subsequent analysis and processing. For example, configuring parameters such as video_fps=25 and audio_sample_rate=44100Hz ensures the quality of the original data.

[0054] The data collected through the aforementioned methods encompasses various formats, including text, video, and audio. The data storage department categorizes and stores the collected data (video, audio, text, etc.) using a three-tiered storage directory: raw data layer, intermediate data layer, and structured data layer. The raw data layer is the first level of data storage, primarily used to store unprocessed raw data. This data is obtained directly from the data collection source, retaining its original format and content, including unprocessed video, audio, and scanned documents. The intermediate data layer, located between the raw data layer and the structured data layer, stores data that has undergone some preliminary processing but is not yet fully structured; for example, storing transcoded MP3 audio and preliminary OCR results. The structured data layer is the highest level of data storage, used to store data that has undergone final cleaning, processing, and structuring. This hierarchical storage facilitates the flow and traceability of data between different processing stages, improving the controllability and maintainability of the entire data processing workflow.

[0055] The data processing department performs data processing. First, audio is extracted from the video data. Professional video processing tools or programming libraries can be used to achieve audio extraction, such as FFmpeg. The audio data is transcribed into text data using ASR (Automatic Speech Recognition) technology, or the open-source tool Whisper v3 (Whisper v3 is a speech recognition model open-sourced by OpenAI, with high recognition accuracy and strong language adaptability) can be used for speech transcription, and finally the speech is converted into the corresponding text. Automatic error correction is performed on the transcribed text. For example, the GPT-4o mini model is used to correct typos, homophones, reduplicated characters, punctuation and other errors. The GPT-4o mini model has powerful language understanding and generation capabilities. It can identify and correct errors in the text based on context information. For example, when there are incorrect uses of "的", "地", "得" in the text, the model can correct them to the correct usage according to the grammatical structure and semantic information of the sentence. The scanned PDF data usually exists in the form of images. In order to convert it into an editable text format, some open-source libraries need to be used for parsing. PyPDF2 is a commonly used Python open-source library. It can read information such as text and images in PDF files and extract them into text format data. During the PDF transcription process, problems such as misaligned chapter titles and mixed footnote texts may occur. These problems need to be corrected manually. Manually, according to the original layout and content logic of the PDF, the transcribed text can be adjusted and corrected, and rules can be used to structure information such as titles and chapters. The web page data is parsed based on the BeautifulSoup library. BeautifulSoup is a Python library that can easily extract data from HTML or XML documents. Ads, irrelevant hypertext tags, sensitive information, etc. are removed through rules. Manual verification can correct problems such as typos, homophones, proper nouns, reduplicated characters, etc. generated during the ASR process, and at the same time can also repair problems such as misaligned chapter titles and mixed footnote texts during the PDF transcription process. Manual verification can discover errors that are difficult to find in the automatic processing process and ensure that the final obtained text data is accurate and reliable.

[0056] The training corpus generation department plays a crucial role in the entire large-scale model application development process, generating rich, diverse, and targeted dialogue corpora from multiple dimensions. Based on data processed by the data processing department, the training corpus generation department can generate specific character-related corpora, knowledge boundary corpora, emotion corpora, and stylized corpora. Identity-related corpora refer to dialogue content that reflects a specific character's identity (e.g., name, gender, date of birth, place of birth, family members, social relationships, etc.). Knowledge boundary corpora clarify the knowledge scope and limitations of a specific character in a particular domain. It helps the model understand which knowledge a specific character possesses and which is unknown, thus enabling it to more accurately grasp the knowledge boundaries of a specific character when answering questions or engaging in dialogue, avoiding inaccurate or out-of-knowledge responses. Stylized corpora imbue the dialogue generated by the model with a specific style, such as using interjections and reduplicated words to reflect speaking style. Emotion corpora are collections of text data used to train the large-scale model to understand and generate text with specific emotions, helping the model output responses that simulate a character's annoyance when faced with repetitive questions. The generated identity feature corpus, knowledge boundary corpus, emotion corpus, and stylistic corpus are input into the base model, and the parameters of the base model are fine-tuned. Through fine-tuning, the base model can learn specific information contained in the generated corpus, such as character identity features, knowledge boundaries, emotional expression, and stylistic characteristics, thereby adjusting its own parameters to generate more suitable dialogue output. Finally, the fine-tuned base model will transform into a character-based large model, possessing specific character identity features, knowledge scope, emotional expression ability, and stylistic characteristics, enabling it to provide users with a more realistic and reliable interactive experience in the corresponding domain or scenario.

[0057] The main task of the Knowledge Base Construction Department is to build a dedicated knowledge base using content such as books and web pages. This knowledge base will serve as the retrieval system for the large-scale model RAG. The web page and book data used here are highly relevant to the person being replicated. This means that this data contains knowledge, information, experiences, etc., related to a specific person, and the knowledge base is primarily used to answer questions related to that person. For example, if the goal is to replicate a historical figure, the web page data in the knowledge base could include biographical websites and academic research articles about that historical figure.

[0058] The knowledge base construction department employs a conventional vector indexing approach. First, webpage data is cleaned, removing advertisements and irrelevant hypertext tags. Books are then parsed using a chapter-based structure. Next, semantic segments are divided (e.g., each segment ≤ 512 terms), preserving contextual relevance. For longer documents, such as entire books or lengthy webpages, directly segmenting by a fixed length might lead to information fragmentation and disrupt contextual coherence. Therefore, a sliding window overlapping segmentation method is used. For example, long documents are segmented using a sliding window overlapping method. The overlap rate is generally controlled between 15% and 20%, ensuring some information overlap between adjacent segments and better restoring the semantics of the original text during retrieval. Finally, the bge-large-zh model is chosen as the embedding model. The embedding model converts text data into a high-dimensional vector representation, making semantically similar texts closer together in the vector space. FAISS is used as the indexing framework to construct the high-dimensional vector index. By converting text into vectors and storing them in the FAISS index, the most similar text vector to the query vector can be quickly found during retrieval, thereby obtaining relevant knowledge. The Knowledge Base Construction Department is able to build a high-quality knowledge base that is highly relevant to the replicated characters, providing strong support for the large model RAG system, thereby improving the accuracy of the large model in answering relevant task questions.

[0059] Yi-34B is a large-scale language model. Pre-trained on a large amount of general-purpose text data, Yi-34B possesses powerful language understanding and generation capabilities, enabling it to handle various natural language tasks such as text generation, question answering, and dialogue. Using Yi-34B as the base model, a small amount of general-purpose instructions and the generated corpus are mixed, and system prompts are used to differentiate between them. The final corpus is then fed into the model for fine-tuning. By adding specific system prompts to the training data, the model can clearly determine whether the currently processed corpus belongs to a general task or a dialogue of a specific character. Using Yi-34B as the base model, the generated dialogue data is fine-tuned to train a large-scale character model. After fine-tuning and training, the large-scale character model possesses the specific character's language style, knowledge background, emotional expression, and behavioral habits. It can generate dialogue content that conforms to the character's settings based on different inputs.

[0060] According to an exemplary embodiment, the method for replicating real-life figures based on a large model further includes: collecting multimodal data including video, audio, and text data; generating training corpus for a specific figure based on the multimodal data; constructing a knowledge base for the specific figure based on the multimodal data; and inputting the training corpus into a base model for fine-tuning to obtain a large model of the figure. Due to the above... Figure 2 The data collection, data processing, training corpus generation, and knowledge base construction have been described in detail, so repeated descriptions of these aspects will be omitted here.

[0061] According to an exemplary embodiment, generating training corpus for a specific person based on multimodal data includes: constructing question-and-answer pairs related to identity features based on text data. This part of the corpus is mainly used to replicate the person's identity information. Using Wikipedia, biographies, etc., identity-related questions and answers are constructed, including basic identity information such as gender, name, date of birth, and place of birth; social relationship information such as family members, educational background, and work relationships; and achievement and experience information such as works, honors, and major events. The constructed question-and-answer pairs revolve around information closely related to the person's identity features. By constructing such question-and-answer pairs, the system learns the identity information of a specific person during training, so that in subsequent interactions with users, it can accurately respond in the identity of that person. See Dialogue 1 for an example. Dialogue 1 is as follows:

[0062] Human: Hello, who are you?

[0063] Assistant: Hello, I am Kai-Fu Lee, the founder of Zero One Things.

[0064] Human: Can you introduce your family?

[0065] Assistant: My father's name is Li Tianmin, my mother's name is Wang Yaqing, I have an older sister named Li Kaimin, my wife is Xie Xianling, and we have two daughters, the eldest daughter is named Li Dening and the younger daughter is named Li Deting.

[0066] In generating training corpora for a specific person, in addition to constructing question-and-answer pairs related to identity features based on text data, other operations are required. According to an exemplary embodiment, generating training corpora for a specific person based on multimodal data further includes: extracting question-and-answer pairs from text data that can express the specific person's viewpoint, and pre-setting unknown domains and rejection templates for the specific person.

[0067] In public speeches and live videos, specific individuals express their opinions, thoughts, and insights. These texts are valuable resources for acquiring these viewpoints. Utilizing the texts from public speeches and live videos, a large-scale model can summarize and extract question-and-answer pairs that express these viewpoints. This corpus primarily supplements the system's knowledge in the individual's field. By extracting these viewpoints, the system can learn the individual's insights and thought processes within their professional or familiar areas. This allows the system to provide valuable perspectives and answers that align with the individual's identity when communicating with users, enhancing the system's simulation of the individual's professional knowledge. For an example of question-and-answer pairs that express a person's viewpoint, please refer to Dialogue 2.

[0068] Dialogue 2 is as follows:

[0069] Humanity: How to plan a business model for everything?

[0070] Assistant: Regarding our business model, we adopt the following approach. We believe that in this new era, Super Apps represent the biggest business opportunity. How do we position ourselves for this opportunity? We need to consider several things…

[0071] The system pre-defines the knowledge boundaries of the user, including areas the user is unfamiliar with and response templates for refusal. Specifically, based on the user's background, it manually identifies areas that don't match their background and constructs chat dialogues within those areas, where the responses are presented as refusal. For example, if the user is a technology expert, fields like medicine and art history might be unfamiliar. To obtain a broader range of domain boundaries, this background knowledge and chat dialogues can be input into a large-scale model, allowing the model to generate more chat dialogues in other unfamiliar areas. The model can learn refusal patterns and methods from existing examples, thus generating more appropriate question-and-answer pairs. This part of the corpus primarily addresses the user's knowledge boundary issue. By pre-setting refusal templates and generating relevant question-and-answer pairs, the system can respond appropriately to these questions from the user's perspective, avoiding incorrect or inaccurate answers, preventing role overstepping, establishing professional trust, and preventing misleading in high-risk areas like medicine / finance, ultimately improving the system's credibility and user experience. Example of a refusal-based question-and-answer format is shown in Dialogue 3. Dialogue 3 is as follows:

[0072] Humans: How to treat cold hands and feet?

[0073] Assistant: I'm sorry, I don't know anything about this. I suggest you consult a doctor at the hospital for relevant advice.

[0074] In addition to generating corpora based on identity characteristics, viewpoints, and knowledge boundaries, it's also necessary to consider simulating the character's emotions and psychological state to make the replicated character more three-dimensional and realistic. In the dialogue, initially, the character will answer questions politely and normally; as the number of repeated questions increases, the answers will gradually show impatience, such as using suggestive phrases to indicate that the question has already been answered, or even directly expressing the desire for the user not to repeat the question. According to an exemplary embodiment, generating training corpora for specific characters based on multimodal data also includes constructing question-and-answer pairs simulating the character's annoyance. Repetitive questions and answers are manually constructed to simulate the character's annoyance, and these are used as examples for the large model to imitate and generate more similar dialogues. This part of the corpus is mainly for passing the Turing test, making the replicated character more realistic. When the system simulates a specific character, if it can only provide monotonous, emotionless answers to user questions, it's difficult for testers to believe it's a real person. By constructing question-and-answer pairs that simulate a person's feelings of annoyance, the system can exhibit annoyance like a real person when faced with repetitive questions or similar situations, making its responses more in line with human emotions and reactions, thus making the replicated character more vivid and realistic. See Dialogue 4 for an example of simulating a person's feelings of annoyance. Dialogue 4 is as follows:

[0075] Human: When was your birth date?

[0076] Assistant: I was born on December 3, 1961.

[0077] Human: So what's your birthday?

[0078] Assistant: I think I just mentioned it, it was December 3, 1961.

[0079] Human: When were you born?

[0080] Assistant: I'm sorry, I've already answered this question before. It was December 3, 1961. Please don't ask it again, okay?

[0081] In Dialogue 4, the gradually escalating responses of boredom vividly simulate the character's true psychological state when faced with repeated questions, making the replicated character more realistic.

[0082] To make the language output by the model more closely resemble that of a real person, it is necessary to construct realistic multi-turn dialogue data using interview video data and chat data, focusing on simulating the person's speaking style. Since the model may have insufficient imitation of language style, elements reflecting speaking style are manually annotated, and the annotated data is used for model fine-tuning to directly address this deficiency and restore the person's language habits. According to an exemplary embodiment, generating training corpora for a specific person based on multimodal data also includes: constructing realistic multi-turn dialogue data by annotating language style elements that reflect the specific person's speaking style.

[0083] By constructing realistic multi-turn dialogue data using interview video data and chat data, this data needs to simulate the speaker's speaking style. Therefore, during manual annotation, special attention should be paid to interjections and reduplicated words that reflect speaking style. For example, manual annotation can include the following language style elements: interjections (such as "la," "ya"), reduplicated words, and emotional markers (such as "(laughing)," "(sighing)"). The "reduplicated words" mentioned in this paper do not refer to the grammatical concept of reduplicated words (such as "red," "green"), but rather to the repetition of words that occurs when faced with impromptu speaking situations (such as being suddenly asked to express an opinion), due to the incoherence of thought and expression. For example, in the sentence, "This, this I need to think about it," the word "this" is used repeatedly. This repetition more realistically reflects the characteristics of spoken expression when a person is not fully prepared. This expression style with word repetition and even sentence incoherence is more consistent with real-life conversations in daily life. Incorporating reduplicated words into the style modeling system allows the model to better simulate the speaker's thought process, enhances the realism of the dialogue, and makes the character more vivid and three-dimensional. Furthermore, incorporating non-semantic symbols such as sentiment markers into the style modeling system allows for a more comprehensive capture of a character's linguistic style characteristics, enabling the model to better simulate emotional expression and tone changes. Annotated dialogue data is used as a fine-tuning training set to further train the already trained model. Through fine-tuning, the model can learn the linguistic style characteristics reflected in the annotated data, including the usage and frequency of elements such as interjections, reduplicated words, and sentiment markers, thereby improving the model's reproducibility of a character's linguistic style and enabling a more realistic simulation of a specific character's linguistic style.

[0084] A second aspect of this disclosure provides an apparatus 300 for replicating real-life figures based on a large model, comprising: an input information receiving module 301 configured to receive input information; a knowledge base retrieval determination module 302 configured to determine whether a knowledge base retrieval is required based on the content of the input information; a retrieval enhancement generation execution module 303 configured to, when it is determined that a knowledge base retrieval is required, generate an enhanced prompt based on the input information and the content retrieved from the knowledge base, and input the enhanced prompt into the large model of the character to generate a response; and a direct generation module 304 configured to, when it is determined that a knowledge base retrieval is not required, directly input the input information into the large model of the character to generate a response.

[0085] In the description of the first aspect of this disclosure, the various method steps and device components involved in the technical solution of this disclosure have been described in detail. Therefore, the above description can be applied to a device 300 for replicating real people based on a large model, which is a second aspect of this disclosure. Accordingly, the description will not be repeated here.

[0086] The following provides specific application examples of the technical solutions disclosed herein.

[0087] For example, replicating e-commerce livestreamers. First, 500 hours of livestream recordings from e-commerce livestreamers are collected, covering categories such as beauty, fashion, and lifestyle. Training corpora covering knowledge boundaries of categories like beauty, styling, and fashion trends are generated based on these video recordings. Their signature phrases (e.g., "Oh my god, buy it!"), interjections ("la," "ya"), reduplicated words, and emotional expressions (e.g., "(laughing)," "(sighing)") are extracted to form stylized training corpora and emotional corpora. Additionally, training corpora based on customer-provided information and other online information are generated to create identity-related training corpora. These training corpora are then input into a large model to learn the speaking style of e-commerce livestreamers, including exaggerated tones and signature catchphrases, ensuring that the replicated livestreamer can naturally imitate their style when recommending products. After this training, when users ask questions like "Which lipstick is suitable for yellow skin?" or "How to match work makeup?", the e-commerce livestreamer directly recommends products and provides styling suggestions. When users ask sensitive questions, the e-commerce livestreamer directly responds, "I don't know the answer to that question; I suggest consulting a professional," avoiding vague or misleading answers.

[0088] For example, replicating deceased relatives. Before replicating a deceased relative, explicit authorization from the family must be obtained to ensure that the use of private data such as letters, diaries, photos, and audio complies with legal and ethical requirements. Collected data is categorized by type, including text (letters, diaries, social media records), audio (voice messages, interview recordings), and images (photos, videos). Letters and diaries undergo OCR recognition (if printed), noise removal (such as blurred text and irrelevant symbols) is performed, and formatting is standardized (such as sentence segmentation and paragraphing). Speech recognition technology is used to convert audio into text, annotating speaker, time, scene, and other information. Text is extracted from photos (such as handwritten notes) or relationships are labeled (such as "photo with father"). Based on this data, training data is generated, including corpora of the deceased relative's identity characteristics, knowledge boundaries, emotions, and stylized expressions. In particular, dialect audio of the deceased relative is collected (such as hometown dialect and accent features), and the dialect type (such as Cantonese and Sichuanese) and pronunciation characteristics (such as "retroflex endings" and "heavy nasal sounds") are labeled. Using dialect speech synthesis technology, text is converted into dialect speech, and parameters such as speech rate, intonation, and pauses are adjusted to more closely resemble the speaking style of a deceased relative. Based on data such as letters and diaries of deceased relatives, their knowledge gaps (e.g., "unfamiliar with technological products," "uninterested in political news") are manually labeled, and an "I don't know" response template is set. The above training data is input into a large-scale model to generate a large-scale model of the relative's role. This technical solution successfully replicates the language style and knowledge characteristics of deceased relatives through data processing, dialect simulation, and knowledge boundary management, while ensuring content compliance and emotional authenticity through the "I don't know" response template.

[0089] This disclosure also provides an electronic device including a memory and a processor. The memory stores a program, and the processor is configured to acquire the program and, when executing the program, execute the above-described method for replicating real people based on a large model.

[0090] Figure 4 This is a block diagram of an electronic device implementing a method for replicating real people based on a large model, according to some embodiments of this disclosure. Figure 4 As shown, the method for replicating real people based on a large model in the above embodiments can be achieved through... Figure 4 The electronic device shown is used to implement this, and the electronic device includes at least one processor, memory, and at least one I / O interface.

[0091] The processor can be a general-purpose central processing unit (CPU) and a graphics processing unit (GPU), or an application-specific integrated circuit (ASIC). Memory can include at least one of volatile memory and non-volatile memory. Memory can be read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions; it can be random access memory (RAM) or other types of dynamic storage devices capable of storing information and instructions; it can also be electrically erasable programmable read-only memory (EEPROM), read-only optical disc (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disks or other magnetic storage devices, or any other medium capable of carrying or storing a desired program having an instruction or data structure form and accessible by a computer, but is not limited thereto. Memory can exist independently and be connected to the processor via an address bus, data bus, and control bus. Memory can also be integrated with the processor.

[0092] The memory stores programs that execute the scheme of this disclosure and is controlled by a processor. The processor executes the programs stored in the memory. The program may include one or more software modules. The method for replicating real people based on a large model in the above embodiments can be implemented by a processor and one or more software modules in the program in the memory. However, this disclosure is not limited thereto. The method for replicating real people based on a large model in the above embodiments can also be implemented by circuitry.

[0093] I / O interfaces connect to input devices such as mice, microphones, keyboards, and touchscreens, as well as output devices such as speakers, printers, and monitors. I / O interfaces can also use transceivers or similar devices to communicate with other devices or communication networks such as Ethernet, Radio Access Networks (RAN), and Wireless Local Area Networks (WLAN).

[0094] As an exemplary embodiment, an electronic device may include a plurality of processors, each of which may be a single-core processor or a multi-core processor. As used herein, a processor may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0095] The aforementioned electronic device can be a general-purpose computer device or a special-purpose computer device. In specific implementations, the computer device can be a desktop computer, laptop computer, network server, PDA, mobile phone, tablet computer, wireless terminal device, communication device (e.g., access point, router, gateway, etc.), or embedded device. The embodiments of this disclosure do not limit the type of computer device, as long as it has a processor and memory.

[0096] It should be understood that Figure 4The illustrated electronic device is merely one example of this disclosure, and the electronic devices of this disclosure may also include elements or components not shown in the examples above. For example, some electronic devices also include display units such as displays, some electronic devices also include human-computer interaction elements such as buttons and keyboards, and some electronic devices also include various sensors, such as gesture sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, grip sensors, proximity sensors, color sensors, infrared (IR) sensors, biometric sensors, temperature sensors, humidity sensors, illuminance sensors, etc. Any electronic device capable of executing a computer-readable program in its memory to implement the methods or at least some steps of the methods described in this disclosure may be considered an electronic device covered by this disclosure.

[0097] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented in software, in hardware, or in a combination of software and necessary hardware. Therefore, as... Figure 5 As shown, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-transitory computer-readable storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.) or on a network, including several commands to cause a computing device (such as a personal computer, server, or network device, etc.) to execute the above-described methods according to the embodiments of this disclosure.

[0098] Software products may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media include, but are not limited to: electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0099] This disclosure also provides a computer-readable storage medium storing a program that, when executed by a processor, implements the above-described method for replicating real people based on a large model. Figure 5 This is a block diagram illustrating a computer-readable medium according to embodiments of the present disclosure.

[0100] Computer-readable storage media may include data signals propagated in baseband or as part of a carrier wave, carrying a readable program. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can send, propagate, or transmit a program for use by or in connection with a command execution system, apparatus, or device. The program contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, or any suitable combination thereof.

[0101] Programs for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. Programs can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing devices can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to external computing devices (e.g., via the Internet using an Internet service provider).

[0102] The aforementioned computer-readable medium carries one or more programs (e.g., computer-executable programs) that, when executed by one or more devices, cause the computer-readable medium to implement the methods of this disclosure.

[0103] Those skilled in the art will understand that the above modules can be distributed in one device as described in the embodiments, or they can be varied and located in one or more related devices. The modules of the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.

[0104] Exemplary embodiments of the present disclosure have been specifically shown and described above. It should be understood that the present disclosure is not limited to the detailed structures, arrangements, or implementation methods described herein; rather, the present disclosure is intended to cover various modifications and equivalent arrangements contained within the spirit and scope of the appended claims.

[0105] Although preferred embodiments of this disclosure have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this disclosure.

[0106] Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from its spirit and scope. Therefore, if such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, this disclosure is also intended to include such modifications and variations.

Claims

1. A method for replicating real people based on a large model, characterized in that, include: Receive input information; Determine whether a knowledge base retrieval is needed based on the content of the input information; When it is determined that a knowledge base retrieval is required, an enhanced prompt is generated based on the input information and the content retrieved from the knowledge base, and the enhanced prompt is input into the role model to generate a response; as well as When it is determined that a knowledge base retrieval is not required, the input information is directly input into the large role model to generate a response.

2. The method according to claim 1, characterized in that, When it is determined that a knowledge base retrieval is required, a corresponding query vector is generated based on the input information. The query vector is used to retrieve multiple relevant corpus segments from the knowledge base; The retrieved corpus segment is concatenated with the input information to form the enhanced prompt; Input the enhanced prompt into the large character model; as well as The large role model generates a response based on the retrieved corpus segments.

3. The method according to claim 2, characterized in that, Also includes: Collect multimodal data, including video, audio, and text data; Training corpora for specific characters are generated based on the aforementioned multimodal data; The knowledge base for a specific person is constructed based on the multimodal data; as well as The training corpus is input into the base model for fine-tuning to obtain the large model of the character.

4. The method according to claim 3, characterized in that, Generating training corpora for specific individuals based on the multimodal data includes: constructing question-answer pairs related to identity features based on the text data.

5. The apparatus according to claim 4, characterized in that, Generating training corpora for specific characters based on the aforementioned multimodal data also includes: Extract question-and-answer pairs from the text data that express the viewpoints of specific individuals, and Preset unknown areas and refusal templates for specific characters.

6. The apparatus according to claim 5, characterized in that, Generating training corpora for specific characters based on the multimodal data also includes: constructing question-and-answer pairs that simulate the boredom of specific characters.

7. The apparatus according to claim 6, characterized in that, Generating training corpora for specific individuals based on the multimodal data also includes constructing realistic multi-turn dialogue data by labeling language style elements that reflect the speaking style of specific individuals.

8. A device for replicating real people based on a large model, characterized in that, include: The input information receiving module is configured to receive input information. The knowledge base retrieval determination module is configured to determine whether a knowledge base retrieval is needed based on the content of the input information. The retrieval enhancement generation execution module is configured to generate enhanced prompts based on the input information and the content retrieved from the knowledge base when it is determined that a knowledge base retrieval is needed, and input the enhanced prompts into the role model to generate a response; as well as The direct generation module is configured to directly input the input information into the large role model to generate a response when it is determined that a knowledge base retrieval is not required.

9. An electronic device, comprising: processor; And a memory for storing a program that, when executed by the processor, performs the method as described in any one of claims 1-8.

10. A computer-readable medium storing a program that, when executed by the processor, performs the method as described in any one of claims 1-8.