Method for constructing cultural relic knowledge natural language dialogue engine based on large model

CN122796089APending Publication Date: 2026-09-22DONGFANG JINDIAN DIGITAL TECH (HUNAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611231873.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-14
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0004]为了解决现有文物对话系统生成的回答缺乏针对性、沉浸感和个性化,难以满足智慧博物馆高质量导览服务的需求的技术问题,本发明的目的在于提供一种基于大模型的文物知识自然语言对话引擎构建方法,所采用的技术方案具体如下:

Benefits of technology

本发明通过构建层层递进的链式分析架构,实现了从文物参观场景触发到对话内容偏重再到对话内容长度控制的完整个性化对话链路。通过量化用户在单次及多次提问中的关键词所属知识维度的关注偏重,使系统能精准识别用户对特定知识维度的兴趣强弱,生成重点突出的回答。基于“驻足”、“快速划过”等场景事件主动发起并设定初始对话策略,使系统响应更贴合用户的实时游览节奏。通过分析用户历史行为中的驻留特征,识别其对信息详略的接受偏好,并据此动态修正回答篇幅,在满足文物信息需求的同时优化阅读、收听体验。通过本发明满足了不同游客用户在接受文物讲解时的针对性、沉浸感和个性化,满足了智慧博物馆高质量导览服务的需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122796089A_ABST
    Figure CN122796089A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of electric digital data processing, in particular to a cultural relic knowledge natural language dialogue engine construction method based on a large model, which comprises the following steps: according to a key word in a current question text, determining a current instant attention bias of a user to a target knowledge dimension; correcting the current instant attention bias by using a historical instant attention bias of the target knowledge dimension to obtain a corrected attention bias; determining an initial length adjustment factor according to a current residence time length of the user to a currently visited cultural relic, and determining a content length correction factor according to a historical residence time length of the user to a visited cultural relic; combining the initial length adjustment factor and the content length correction factor to obtain a final length adjustment factor; and combining the final length adjustment factor of dialogue output content of the model and the corrected attention bias of each knowledge dimension to determine actual dialogue output content. Through the technical scheme of the application, the cultural relic interpretation can be targeted, immersive and personalized for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing technology, specifically to a method for constructing a natural language dialogue engine for cultural relic knowledge based on a large model. Background Technology

[0002] With the popularization of museum digitization and smart cultural tourism, cultural relic knowledge dialogue engines based on LLM (Large Language Model) have become an important means to enhance the visitor experience. These systems aim to provide visitors with in-depth, interesting, and personalized cultural relic explanation services through natural language interaction.

[0003] However, existing cultural relic dialogue systems mostly adopt a passive "question-and-answer" model, generating responses solely based on the user's input text. This model has several significant limitations: First, it fails to perceive the user's physical environment and behavioral state within the exhibition hall (such as lingering or quickly scrolling), leading to a disconnect between the timing of the dialogue and the level of detail in the response and the user's actual state. Second, when generating responses, the system's emphasis on various knowledge dimensions related to the cultural relic (such as historical context, craftsmanship, and historical stories) is static and fixed, unable to dynamically adjust based on the user's keyword focus in a single question and their interest in the historical dialogue. Third, it ignores individual differences in reading habits and patience levels among users, outputting responses of the same length to all users, resulting in some users finding the information tedious while others feel the explanations are not in-depth enough. These shortcomings collectively lead to generated responses lacking relevance, immersion, and personalization, failing to meet the demands of high-quality guided tours in smart museums. Summary of the Invention

[0004] To address the technical problem that existing cultural relic dialogue systems lack targeted, immersive, and personalized responses, failing to meet the demands of high-quality guided tours in smart museums, this invention aims to provide a method for constructing a natural language dialogue engine for cultural relic knowledge based on a large-scale model. The specific technical solution adopted is as follows: This invention provides a method for constructing a natural language dialogue engine for cultural relic knowledge based on a large model, the method comprising: Based on the knowledge dimension to which the keywords in the user's current question text about the cultural relic belong, determine the user's current immediate focus on the target knowledge dimension; By using the historical real-time attention bias of the target knowledge dimension to correct the current real-time attention bias, we obtain the corrected attention bias of the target knowledge dimension. The initial length adjustment factor of the model dialogue output content is determined based on the user's current dwell time on the currently visited cultural relic, and the content length correction factor is determined based on the user's historical dwell time on the visited cultural relic. The final length adjustment factor is obtained by combining the initial length adjustment factor and the content length correction factor. The actual output content of the dialogue is determined by combining the final length adjustment factor of the model dialogue output content and the adjusted focus of each knowledge dimension.

[0005] Furthermore, determining the user's current immediate focus on the target knowledge dimension based on the knowledge dimension to which the keywords in the user's current question text about the currently visited cultural relic belong includes: Based on the knowledge dimension to which each keyword in the user's current question text about the cultural relic belongs, determine the number of matching keywords belonging to the target knowledge dimension; The frequency ratio of the target knowledge dimension in the current question is obtained by comparing the number of matched keywords with the total number of keywords in the current question text. The user's current immediate focus on the target knowledge dimension is then determined based on the frequency ratio.

[0006] Furthermore, determining the user's current immediate focus on the target knowledge dimension based on the word frequency ratio includes: Identify the matching keywords in the current question text that belong to the target knowledge dimension, and determine the position weighting factor based on the sentence position of the matching keywords in the current question text; By combining the positional weighting factor and word frequency ratio of each matching keyword, we can obtain the user's current immediate focus on the target knowledge dimension.

[0007] Furthermore, the step of using historical real-time attention bias of the target knowledge dimension to correct the current real-time attention bias to obtain the corrected attention bias of the target knowledge dimension includes: The average historical attention bias is calculated by utilizing the user's real-time historical attention bias in the target knowledge dimension of each historical question text of the currently visited cultural relic. The current real-time attention bias is corrected by using the historical real-time attention bias average, resulting in the corrected attention bias for the target knowledge dimension.

[0008] Furthermore, the step of calculating the average historical real-time attention bias by utilizing the user's historical attention bias in the target knowledge dimension of each historical question text of the currently visited cultural relic includes: Determine the historical and immediate focus of the target knowledge dimension in the historical question text and the time interval between the historical question moment and the current question moment; A time decay factor is constructed using the time interval, and the weighted average of historical real-time attention bias is calculated using the time decay factor as a weight.

[0009] Furthermore, the step of determining the initial length adjustment factor of the model dialogue output content based on the user's current dwell time on the currently visited cultural relic includes: When the user's current dwell time on the currently visited cultural relic is greater than or equal to a preset first threshold, the initial length adjustment factor of the model dialogue output content is determined to be the first preset value. When the current dwell time of the user on the currently visited cultural relic is less than or equal to the preset second threshold, the initial length adjustment factor of the model dialogue output content is determined to be the second preset value. When the user's current dwell time on the currently visited cultural relic is greater than the preset second threshold but less than the preset first threshold, the initial length adjustment factor of the model dialogue output content is determined to be the third preset value. Among them, the first preset value is used as the base value, the second preset value is less than the third preset value, and the third preset value is less than the first preset value.

[0010] Furthermore, the determination of the content length correction factor based on the user's historical dwell time on the visited cultural relics includes: Based on the user's historical dwell time and suggested dwell time for the visited cultural relics, the user's dwell time investment ratio for the visited cultural relics is obtained; The average retention rate is calculated based on the user's retention rate for each visited artifact, and this average retention rate is used as a content length correction factor.

[0011] Furthermore, the step of combining the initial length adjustment factor and the content length correction factor to obtain the final length adjustment factor includes: The initial length adjustment factor is corrected using the content length correction factor to obtain the final length adjustment factor; Determine whether the final length adjustment factor exceeds the preset reasonable range. If it does, perform truncation.

[0012] Furthermore, the determination of the actual dialogue output content by combining the final length adjustment factor of the combined model dialogue output content and the adjusted focus bias of each knowledge dimension includes: The length constraint instruction is generated based on the final length adjustment factor of the output content of the model dialogue, and the content emphasis constraint instruction for different knowledge dimensions is determined based on the correction of each knowledge dimension. Input the length constraint and content emphasis constraint into the preset large language model to determine the actual output content of the dialogue.

[0013] Furthermore, the step of inputting the length constraint instruction and content emphasis constraint instruction into a preset large language model to determine the actual output content of the dialogue includes: Based on the cultural relic currently being visited by the user, retrieve knowledge entries related to multiple knowledge dimensions of the currently visited cultural relic; The length constraint, content emphasis constraint, and knowledge items are integrated into prompt words and input into a pre-defined large language model to determine the actual output content of the dialogue.

[0014] The present invention has the following beneficial effects: This invention constructs a progressively layered chain-like analysis architecture, realizing a complete personalized dialogue chain from triggering the visit to emphasizing dialogue content and controlling dialogue length. By quantifying the user's focus on the knowledge dimensions associated with keywords in single and multiple questions, the system can accurately identify the user's interest in specific knowledge dimensions and generate answers that highlight key points. Based on scenarios such as "pausing" and "quickly scrolling," the system proactively initiates and sets initial dialogue strategies, making the system's response more aligned with the user's real-time tour pace. By analyzing the user's dwell characteristics in historical behavior, the system identifies their preference for information detail and dynamically adjusts the length of responses accordingly, optimizing the reading and listening experience while meeting the information needs of cultural relics. This invention satisfies the needs of different visitors for targeted, immersive, and personalized cultural relic explanations, meeting the requirements of high-quality guided tour services in smart museums. Attached Figure Description

[0015] Figure 1 The flowchart shows the steps of a method for constructing a natural language dialogue engine for cultural relics knowledge based on a large model, as provided in one embodiment of the present invention. Figure 2 A detailed flowchart of step S1 in a method for constructing a natural language dialogue engine for cultural relics knowledge based on a large model, provided in an embodiment of the present invention. Figure 3 A detailed flowchart of step S2 in a method for constructing a natural language dialogue engine for cultural relics knowledge based on a large model, provided in an embodiment of the present invention; Figure 4 A detailed flowchart of step S3 in a method for constructing a natural language dialogue engine for cultural relics knowledge based on a large model, provided in an embodiment of the present invention; Figure 5 This is a detailed flowchart of step S4 in a method for constructing a natural language dialogue engine for cultural relics knowledge based on a large model, provided as an embodiment of the present invention. Detailed Implementation

[0016] To facilitate understanding of the various embodiments of the present invention, the circumstances and main objectives of the present invention are described herein: Existing solutions fail to utilize IoT or visual sensors to perceive user physical location, dwell time, and other contextual events, resulting in rigid dialogue patterns. They also fail to analyze the immediate focus reflected in the keyword types within a single user question, nor do they track the sustained interest trends reflected in the distribution of keyword types across multiple interactions. Furthermore, they fail to identify the specific preferences of different users regarding the level of detail required for reading, leading to an inability to adaptively match the output text length to the user's comprehension capacity. To address these shortcomings, this invention proposes a dialogue method for cultural relic knowledge that integrates contextual event triggering, question emphasis analysis, and user length preference adaptation. Through a layered, chain-like computational architecture, the dialogue triggering strategy and the dimensional emphasis of the response content are determined sequentially, and finally, the response length is adaptively adjusted based on the user's historical behavior, achieving a highly personalized dialogue experience.

[0017] The following description, in conjunction with the accompanying drawings, details a specific scheme for constructing a natural language dialogue engine for cultural relic knowledge based on a large model, provided by this invention.

[0018] Example 1: Please see Figure 1 , Figure 1 The flowchart illustrates the steps of a method for constructing a natural language dialogue engine for cultural relics knowledge based on a large model, according to an embodiment of the present invention.

[0019] The method for constructing a natural language dialogue engine for cultural relic knowledge based on a large model includes the following steps: Step S1: Determine the user's current immediate focus on the target knowledge dimension based on the knowledge dimension to which the keywords in the user's current question text about the currently visited cultural relic belong; In this embodiment, firstly, the ultra-wideband positioning base station or Bluetooth beacon array deployed in the exhibition hall is combined with the signal strength indication value received by the user's smart terminal (such as a mobile phone or a smart guide terminal distributed in the museum exhibition hall). The three-point positioning algorithm is used to obtain the timestamps of the user entering and leaving each cultural relic exhibition area, calculate the actual stay time, and locate the cultural relic that the user is currently visiting, i.e. the currently visited cultural relic.

[0020] The user's natural language question text can be obtained through the exhibition hall or the user's terminal interactive interface. The question text for the currently visited cultural relic is the current question text, and the question text for the cultural relic that has been visited is the historical question text. The Jieba word segmentation tool is used in combination with a custom cultural relic domain dictionary for word segmentation processing, and the keyword extraction technology based on the word frequency-inverse document frequency algorithm is used to identify core words, and the absolute position index of each keyword in the question text is recorded simultaneously.

[0021] After extracting keywords, a pre-defined knowledge dimension keyword database, historical question records from user browsing sessions, and corresponding keyword attribution tags are retrieved from the backend database using structured query statements.

[0022] The knowledge dimensions here include, but are not limited to, the historical background, craftsmanship, historical stories, cultural significance, and functions of cultural relics. They can also be customized as needed. Each knowledge dimension often includes multiple matching keywords. Therefore, a knowledge dimension keyword library can be pre-built, which includes multiple knowledge dimensions and the keywords contained in each knowledge dimension.

[0023] For the current question text, after extracting the knowledge dimensions to which each keyword belongs, for any knowledge dimension, i.e. the target knowledge dimension, it contains one or more keywords in the current question text, that is, one or more keywords belong to the target knowledge dimension. By statistically analyzing the keywords corresponding to the target knowledge dimension, we can determine the user's current immediate focus on the target knowledge dimension. The greater the current immediate focus, the more the user is paying attention to this knowledge dimension and the more interested they are in this knowledge dimension.

[0024] Specifically, please refer to Figure 2 Step S1 includes: Step S11: Determine the number of matching keywords belonging to the target knowledge dimension based on the knowledge dimension to which each keyword in the user's current question text about the currently visited cultural relic belongs; Step S12: Based on the number of matched keywords and the total number of keywords in the current question text, obtain the word frequency ratio of the target knowledge dimension in the current question, and determine the user's current immediate focus on the target knowledge dimension based on the word frequency ratio.

[0025] More specifically, step S12, determining the user's current immediate focus on the target knowledge dimension based on the word frequency ratio, includes: Identify the matching keywords in the current question text that belong to the target knowledge dimension, and determine the position weighting factor based on the sentence position of the matching keywords in the current question text; By combining the positional weighting factor and word frequency ratio of each matching keyword, we can obtain the user's current immediate focus on the target knowledge dimension.

[0026] In this embodiment, the knowledge dimension to which the keywords used by the user in the current single question belong directly reflects their current focus. By analyzing the distribution of keywords belonging to each knowledge dimension in the current question text, the user's current immediate focus is quantified.

[0027] Based on the above embodiments, the user's current question text is obtained, and word segmentation and keyword extraction are performed.

[0028] The extracted keywords are matched with keyword databases for each knowledge dimension. The keywords belonging to each knowledge dimension in the current question text are then analyzed. Number of keywords Here will This is recorded as the number of matched keywords; the corresponding keywords are the matched keywords. This retrieves the total number of keywords extracted from the current query text. .

[0029] Computational knowledge dimension (As a target knowledge dimension) the proportion of word frequency in the current question : It should be noted that the total number of keywords... The value should not be 0, as this forms the data basis for various embodiments of the present invention; otherwise, the calculation is meaningless. When the value is 0, the analysis of the current question text can be stopped, assuming that the user has not entered any question text; it can also skip the weight calculation step, set the attention weight of all dimensions to the default mean, and directly pass the user's original question to the large language model with the default length adjustment factor. The large language model will then use its general knowledge to provide a fallback answer, ensuring the closed loop and smoothness of the dialogue process.

[0030] Considering that keywords should have different positional weights in different sentences within the question text (for example, keywords at the beginning of a sentence or near the question word are considered more important and can be considered to be in a key position; the weight of different sentence positions can be customized according to the actual situation), a positional weighting factor is set here. As a positional weight, matching keywords When in a critical position For example, 1.2, otherwise .

[0031] This allows us to obtain the knowledge dimension from the user's current question text. Current immediate focus : in, For all knowledge dimensions in the current question The sum of positional weighting factors for matching keywords.

[0032] Knowledge Dimension of Perform maximum and minimum value normalization to obtain the normalized real-time attention bias. Its range is The maximum and minimum values ​​used for normalization here can represent the knowledge dimensions of historical visitor users. The system prioritizes both the maximum and minimum values ​​for immediate focus. For normalized values ​​exceeding the range, truncation is performed to limit them to the range. This is a standard mathematical process and will not be elaborated further. Specifically, in this embodiment, truncation for normalized values ​​exceeding the range means that when the normalized object is greater than the corresponding maximum value, the normalization result is 1; when the normalized object is less than the corresponding minimum value, the normalization result is 0. It should be noted that in this embodiment, during the maximum and minimum value normalization process, as much historical data as possible should be obtained to ensure a significant difference between the maximum and minimum values. If, in extreme cases, the maximum and minimum values ​​are equal, the maximum and minimum values ​​used for normalization are taken from theoretical boundary values. For example, the minimum value is 0 (i.e., the number of matched keywords is 0), and the maximum value is the theoretical extreme value when all keywords belong to this dimension and are in the highest weight position (e.g., 1.2).

[0033] Step S2: Use the historical real-time attention bias of the target knowledge dimension to correct the current real-time attention bias to obtain the corrected attention bias of the target knowledge dimension. Specifically, please refer to Figure 3 Step S2 includes: Step S21: Calculate the average historical real-time attention bias by utilizing the user's historical real-time attention bias in the target knowledge dimension of each historical question text of the currently visited cultural relic. More specifically, step S21 includes: Determine the historical and immediate focus of the target knowledge dimension in the historical question text and the time interval between the historical question moment and the current question moment; A time decay factor is constructed using the time interval, and the weighted average of historical real-time attention bias is calculated using the time decay factor as a weight.

[0034] Step S22: Use the historical real-time attention bias average to correct the current real-time attention bias, and obtain the corrected attention bias of the target knowledge dimension.

[0035] In this embodiment, the user's historical question text sequence during this visit reflects their ongoing interest tendencies. This embodiment uses the current immediate focus as the basis for calculation and utilizes historical information to progressively correct it, resulting in a more robust content focus.

[0036] Retrieve all historical questions asked by the user during the current visit session that are related to the currently visited artifact. This is essentially a sequence of all historical question texts. Each element in the sequence represents a historical question text, and m represents the index of the last element in the sequence, which also corresponds to the number of elements and the total number of historical question texts.

[0037] For each historical question text Similar to calculating the current immediate focus, we can also calculate its corresponding knowledge dimension. Historical real-time focus Here, k represents the index of an element in the sequence of historical question texts. Set time decay factor ,in Questioning texts about history The time interval since the event occurred, specifically the historical question text. The time interval between the corresponding historical question time and the current question time corresponding to the current question text (the unit of time interval in this embodiment can be hours). This is the attenuation coefficient, which can be based on test calibration (the unit of the attenuation coefficient is opposite to the unit of the time interval to ensure that the product of the two is unitless), for example, an empirical value of 0.6. The more recent the time, the shorter the time interval. The larger.

[0038] This leads to the knowledge dimension. Historical weighting, meaning that historical real-time focus leans towards the average value. Specifically, it can be a weighted average of historical data with specific focus: It should be noted that, according to the defined time decay factor , Not equal to 0, It is also not 0. It should be understood that in the special case of the cold start phase, i.e., when m equals 0, a weighted average calculation is not performed; instead, the value is directly set. It equals 0.

[0039] Focus on the present moment As a primary factor, historical real-time data focuses on the average value. As a correction factor, the knowledge dimension is calculated. The focus of the corrected generation When historical data focuses on average values... The larger the value, the more sustained the user's interest in that knowledge dimension; the current focus should be increased accordingly. The effect of this. Therefore, the corrected formula is: in, This is an amplification factor for historical interest (referring to the average of current historical attention), for example, taking... This formula prioritizes current, immediate concerns and then reinforces those concerns based on the strength of historical interests, ensuring more accurate responses across different knowledge dimensions. of After performing maximum and minimum value normalization, the final corrected value is then adjusted to focus on the bias. Its range is The maximum and minimum values ​​used for normalization can be taken from the knowledge dimensions of historical visitor users, respectively. The focus is on Maximum value and focus Minimum value.

[0040] Step S3: Determine the initial length adjustment factor of the model dialogue output content based on the user's current dwell time on the currently visited cultural relic, and determine the content length correction factor based on the user's historical dwell time on the visited cultural relic. Specifically, step S3, which determines the initial length adjustment factor of the model dialogue output content based on the user's current dwell time on the currently visited cultural relic, includes: When the user's current dwell time on the currently visited cultural relic is greater than or equal to a preset first threshold, the initial length adjustment factor of the model dialogue output content is determined to be the first preset value. When the current dwell time of the user on the currently visited cultural relic is less than or equal to the preset second threshold, the initial length adjustment factor of the model dialogue output content is determined to be the second preset value. When the user's current dwell time on the currently visited cultural relic is greater than the preset second threshold but less than the preset first threshold, the initial length adjustment factor of the model dialogue output content is determined to be the third preset value. Among them, the first preset value is used as the base value, the second preset value is less than the third preset value, and the third preset value is less than the first preset value.

[0041] In this embodiment, different users exhibit varying levels of acceptance and patience when reading text of the same length. To avoid this difference affecting the answers to questions, it can be quantified by analyzing the duration of user dwell time in front of different cultural relics. Using the initial length adjustment factor of the model's dialogue output as the basis for calculation, it is progressively adjusted using the user's historical dwell time characteristics to obtain a personalized final length adjustment factor.

[0042] Environmental sensors (such as Bluetooth beacons and ultra-wideband positioning base stations) are deployed within the museum's exhibition halls. The system acquires real-time location information from the smart tour guide terminals carried by users.

[0043] Referring to Table 1 below, define the following scenario events, corresponding dialogue strategy modes, and initial output length adjustment factors. : The shorter the time users spend in front of cultural relics, the less they think about them, the more superficial their understanding is, and the less demand they have for in-depth explanations; conversely, the longer they stay in front of cultural relics, the greater their demand for in-depth explanations.

[0044] Specifically, please refer to Figure 4 Step S3, determining the content length correction factor based on the user's historical dwell time on the visited cultural relics, includes: Step S31: Based on the user's historical dwell time and suggested dwell time for the visited cultural relics, obtain the user's dwell time input ratio for the visited cultural relics. Step S32: Obtain the average dwell time ratio based on the user's dwell time ratio for each visited cultural relic, and use the average dwell time ratio as a content length correction factor.

[0045] In this embodiment, the system obtains the historical dwell time sequence of users in front of the cultural relics they have visited by querying the behavior log data table, and obtains the system's preset suggested dwell time for each cultural relic by accessing the cultural relic configuration information table.

[0046] Determine if the user has sufficient historical visit data. Set a minimum sample size threshold of three visited artifacts for valid data (the specific threshold can be adjusted according to actual conditions). The number of historical visit records, i.e., the number of visited artifacts, is considered. If the value is less than 3, the length adjustment of the model dialogue output content will not be performed, and the initial length adjustment factor will be directly adjusted. As the final length adjustment factor Output.

[0047] like If the value is greater than or equal to 3, retrieve the user's activity during this visit session. The actual duration of historical stay in front of the visited cultural relics. .

[0048] The system is for this Recommended stay duration for each cultural relic (Can be preset according to the importance of the cultural relic, not 0).

[0049] Calculate the number of times a user has visited a single artifact. Previous residency investment ratio This ratio reflects how much time users are willing to invest in an artifact relative to the suggested time.

[0050] Calculate the average time spent by users in front of all the artifacts they have visited. .

[0051] when This indicates that the user prefers in-depth browsing and is willing to spend more time reading detailed information than suggested.

[0052] when This indicates that the user tends to browse quickly and prefers concise and refined information.

[0053] Adjust the initial length factor As the main factor, the average retention ratio As a content length correction factor for the main factor.

[0054] Step S4: Combine the initial length adjustment factor and the content length correction factor to obtain the final length adjustment factor. Combine the final length adjustment factor of the model dialogue output content and the corrected attention bias of each knowledge dimension to determine the actual output content of the dialogue.

[0055] Specifically, step S4, which combines the initial length adjustment factor and the content length correction factor to obtain the final length adjustment factor, includes: The initial length adjustment factor is corrected using the content length correction factor to obtain the final length adjustment factor; Determine whether the final length adjustment factor exceeds the preset reasonable range. If it does, perform truncation.

[0056] In this embodiment, when the content length correction factor The larger the value, the more users prefer longer texts; therefore, the initial length adjustment factor should be used as a base. Increase the output length on ). When The smaller the value, the more users prefer shorter text; therefore, the output length should be reduced from the initial settings. This leads to the final length adjustment factor. : in, This is a user-specific adjustment coefficient; the preset value can be... This is used to control the degree of user-specific correction to the length of the model's dialogue output content.

[0057] It is still necessary to Perform amplitude limiting to ensure it falls within a preset reasonable range (e.g.) To avoid generating answers that are too short or too long, if the calculated result is less than 0.3, it will be forced to a value of 0.3 to ensure that the answer content retains at least the most essential information; if the calculated result is greater than 1.5, it will be forced to a value of 1.5 to avoid the answer content being too long and affecting the user experience.

[0058] Specifically, please refer to Figure 5 Step S4, combining the final length adjustment factor of the model dialogue output content and the adjusted focus bias of each knowledge dimension, determines the actual dialogue output content, including: Step S41: Generate length constraint instructions based on the final length adjustment factor of the model dialogue output content, and determine the content emphasis constraint instructions for different knowledge dimensions based on the corrected focus of each knowledge dimension. Step S42: Input the length constraint instruction and content emphasis constraint instruction into the preset large language model to determine the actual output content of the dialogue.

[0059] More specifically, step S42 includes: Based on the cultural relic currently being visited by the user, retrieve knowledge entries related to multiple knowledge dimensions of the currently visited cultural relic; The length constraint, content emphasis constraint, and knowledge items are integrated into prompt words and input into a pre-defined large language model to determine the actual output content of the dialogue.

[0060] In this embodiment, after completing the calculations of the aforementioned embodiments and obtaining the corrections for each knowledge dimension, the focus is shifted to the bias. and the final length adjustment factor Next, the system enters the large language model response generation stage. The preset large language models here include DeepSeek series, GPT series, Gemini series, etc., and no restrictions are imposed here.

[0061] First, the system constructs a structured large language model prompt, which is composed of several functionally defined components. The first part is the task context setting domain, which clearly informs the large language model that its current role is a professional museum guide, to ensure that the tone and style of the generated response are consistent with the tour guiding scenario.

[0062] The second part addresses the length control domain, directly presenting the calculated final length adjustment factor. This translates into a length constraint instruction that a large language model can understand. Typically, this means requiring the large language model to control the total number of words in the generated response (i.e., the actual output of the dialogue) to a baseline content length (corresponding to the aforementioned in-depth explanation mode, specifically a total word count of 1000-1200 words). times, for example when When the value is 1.2, the large language model should output an answer that is approximately 120% of the baseline content length, while when... When the value is 0.5, the output should be a concise response that is about 50% of the length of the baseline content. It should be noted that the actual output of the dialogue is generally within a certain word range.

[0063] The third part focuses on the domain of emphasis in the response. This part uses tables or structured text to present the calculated and corrected emphasis. Each knowledge dimension is listed with its corresponding name and a clear content emphasis constraint. The large language model is required to strictly allocate detail based on the adjusted emphasis of each knowledge dimension when organizing the response. Knowledge dimensions with higher adjusted emphasis should be given more space and detail for focused explanation, while those with lower adjusted emphasis can be briefly mentioned or omitted. It should be noted that in extreme cases where all calculated adjusted emphasis is 0, the above operation can be canceled, and all responses will be directly explained according to the same preset allocation mechanism.

[0064] The fourth part is the user question text field, where you can directly enter the original question that the user is currently asking.

[0065] The fifth part is the cultural relic knowledge base domain, which is filled with knowledge entries related to the currently visited cultural relic from the system's backend database, providing factual basis for the large language model to generate answers. Based on the cultural relic identifier of the currently visited cultural relic, key-value matching is used to retrieve detailed content knowledge entries for that cultural relic under each knowledge dimension from the cultural relic knowledge base.

[0066] After the structured prompts are constructed, the system uses them as input parameters to call the pre-deployed large language model service through the application programming interface. Upon receiving the prompts, the large language model generates text based on various constraints, including length constraints and content emphasis constraints, along with relevant knowledge items. Because the prompts contain both content emphasis vectors and length control factors, the large language model performs semantic planning under dual constraints during generation. Within a given total length limit, it dynamically allocates the expansion of each knowledge dimension according to its weight ratio (adjusted emphasis proportion), ensuring that the final output (i.e., the actual dialogue output) conforms to the user's personalized reading preferences in terms of macro-length while accurately responding to the user's immediate focus and ongoing interest in terms of micro-content structure. After the dialogue content is generated in the large language model, the system receives the returned text content and can selectively convert it into an audio stream through the speech synthesis module. Finally, it is presented to the user through the user terminal device (such as a mobile phone or a smart guide terminal distributed in the exhibition hall) or the large screen terminal in the exhibition hall in the form of voice playback and / or text display, completing a complete round of personalized cultural relic knowledge dialogue interaction.

[0067] This invention constructs a progressively layered chain-like analysis architecture, realizing a complete personalized dialogue chain from triggering the visit to emphasizing dialogue content and controlling dialogue length. By quantifying the user's focus on the knowledge dimensions associated with keywords in single and multiple questions, the system can accurately identify the user's interest in specific knowledge dimensions and generate answers that highlight key points. Based on scenarios such as "pausing" and "quickly scrolling," the system proactively initiates and sets initial dialogue strategies, making the system's response more aligned with the user's real-time tour pace. By analyzing the user's dwell characteristics in historical behavior, the system identifies their preference for information detail and dynamically adjusts the length of responses accordingly, optimizing the reading and listening experience while meeting the information needs of cultural relics. This invention satisfies the needs of different visitors for targeted, immersive, and personalized cultural relic explanations, meeting the requirements of high-quality guided tour services in smart museums.

[0068] Example 2: This invention also proposes a device for constructing a natural language dialogue engine for cultural relics knowledge based on a large model. The device can be a mobile phone, a navigation terminal, a computer, a server, or a combination of multiple devices. Based on the hardware structure of the aforementioned device for constructing a natural language dialogue engine for cultural relics knowledge based on a large model, various embodiments of the method for constructing a natural language dialogue engine for cultural relics knowledge based on a large model of this invention are implemented.

[0069] Furthermore, the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a program for constructing a natural language dialogue engine for cultural relics knowledge based on a large model. When executed by a processor, the program implements the method for constructing a natural language dialogue engine for cultural relics knowledge based on a large model, as described above.

[0070] The method implemented when the large-model-based cultural relic knowledge natural language dialogue engine construction program is executed can refer to the method of the large-model-based cultural relic knowledge natural language dialogue engine construction in this invention, and will not be repeated here.

Claims

1. A method for constructing a natural language dialogue engine for cultural relic knowledge based on a large model, characterized in that, The method includes: Based on the knowledge dimension to which the keywords in the user's current question text about the cultural relic belong, determine the user's current immediate focus on the target knowledge dimension; By using the historical real-time attention bias of the target knowledge dimension to correct the current real-time attention bias, we obtain the corrected attention bias of the target knowledge dimension. The initial length adjustment factor of the model dialogue output content is determined based on the user's current dwell time on the currently visited cultural relic, and the content length correction factor is determined based on the user's historical dwell time on the visited cultural relic. The final length adjustment factor is obtained by combining the initial length adjustment factor and the content length correction factor. The actual output content of the dialogue is determined by combining the final length adjustment factor of the model dialogue output content and the adjusted focus of each knowledge dimension.

2. The method for constructing a natural language dialogue engine for cultural relic knowledge based on a large model according to claim 1, characterized in that, The process of determining the user's current immediate focus on a target knowledge dimension based on the knowledge dimension to which the keywords in the user's current question text about the currently visited cultural relic belong includes: Based on the knowledge dimension to which each keyword in the user's current question text about the cultural relic belongs, determine the number of matching keywords belonging to the target knowledge dimension; The frequency ratio of the target knowledge dimension in the current question is obtained by comparing the number of matched keywords with the total number of keywords in the current question text. The user's current immediate focus on the target knowledge dimension is then determined based on the frequency ratio.

3. The method for constructing a natural language dialogue engine for cultural relic knowledge based on a large model according to claim 2, characterized in that, The method of determining the user's current immediate focus on the target knowledge dimension based on word frequency ratio includes: Identify the matching keywords in the current question text that belong to the target knowledge dimension, and determine the position weighting factor based on the sentence position of the matching keywords in the current question text; By combining the positional weighting factor and word frequency ratio of each matching keyword, we can obtain the user's current immediate focus on the target knowledge dimension.

4. The method for constructing a natural language dialogue engine for cultural relic knowledge based on a large model according to claim 1, characterized in that, The process of adjusting the current attention bias based on the historical real-time attention bias of the target knowledge dimension to obtain the corrected attention bias of the target knowledge dimension includes: The average historical attention bias is calculated by utilizing the user's real-time historical attention bias in the target knowledge dimension of each historical question text of the currently visited cultural relic. The current real-time attention bias is corrected by using the historical real-time attention bias average, resulting in the corrected attention bias for the target knowledge dimension.

5. The method for constructing a natural language dialogue engine for cultural relic knowledge based on a large model according to claim 4, characterized in that, The method of calculating the average historical attention bias by utilizing the user's real-time historical attention bias in the target knowledge dimension of each historical question text of the currently visited cultural relic includes: Determine the historical and immediate focus of the target knowledge dimension in the historical question text and the time interval between the historical question moment and the current question moment; A time decay factor is constructed using the time interval, and the weighted average of historical real-time attention bias is calculated using the time decay factor as a weight.

6. The method for constructing a natural language dialogue engine for cultural relic knowledge based on a large model according to claim 1, characterized in that, The initial length adjustment factor for determining the model dialogue output content based on the user's current dwell time on the currently visited cultural relic includes: When the user's current dwell time on the currently visited cultural relic is greater than or equal to a preset first threshold, the initial length adjustment factor of the model dialogue output content is determined to be the first preset value. When the current dwell time of the user on the currently visited cultural relic is less than or equal to the preset second threshold, the initial length adjustment factor of the model dialogue output content is determined to be the second preset value. When the user's current dwell time on the currently visited cultural relic is greater than the preset second threshold but less than the preset first threshold, the initial length adjustment factor of the model dialogue output content is determined to be the third preset value. Among them, the first preset value is used as the base value, the second preset value is less than the third preset value, and the third preset value is less than the first preset value.

7. The method for constructing a natural language dialogue engine for cultural relic knowledge based on a large model according to claim 1, characterized in that, The method of determining the content length correction factor based on the user's historical dwell time on the visited cultural relics includes: Based on the user's historical dwell time and suggested dwell time for the visited cultural relics, the user's dwell time investment ratio for the visited cultural relics is obtained; The average retention rate is calculated based on the user's retention rate for each visited artifact, and this average retention rate is used as a content length correction factor.

8. The method for constructing a natural language dialogue engine for cultural relic knowledge based on a large model according to claim 1, characterized in that, The process of combining the initial length adjustment factor and the content length correction factor to obtain the final length adjustment factor includes: The initial length adjustment factor is corrected using the content length correction factor to obtain the final length adjustment factor; Determine whether the final length adjustment factor exceeds the preset reasonable range. If it does, perform truncation.

9. The method for constructing a natural language dialogue engine for cultural relic knowledge based on a large model according to claim 1, characterized in that, The final length adjustment factor of the combined model dialogue output content and the adjusted focus of each knowledge dimension are used to determine the actual dialogue output content, including: The length constraint instruction is generated based on the final length adjustment factor of the output content of the model dialogue, and the content emphasis constraint instruction for different knowledge dimensions is determined based on the correction of each knowledge dimension. Input the length constraint and content emphasis constraint into the preset large language model to determine the actual output content of the dialogue.

10. The method for constructing a natural language dialogue engine for cultural relic knowledge based on a large model according to claim 9, characterized in that, The process of inputting length constraint instructions and content emphasis constraint instructions into a preset large language model to determine the actual output content of the dialogue includes: Based on the cultural relic currently being visited by the user, retrieve knowledge entries related to multiple knowledge dimensions of the currently visited cultural relic; The length constraint, content emphasis constraint, and knowledge items are integrated into prompt words and input into a pre-defined large language model to determine the actual output content of the dialogue.