Content generation method, apparatus, and electronic device

By receiving preset knowledge base and wide area network data from the intelligent assistant application, and combining them with priority generation rules, the target content is generated using a content generation model. This solves the problem of insufficient data within user groups and improves content acquisition efficiency and matching accuracy.

CN118861225BActive Publication Date: 2026-07-21BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-06-21
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, intelligent assistant applications fail to effectively utilize preset knowledge bases and wide area network data within user groups, resulting in the inability to meet users' content acquisition needs and increasing the number of operations required for users to re-initiate data acquisition requests.

Method used

By receiving and generating data acquisition results from a preset knowledge base and a wide area network, and combining the priorities in the content generation rules, the target content is generated using a content generation model, prioritizing the use of preset knowledge base data to ensure that the content matches user needs.

Benefits of technology

It improves the efficiency for users to obtain relevant content from different data sources, reduces the need to re-initiate data retrieval requests, and generates content that is more closely matched to user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118861225B_ABST
    Figure CN118861225B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a content generation method, device and electronic equipment, the method comprising: receiving a first data acquisition result and a second data acquisition result; generating prompt information comprising the first data acquisition result and the second data acquisition result; wherein the prompt information further comprises a content generation rule; the content generation rule comprises a priority of referring to the first data acquisition result and the second data acquisition result when generating a target content; inputting the prompt information to a pre-trained content generation model, and generating the target content by the content generation model according to the content generation rule, the first data acquisition result and the second data acquisition result. The target content generated by the content generation model according to the prompt information matches the user demand. The efficiency of the user obtaining reply content from different data sources can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of Internet technology, and in particular to a content generation method, apparatus, and electronic device. Background Technology

[0002] When users encounter questions and need knowledge, they can ask the intelligent assistant application a question. The intelligent assistant application has the ability to obtain knowledge from the internet, and can retrieve relevant data from the internet based on the user's question, and then respond to the question based on the relevant data retrieved from the internet. Summary of the Invention

[0003] This disclosure provides a content generation method, apparatus, and electronic device.

[0004] In a first aspect, embodiments of this disclosure provide a content generation method, the method comprising: receiving a first data acquisition result and a second data acquisition result; wherein the first data acquisition result is obtained by acquiring data based on knowledge content in a preset knowledge base in response to a content acquisition request, and the second data acquisition result is obtained by acquiring data based on a wide area network in response to the content acquisition request, or the second data acquisition result is generated by a content generation model according to the content acquisition request; generating prompt information including the first data acquisition result and the second data acquisition result; wherein the prompt information further includes content generation rules; the content generation rules include a priority of referring to the first data acquisition result and referring to the second data acquisition result when generating target content; inputting the prompt information into a pre-trained content generation model, and having the content generation model generate target content according to the content generation rules, the first data acquisition result, and the second data acquisition result. Secondly, embodiments of this disclosure provide a content generation apparatus, comprising: a receiving unit, configured to receive a first data acquisition result and a second data acquisition result; wherein the first data acquisition result is obtained by acquiring data based on knowledge content in a preset knowledge base in response to a content acquisition request, and the second data acquisition result is obtained by acquiring data based on a wide area network in response to the content acquisition request, or the second data acquisition result is generated by a content generation model according to the content acquisition request; a first generation unit, configured to generate prompt information including the first data acquisition result and the second data acquisition result; wherein the prompt information further includes content generation rules; the content generation rules include a priority of referring to the first data acquisition result and referring to the second data acquisition result when generating target content; and a second generation unit, configured to input the prompt information into a pre-trained content generation model, and the content generation model generates target content according to the content generation rules, the first data acquisition result, and the second data acquisition result.

[0005] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor and a memory; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory, causing the at least one processor to perform the first aspect and various possible methods described above.

[0006] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the first aspect and various possible methods described above.

[0007] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and various possible methods described above.

[0008] The content generation method, apparatus, and electronic device provided in this embodiment, upon receiving a content acquisition request, obtains a first data acquisition result based on knowledge content in a preset database, and a second data acquisition result based on a wide area network, or generates a second data acquisition result by a content generation model. Then, a prompt message including the first data acquisition result, the second data acquisition result, and content generation rules is generated. The priority in the content generation rules indicates the priority between referencing the first and second data acquisition results when generating the target content. Upon receiving the prompt message, the content generation model generates the target content according to the priority. The priority set in the prompt message can match the source of the response the user expects to receive. Therefore, the target content generated by the content generation model based on the prompt message matches the user's needs. Furthermore, based on the prompt message, the content generation model can also generate the target content from another data acquisition result when the priority-indicated data acquisition result cannot generate the target content. Therefore, it can reduce the user's need to re-initiate the content acquisition request to other data sources and improve the efficiency of obtaining response content from different data sources. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 Flowchart of the content generation method provided in this disclosure Figure 1 ;

[0011] Figure 2 Flowchart of the content generation method provided in this disclosure Figure 2 ;

[0012] Figure 3 A structural block diagram of the content generation apparatus provided in the embodiments of this disclosure;

[0013] Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0015] In some embodiments, when a user initiates a content retrieval request to a smart assistant application (e.g., asks a question), the smart assistant application can retrieve relevant information from its own knowledge data and generate a response. If the smart assistant application's own available knowledge data does not contain relevant information, it can retrieve external data corresponding to the content retrieval request from a wide area network and generate a response based on the external data.

[0016] In some application scenarios, users may belong to a user group. For a user group, it's common practice to build its own internal knowledge base, also known as a user group knowledge base. This knowledge base can include various knowledge data generated within the user group.

[0017] When users within a user group initiate content retrieval requests to the intelligent assistant application, they may need to obtain information from the user group's internal knowledge base. For example, a request to retrieve content related to the user group's internal transaction processing flow would require generating the response content from data in the user group's knowledge base. In this scenario, if the intelligent assistant application generates the response content based on data obtained from the wide area network, it may not be able to meet the user's needs.

[0018] In some implementations, data can be retrieved from a user group based on the user's content request, and a response can be generated based on the data in the user group. However, if there is no relevant data in the user group, and the smart assistant application itself does not have the relevant information, a response cannot be generated. The user needs to initiate a search on a wide area network to obtain the relevant information, which is inconvenient for the user.

[0019] The solution provided in this disclosure combines a first data acquisition result obtained from a preset database and a second data acquisition result obtained from a wide area network, and generates target content based on the priority of referring to the first and second data acquisition results. When the priority indicates that referring to the first data acquisition result has a higher priority, the target content is generated based on the first data acquisition result. Therefore, it can meet the user's need to obtain response content from a preset knowledge base. In addition, when the first data acquisition result is empty, that is, when the content acquisition request is unrelated to the preset knowledge base, the target content can also be generated based on the second data acquisition result to provide response content for the above content acquisition request. This improves the efficiency for users to obtain relevant data from different data sources.

[0020] Please refer to Figure 1 , Figure 1 Flowchart of the content generation method provided in this disclosure Figure 1 ,like Figure 1 As shown, the method includes the following steps:

[0021] S101: Receive the first data acquisition result and the second data acquisition result; wherein, the first data acquisition result is obtained by acquiring data based on knowledge content in a preset knowledge base in response to the content acquisition request, and the second data acquisition result is obtained by acquiring data based on a wide area network in response to the content acquisition request, or the second data acquisition result is generated by the content generation model according to the content acquisition request.

[0022] In this embodiment, the entity executing the content generation method can be a server (e.g., a server that provides services for a smart assistant application).

[0023] The content retrieval request can be sent from a user terminal, for example, a user sending a content retrieval request to the server using their terminal. Specifically, a smart assistant application can be running on the user terminal. Within this smart assistant application, the user can input the content retrieval request. The smart assistant application can then send the content retrieval request to the server.

[0024] After receiving a content retrieval request, the aforementioned server can retrieve data based on the knowledge content in the preset database to obtain the first data retrieval result.

[0025] In one implementation, the preset database here can be a knowledge base built based on data generated within the user group.

[0026] In some implementations, after receiving a content retrieval request, the server can also retrieve data from the wide area network in response to the content retrieval request, and obtain a second data retrieval result based on the retrieved data.

[0027] For example, the server initiates a search on the wide area network in response to a content retrieval request, obtaining multiple candidate second data retrieval results; then, the multiple candidate second data retrieval results are sorted and filtered to obtain the second data retrieval result.

[0028] In some other implementations, after receiving a content retrieval request, the server may also send the content retrieval request to a content generation model, which will then generate a first data retrieval result based on the content retrieval request.

[0029] The aforementioned executing entity can receive a first data acquisition result based on data from a preset database, and a second data acquisition result based on data from a wide area network.

[0030] S102: Generate a prompt message including the first data acquisition result and the second data acquisition result; wherein, the prompt message also includes content generation rules; the content generation rules include the priority of referring to the first data acquisition result and referring to the second acquisition result when generating content.

[0031] Upon receiving the first and second data acquisition results, a prompt message can be generated. This prompt message may include the first and second data acquisition results, as well as the content generation rules. The content generation rules may include prioritizing the first and second data acquisition results when generating content.

[0032] As an example, the above content generation rule may include "when a first data acquisition result related to the content acquisition request exists in the preset database, the first data acquisition result shall be used first to generate the response content." Alternatively, the above content generation rule may also include "when no data related to the content acquisition request exists in the preset database, the second data acquisition result shall be used to generate the response content." That is, the first data acquisition result from the preset database has the highest priority.

[0033] It is understandable that the first data acquisition result and the second data acquisition result in the above prompt information are distinguishable. For example, a first label can be set for the first data acquisition result and a second label can be set for the second data acquisition result in the above prompt information. The first label and the second label are different. Another example is that a preset separator can be used to separate the first data acquisition result and the second data acquisition result. By setting a first label and a second label for the first data acquisition result and the second data acquisition result respectively, or by setting a preset separator between the first data acquisition result and the second data acquisition result, the content generation model can distinguish between the first data acquisition result and the second data acquisition result.

[0034] S103: Input the prompt information into the pre-trained content generation model, which then generates the target content based on the content generation rules, the first data acquisition result, and the second data acquisition result.

[0035] The content generation model here can be any learning model with language processing capabilities or content generation capabilities.

[0036] The aforementioned prompt includes the first data acquisition result, the second data acquisition result, and the content generation rules, and may also include a content acquisition request. Therefore, the aforementioned content generation model will generate the target content according to the content generation rules, based on the first data acquisition result and / or the second data acquisition result.

[0037] Since the above content generation rules indicate the priority of referring to the first data acquisition result and referring to the second data acquisition result, the above content generation model prioritizes using the data acquisition result indicated by the priority to generate the target content, provided that the above priority indicates which data acquisition result is preferred to generate the target content.

[0038] In one example, the content generation rule is as follows: the target content is generated using the first data retrieval result from the preset database first; if there is no data related to the content retrieval request in the preset database, the target content is generated using the second data retrieval result.

[0039] In this example, the content generation rule states that if the prompt message contains data related to the content retrieval request from a preset database—that is, if the first data retrieval result is not empty—the response content can be generated from the first data retrieval result. Therefore, when the first data retrieval result in the prompt message is not empty, the content generation model can generate the target content based on the first data retrieval result. In other words, when the prompt message includes both a non-empty first data retrieval result and a non-empty second data retrieval result, based on the above content generation rule, the content generation model can generate the target content based on the first data retrieval result.

[0040] If the first data acquisition result is empty, that is, there is no data related to the content acquisition request in the preset knowledge base, the content generation model can generate the target content based on the second data acquisition result according to the above prompt information.

[0041] In another example, the above content generation rule can also instruct that the second data acquisition result be used first to generate the target content, and when the second data acquisition result is empty, the first data acquisition result from the preset database be used to generate the target content.

[0042] According to the content generation rules provided in this example, the content generation model can use the second data acquisition result to generate the target content corresponding to the content acquisition request when the second data acquisition result is not empty; and can use the first data acquisition result to generate the target content when the second data acquisition result is empty.

[0043] As another example, the above prompt could also suggest first generating the target content based on the first data acquisition result, and then generating the target content based on the second data acquisition result, displaying the two target contents separately. The content generation model can then generate the first target content based on the first data acquisition result and the second target content based on the second data acquisition result, respectively, according to this prompt. It then outputs the first and second target contents separately, labeling the first target content as originating from a preset knowledge base and the second target content as originating from a wide area network or the content generation model itself.

[0044] As one implementation, the content generation model described above can send the target content to the execution entity, which then responds to the content retrieval request using the target content. For example, the execution entity can send the target content to the user terminal, enabling the user terminal to display the target content in a smart assistant application.

[0045] In this embodiment, upon receiving a content retrieval request, a first data retrieval result is obtained by retrieving data from a preset database based on the knowledge content in the request. A second data retrieval result is obtained by retrieving data from a wide area network (WAN) based on the request, or the second data retrieval result is generated by a content generation model. Then, a prompt message is generated including the first data retrieval result, the second data retrieval result, and content generation rules. These content generation rules include the priority of referencing the first and second data retrieval results when generating target content. Upon receiving the prompt message, the content generation model generates the target content based on the priority, the first data retrieval result, and the second data retrieval result. The priority set in the content generation rules can match the source of the response the user expects to receive. Therefore, the target content generated by the content generation model based on the prompt message matches the user's needs. Furthermore, based on the prompt message, the content generation model can also generate the target content from other data retrieval results if the data retrieval result indicated by the priority cannot generate the target content. Therefore, this reduces the need for the user to re-initiate the operation of retrieving response content from other data sources for the same content retrieval request, improving the efficiency of the user obtaining relevant content from different data sources.

[0046] In some embodiments of this example, the content generation method further includes the following steps:

[0047] First, multiple first data fragments are retrieved from a preset knowledge base based on the content retrieval request.

[0048] Secondly, the multiple first data fragments are sorted according to their relevance to the content retrieval request, and the sorting results are filtered to obtain the first data retrieval results.

[0049] In these embodiments, content retrieval requests can be matched against a preset database to obtain multiple first data fragments.

[0050] As one implementation, the above retrieves multiple first data fragments from a preset knowledge base based on the content retrieval request, including:

[0051] The content retrieval request is matched with multiple third data fragments in a preset knowledge base, and multiple first data fragments are obtained from the multiple third data fragments based on the matching results; wherein, the multiple third data fragments are obtained by pre-segmenting the data in the preset knowledge base.

[0052] Specifically, the data in the preset knowledge base can be segmented in the following ways to obtain multiple third data segments: rule-based segmentation, segmentation of data in the knowledge base based on semantic segmentation models, etc.

[0053] The rule-based segmentation described above can include: setting segmentation rules, and then segmenting data in a preset knowledge base according to the segmentation rules. These segmentation rules include segmentation based on paragraphs, sentences, topics, or fixed lengths.

[0054] The above segmentation is based on paragraphs. For the text in the preset knowledge base, it is segmented by paragraph, and each paragraph is treated as an independent third data segment.

[0055] The above sentence-based segmentation method segments the text in the pre-defined knowledge base into sentences, with each sentence serving as an independent third-party data segment.

[0056] The above-mentioned topic-based segmentation uses topic models or clustering algorithms to segment text by topic. Data under each topic is treated as a third data segment.

[0057] Fixed-length segmentation divides the text into segments of a fixed length (e.g., number of characters, number of words). Each segment of fixed length comprises a third data fragment.

[0058] For multiple third-party data segments in a pre-defined knowledge base, each segment can be vectorized to obtain its corresponding vector. Upon receiving a content retrieval request, the same vectorization method can be used to vectorize the request, resulting in a content retrieval request vector. Then, the similarity between the content retrieval request and each third-party data segment can be calculated using the content retrieval request vector and the vectors of each segment. The third-party data segments with a similarity greater than a pre-defined similarity threshold are designated as the first data segments. Alternatively, multiple third-party data segments can be sorted in descending order of similarity to the content retrieval request, and the segments with sort numbers less than or equal to N are designated as the first data segments.

[0059] After obtaining multiple first data fragments, these fragments can be sorted according to their relevance to the content retrieval request. This relevance can be characterized by the similarity between the first data fragment and the content retrieval request. The multiple first data fragments are then sorted in descending order of their similarity to the content retrieval request, resulting in a sorting result. From this sorting result, the first data fragments with sort numbers less than a preset threshold (the preset threshold could be, for example, K, where K is an integer greater than 0, and the specific value of K can be determined based on the application scenario; illustratively, K is 10) can be extracted as second data fragments. In other words, the top K first data fragments from the sorting result are selected as K second data fragments.

[0060] To quickly find the first data acquisition result, after obtaining multiple second data segments, these segments need to be filtered. Alternatively, the multiple second data segments can be deduplicated first, and then filtered to obtain the first data acquisition result.

[0061] In some implementations of this embodiment, the above-described sorting of multiple first data segments according to their relevance to the content retrieval request includes:

[0062] First, multiple first data fragments and content retrieval requests are input into a pre-trained ranking model, and the relevance scores between each first data fragment and content retrieval request are output by the ranking model.

[0063] Secondly, the multiple first data segments are sorted according to their relevance scores.

[0064] In these embodiments, the ranking model described above can be any model with text processing capabilities. After the content retrieval request and multiple first data fragments are input into the ranking model, the ranking model outputs a relevance score between the multiple first data fragments and the content retrieval request. Then, the multiple first data fragments are ranked in descending order of the relevance scores to obtain the ranking result.

[0065] Specifically, the above ranking model can be trained through the following steps:

[0066] First, training samples are constructed. For each third data segment within a predefined knowledge base, a sample content retrieval request is created. The constructed sample content retrieval request and the current third data segment are treated as positive samples, with a relevance score of the first relevance score. A score indicating the relevance between the sample content retrieval request and the current third data segment is also provided. Then, multiple other third data segments are recalled from the predefined knowledge base based on their text similarity to the sample content retrieval request. These multiple other third data segments are treated as negative samples corresponding to the sample content retrieval request, with a relevance score of the second relevance score for each negative sample. This process yields one positive sample and multiple negative samples corresponding to each of the multiple sample content retrieval requests.

[0067] The first relevance score here can be, for example, 1, and the second relevance score can be, for example, 0.

[0068] Secondly, for each training sample among multiple training samples, the third data segment in the training sample and the sample content acquisition request are input into the ranking model, and the relevance score corresponding to the training sample is used as the output to train the ranking model.

[0069] During the training process described above, a loss function can be selected, such as a contrastive loss function or a ternary loss function.

[0070] After training with multiple training samples, a trained ranking model can be obtained. This trained ranking model, given an input content retrieval request and a data fragment, can output a relevance score between the data fragment and the content retrieval request.

[0071] In this embodiment, the correlation score between the content retrieval request and each first data segment is determined by the above sorting model, and then the multiple first data segments are sorted according to the above correlation score. This can quickly complete the sorting of multiple first segments and facilitate the rapid acquisition of first data results.

[0072] In one implementation of these embodiments, the above-described filtering of multiple second data segments to obtain a first data acquisition result includes:

[0073] Filter multiple second data segments based on one or more of the following:

[0074] Timeliness, authority, and relevance to the content retrieval request.

[0075] As an application scenario, multiple second data segments can be filtered based on timeliness. For example, a timeliness score can be calculated based on the time difference between the timestamps of multiple second data segments and the target time when the content retrieval request is received. The larger the time difference, the lower the timeliness score of the historical second data segment; the smaller the time difference, the higher the timeliness score. A timeliness threshold can be set, and second data segments with a timeliness score greater than the threshold can be regarded as the first data retrieval result.

[0076] As an application scenario, multiple second data segments can be filtered based on their authority. Authority metrics can be predefined, including but not limited to: data source and citation count. Data source includes the credibility and reputation of the data source. Citation count includes the number or frequency of times the data is cited. Then, scoring rules for the authority metrics are defined. Taking data source as an example, a second data segment from data source A has a score of 1; a second data segment from data source B has a score of 0.9; a second data segment from data source C has a score of 0.5, and so on. For different authority metrics, the weighted sum of the scores for each authority metric can be calculated to obtain the authority score for each second data segment. A preset authority threshold can be set, and second data segments with an authority score greater than the preset authority threshold are considered as the first data acquisition result.

[0077] The relevance of the second data segment to the content retrieval request can be characterized by the similarity between the two. A higher similarity between the second data segment and the content retrieval request corresponds to a higher relevance score, while a lower similarity corresponds to a lower relevance score. A preset relevance threshold can be set, and second data segments with relevance scores greater than the preset threshold can be considered as the first data retrieval result.

[0078] In one application scenario, multiple second data points can be filtered by comprehensively considering timeliness, authority, and relevance. Specifically, for each second data segment, the timeliness score, authority score, and relevance score corresponding to that second data segment are weighted and summed, and the sum is used as the comprehensive score for that second data segment. A comprehensive threshold can be set, and second data segments with a comprehensive score greater than the threshold are used as the first data points.

[0079] The first data acquisition result obtained through the above sorting and filtering has high accuracy and reliability, and the target content generated from the first data acquisition result has high quality.

[0080] Please refer to Figure 2 , Figure 2 Illustrative flow of the content generation method provided in this disclosure Figure 2 ,like Figure 2 As shown, the method includes the following steps:

[0081] S201: Parse the response to the content retrieval request to determine whether external data needs to be retrieved. External data refers to data retrieved from outside the content generation model to assist the content generation model in generating the target content.

[0082] In this embodiment, the entity executing the content generation method can be a server (e.g., a server that provides services for a smart assistant application).

[0083] The request to obtain the above content may be sent by the user terminal, for example.

[0084] The above-mentioned server can be pre-configured with parsing rules, which are used to determine whether external data needs to be retrieved in response to a content retrieval request.

[0085] For example, the parsing rules mentioned above could include: if the content retrieval request is for weather information, then there is no need to retrieve external data; if the content retrieval request is for information about historical figures, then there is no need to retrieve external data, etc. It is understood that the parsing rules can also include other rules, which can be determined based on the data range covered by the content generation model's own database.

[0086] Upon receiving a content retrieval request, semantic understanding can be performed on the request. Based on the semantic understanding result and the aforementioned parsing rules, it can be determined whether external data needs to be retrieved. For example, if the content retrieval request is "What is the weather trend this week?", the semantic understanding result indicates that the request is for obtaining weather information for this week, so external data retrieval is not required. If the content retrieval request is "How to pick up an office computer?", the semantic understanding result indicates that the request is related to picking up an office computer. This semantic understanding result is matched against the parsing rules; if no match is found, it is determined that external data needs to be retrieved.

[0087] In some embodiments, step S201 includes:

[0088] The content retrieval request is input into the prediction model, which outputs the prediction result; the prediction result indicates whether the response to the content retrieval request requires the retrieval of external data.

[0089] For example, a prediction result of "0" indicates that no external data is needed. A prediction result of "1" indicates that external data is needed.

[0090] The above prediction model can be trained based on the following steps:

[0091] First, from the historical content retrieval requests, multiple first historical content retrieval requests that require external data to respond to historical content requests are identified as positive samples, with their corresponding first label being "1". From the historical content retrieval requests, multiple second historical content retrieval requests that do not require external data to respond to historical content requests are identified as negative samples, with their corresponding second label being "0".

[0092] Then, the first historical content retrieval request in the positive sample is used as input, and the first label is used as the output ground truth. The second historical content retrieval request in the negative sample is used as input, and the second label is used as the output. The prediction model is trained to obtain the trained prediction model.

[0093] In these embodiments, predictive models are used to determine whether external data is needed to respond to content retrieval requests, providing a rapid conclusion. This improves the efficiency of target content generation.

[0094] S202: In response to the parsing result indicating that external data needs to be obtained, obtain the first data acquisition result and the second data acquisition result.

[0095] If the parsing result in step S201 indicates that external data needs to be obtained, then the steps of obtaining the first data acquisition result and obtaining the second data acquisition result are executed.

[0096] For example, multiple first data fragments are retrieved from a preset knowledge base based on a content retrieval request; the multiple first data fragments are sorted according to their relevance to the content retrieval request, and the sorting results are filtered to obtain the first data retrieval result.

[0097] In some embodiments, obtaining the second data result as described above includes the following steps:

[0098] First, a search was initiated on the wide area network in response to the content retrieval request, resulting in multiple fourth data fragments.

[0099] Secondly, the multiple fourth data segments are sorted according to their relevance to the content retrieval request, and the sorting results are filtered to obtain the second data retrieval result.

[0100] In these embodiments, a search can be performed on a wide area network in response to a content retrieval request to obtain multiple fourth data segments whose similarity to the content retrieval request meets a preset similarity requirement.

[0101] Then, the data is sorted in descending order of its relevance to the content retrieval request (represented by similarity). The top L (L is an integer greater than or equal to 1) first data segments in the sorted results are filtered (e.g., by deduplication, and by combining timeliness, authority, and relevance to the content retrieval request) to obtain the second data retrieval results.

[0102] The second data acquisition result obtained through the above steps has a high degree of matching with the content acquisition request, and the target content generated based on the second data acquisition result is also highly accurate.

[0103] The operations described above for obtaining the first data acquisition result and the second data acquisition result can be performed simultaneously.

[0104] S203: Receive the first data acquisition result and the second data acquisition result; wherein, the first data acquisition result is obtained by acquiring data based on knowledge content in a preset knowledge base for the content acquisition request, and the second data acquisition result is obtained by acquiring data based on a wide area network for the content acquisition request.

[0105] S204: Generate a prompt message including the first data acquisition result and the second data acquisition result; wherein, the prompt message also includes content generation rules; the content generation rules include the priority of referring to the first data acquisition result and referring to the second data acquisition result when generating target content.

[0106] S205: Input the prompt information into the pre-trained content generation model, which then generates the target content based on the content generation rules, the first data acquisition result, and the second data acquisition result.

[0107] For specific implementation details of steps S203 to S205 above, please refer to [reference needed]. Figure 1 The descriptions of steps S101 to S103 in the illustrated embodiment are not repeated here.

[0108] It is understandable that if the parsing result obtained in step S201 indicates that there is no need to obtain external data, the content acquisition request can be directly sent to the content generation model, and the content generation model can generate the target content based on the data in its own database.

[0109] In this embodiment, the system first analyzes whether external data needs to be acquired. If no external data is needed, the content generation model can quickly generate the target content. This reduces the amount of data processing required to acquire and process external data, thus improving the speed of generating the target content. When it is determined that external data needs to be acquired, the system obtains a first data acquisition result from a preset database and a second data acquisition result from a wide area network. The content generation model then generates the target content based on the first and second data acquisition results and their priorities as indicated in the prompt information. Since both the first and second data acquisition results are data fragments, the generated target content has a high degree of matching with user needs and can also improve the efficiency of users acquiring relevant data from different data sources.

[0110] exist Figure 1 and Figure 2 In some implementations of the illustrated embodiments, the content generation method further includes:

[0111] The target content is displayed on the page where the content retrieval request is made.

[0112] In these implementations, after the target content is generated, it can be displayed on the page where the content retrieval request is located.

[0113] Specifically, the aforementioned executing entity can send the target content as a response to the content retrieval request to the terminal device, which will then display the target content on the page where the content retrieval request is located.

[0114] In these implementations, by displaying the target content on the page where the content retrieval request is located, users can promptly view the target content that responds to the content retrieval request on the page.

[0115] exist Figure 1 and Figure 2 In some implementations of the illustrated embodiments, the content generation model is generated based on the following steps:

[0116] First, a first training sample pair is constructed, which includes a first sample prompt message as input and a first sample response content as output. The first sample prompt message includes a first sample content retrieval request, a first sample data fragment, a second sample data fragment, and a first content generation rule. The first sample data fragment is a data fragment from a preset knowledge base, and the first sample response content is generated from the first sample data fragment. The second sample data fragment is data related to the first sample content retrieval request from the wide area network. The first priority in the first content generation rule indicates that the response content is generated with priority reference to the first sample data fragment.

[0117] Secondly, a second training sample pair is constructed, which includes a second sample prompt message as input and a second sample response content as output. The second sample prompt message includes a second sample content retrieval request, a third sample data fragment, a fourth sample data fragment, and a second content generation rule. The third sample data fragment is a data fragment from a preset knowledge base, and the fourth sample data fragment is data related to the second sample content retrieval request from the wide area network. The second sample response content is generated from the fourth sample data fragment. The second priority in the second content generation rule indicates that the response content is generated with priority reference to the fourth sample data fragment.

[0118] Finally, the initial content generation model is trained using multiple first training sample pairs and / or multiple second training samples to obtain the trained content generation model.

[0119] Understandably, the initial content generation model can be a model with language processing capabilities. The initial content generation model can be fine-tuned using the first training sample pair and / or the second training sample pair to obtain a fine-tuned content generation model.

[0120] In some application scenarios, multiple first training samples are used to fine-tune the initial content generation model. The fine-tuned content generation model can have the ability to generate content based on the first priority in the prompt information and data fragments from a preset knowledge base to obtain the response content of the request.

[0121] In some application scenarios, multiple first training samples are used to fine-tune the initial content generation model. The fine-tuned content generation model can then generate content based on the second priority in the prompt information and data fragments from the wide area network to obtain the response content of the request.

[0122] During training, the loss value for each training iteration can be calculated based on a preset loss function, and the parameters of the content generation model can be adjusted accordingly. The preset loss function could be, for example, the cross-entropy loss function or the mean squared error loss function.

[0123] After the above training, a trained content generation model can be obtained. This model can accurately generate target content based on the aforementioned content generation rules and the first and second data acquisition results in the prompt information.

[0124] Corresponding to the content generation method in the above embodiments, Figure 3 This is a structural block diagram of a content generation apparatus provided for embodiments of this disclosure. For ease of explanation, only the parts relevant to embodiments of this disclosure are shown. (Refer to...) Figure 3 The device 30 includes: a receiving unit 301, a first generating unit 302, and a second generating unit 303. Wherein,

[0125] The receiving unit 301 is used to receive a first data acquisition result and a second data acquisition result; wherein, the first data acquisition result is obtained by acquiring data based on knowledge content in a preset knowledge base in response to the content acquisition request, and the second data acquisition result is obtained by acquiring data based on a wide area network in response to the content acquisition request, or the second data acquisition result is generated by a content generation model according to the content acquisition request.

[0126] The first generation unit 302 is used to generate a prompt message including a first data acquisition result and a second data acquisition result; wherein, the prompt message also includes content generation rules; the content generation rules include the priority of referring to the first data acquisition result and referring to the second data acquisition result when generating target content;

[0127] The second generation unit 303 is used to input the prompt information into the pre-trained content generation model, and the content generation model generates the target content according to the content generation rules, the first data acquisition result and the second data acquisition result.

[0128] In some embodiments, the device 30 further includes a first data acquisition unit (not shown in the figures), the first data acquisition unit being used for:

[0129] Based on the content retrieval request, retrieve multiple first data fragments from the preset knowledge base;

[0130] Multiple first data fragments are sorted according to their relevance to the content retrieval request, and the sorting results are filtered to obtain the first data retrieval result.

[0131] In some embodiments, the first data acquisition unit is further configured to:

[0132] Sort the multiple first data fragments in descending order of their relevance to the content retrieval request;

[0133] Multiple first data segments with sort numbers less than a first preset threshold are used as second data segments. These multiple second data segments are then filtered to obtain the first data acquisition result.

[0134] In some embodiments, the first data acquisition unit is further configured to:

[0135] Multiple first data fragments and content retrieval requests are input into a pre-trained ranking model, and the ranking model outputs a ranking of the multiple first data fragments.

[0136] In some embodiments, the first data acquisition unit is further configured to:

[0137] Filter multiple second data segments based on one or more of the following:

[0138] Timeliness, authority, and relevance to the content retrieval request.

[0139] In some embodiments, the first data acquisition unit is further configured to:

[0140] The content retrieval request is matched with multiple third data fragments in a preset knowledge base, and multiple first data fragments are obtained from the multiple third data fragments based on the matching results; wherein, the multiple third data fragments are obtained by pre-segmenting the data in the preset knowledge base.

[0141] In some embodiments, the apparatus 30 further includes a parsing unit (not shown in the figures), the parsing unit being used for:

[0142] The response to the content retrieval request is analyzed to determine whether external data needs to be retrieved. External data refers to data obtained from outside the content generation model to assist the content generation model in generating the target content.

[0143] In response to the parsing result indicating the need to obtain external data, the first data acquisition result and the second data acquisition result are obtained.

[0144] In some embodiments, the parsing unit is further configured to:

[0145] The content retrieval request is input into the prediction model, which outputs a prediction result indicating whether external data needs to be retrieved in response to the content retrieval request.

[0146] In some embodiments, the device 30 further includes a second data acquisition unit (not shown in the figures), the second data acquisition unit being used for:

[0147] A search was initiated on the wide area network in response to the content retrieval request, resulting in multiple candidate second data retrieval results;

[0148] The multiple candidate second data acquisition results are sorted and filtered to obtain the second data acquisition result.

[0149] In some embodiments, the device 30 further includes a display unit (not shown), which is used to display the target content on the page where the content retrieval request is located.

[0150] To implement the above embodiments, this disclosure also provides an electronic device.

[0151] refer to Figure 4The diagram illustrates a structural schematic of an electronic device 400 suitable for implementing embodiments of the present disclosure. The electronic device 400 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0152] like Figure 4 As shown, electronic device 400 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of electronic device 400. The processing unit 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.

[0153] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0154] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program (computer execution instructions) is executed by processing device 401, the functions defined above in the methods of embodiments of this disclosure are performed.

[0155] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0156] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0157] The aforementioned computer-readable medium carries one or more programs (computer execution instructions) that, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.

[0158] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0160] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0161] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0162] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0163] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0164] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0165] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A content generation method, comprising: Receive a first data acquisition result and a second data acquisition result; wherein, the first data acquisition result is obtained by acquiring data based on knowledge content in a preset knowledge base in response to the content acquisition request, and the second data acquisition result is obtained by acquiring data based on a wide area network in response to the content acquisition request, or the second data acquisition result is generated by a content generation model according to the content acquisition request; Generate a prompt message including the first data acquisition result and the second data acquisition result; wherein, the prompt message further includes content generation rules; the content generation rules include the priority of referring to the first data acquisition result and referring to the second data acquisition result when generating target content; The prompt information is input into a pre-trained content generation model, which then generates target content based on the content generation rules, the first data acquisition result, and the second data acquisition result.

2. The method according to claim 1, characterized in that, The method further includes: According to the content retrieval request, multiple first data fragments are retrieved from the preset knowledge base; The plurality of first data fragments are sorted according to their relevance to the content acquisition request, and the sorting results are filtered to obtain the first data acquisition result.

3. The method according to claim 2, characterized in that, The step of sorting the plurality of first data segments according to their relevance to the content acquisition request, filtering the sorting results, and obtaining the first data acquisition result includes: The plurality of first data fragments are sorted in descending order of their relevance to the content retrieval request; Multiple first data segments with sort numbers less than a first preset threshold are used as second data segments. These multiple second data segments are then filtered to obtain the first data acquisition result.

4. The method according to claim 3, characterized in that, The step of filtering multiple second data segments to obtain the first data acquisition result includes: Filter multiple second data segments based on one or more of the following: Timeliness, authority, and relevance to the content retrieval request.

5. The method according to claim 2, characterized in that, The step of retrieving multiple first data fragments from a preset knowledge base based on a content retrieval request includes: The content retrieval request is matched with multiple third data fragments in the preset knowledge base, and the multiple first data fragments are obtained from the multiple third data fragments according to the matching results; wherein, the multiple third data fragments are obtained by pre-segmenting the data in the preset knowledge base.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: The response to the content retrieval request is analyzed to determine whether external data needs to be obtained. The external data refers to data obtained from outside the content generation model to assist the content generation model in generating the target content. In response to the parsing result indicating that external data needs to be obtained, the first data acquisition result and the second data acquisition result are obtained.

7. The method according to claim 6, characterized in that, The parsing of the response to the content retrieval request, including whether external data needs to be obtained, includes: The content retrieval request is input into a prediction model, which outputs a prediction result, wherein the prediction result indicates whether external data needs to be retrieved in response to the content retrieval request.

8. The method according to any one of claims 1-5, characterized in that, The method further includes: In response to the content retrieval request, a search was initiated on the wide area network, resulting in multiple candidate second data retrieval results; The multiple candidate second data acquisition results are sorted and filtered to obtain the second data acquisition result.

9. The method according to any one of claims 1-5, characterized in that, Also includes: The target content is displayed on the page where the content retrieval request is located.

10. A content generation apparatus, comprising: A receiving unit is configured to receive a first data acquisition result and a second data acquisition result; wherein the first data acquisition result is obtained by acquiring data based on knowledge content in a preset knowledge base in response to a content acquisition request, and the second data acquisition result is obtained by acquiring data based on a wide area network in response to the content acquisition request, or the second data acquisition result is generated by a content generation model according to the content acquisition request; The first generation unit is used to generate a prompt message including the first data acquisition result and the second data acquisition result; wherein, the prompt message further includes content generation rules; the content generation rules include the priority of referring to the first data acquisition result and referring to the second data acquisition result when generating target content; The second generation unit is used to input the prompt information into a pre-trained content generation model, and the content generation model generates target content according to the content generation rules, the first data acquisition result, and the second data acquisition result.

11. An electronic device, characterized in that, include: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1 to 9.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 9.