Literature adding method, device and equipment, computer readable medium and program product
By generating and displaying a sequence of candidate citation information during the text creation process, the problem of frequent interruptions in querying during the creation process in existing technologies is solved, realizing an efficient and continuous citation method and improving creation efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-04-03
AI Technical Summary
During the text creation process, existing technologies require interrupting the creation process to search for literature on a literature search page, resulting in low literature search efficiency and affecting the normal execution of the creation process.
By generating a sequence of candidate citation information in the target format, the candidate citation information is displayed on the text creation page in real time. It supports the selection and addition of relevant citations without interrupting the creation process, and uses a large language model and a text content association model to match and sort citations.
It improves the efficiency of citation, reduces the number of page switches, ensures the continuity and accuracy of the creation process, and improves the efficiency of obtaining and adding citation information.
Smart Images

Figure CN121787375A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to the field of artificial intelligence, and specifically to document addition methods, apparatus, devices, computer-readable media, and program products. Background Technology
[0002] Currently, in the process of text creation, relevant literature is often cited to verify the accuracy of the text content. Therefore, citing literature is crucial to ensuring the reliability of the created text. The usual method for citing literature is to search for and cite relevant documents on a literature search page after interrupting the creation process.
[0003] However, the inventors discovered that the following technical problems often arise when using the above method: The creative process needs to be interrupted to conduct further searches on the literature search page, which requires switching back and forth between pages, resulting in low efficiency in literature retrieval and affecting the normal execution of the creative process.
[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0006] Some embodiments of this disclosure provide methods, apparatus, devices, computer-readable media, and program products for adding documents to address the technical problems mentioned in the background section above.
[0007] In a first aspect, some embodiments of this disclosure provide a method for adding references, including: in response to receiving a reference citation request for a target created text during the text creation process, generating a sequence of candidate reference citation information in a target format based on the target created text; displaying the candidate reference sequence on a target page corresponding to the target created text for selecting candidate references; in response to receiving selection information for target candidate references in the candidate reference sequence, adding the target candidate references to the target created text to obtain reference-added text; and displaying the reference-added text on the target page.
[0008] Optionally, the above-mentioned generation of a candidate literature citation information sequence in the target format based on the target creative text includes: using a large language model to generate a keyword set and text summary corresponding to the target creative text; for each literature data in the literature database, performing the generation steps: determining the first matching information between the literature data and the keyword set, and determining the second matching information between the literature data and the text summary; generating literature matching information corresponding to the literature data based on the first matching information and the second matching information; selecting a first literature dataset from the literature database whose corresponding literature matching information meets the first matching condition; and generating a candidate literature citation information sequence in the target format based on the first literature dataset.
[0009] Optionally, generating the candidate citation information sequence in the target format based on the first document dataset includes: for each first document data in the first document dataset, performing a second generation step: inputting the text summary and the target creative text included in the first document data into a pre-trained text content association information generation model to obtain text content association information; generating text content association information based on the text content association information and the document matching information corresponding to the first document data; selecting a subset of document data whose corresponding text content association information satisfies the second matching condition from the first document dataset as the second document dataset; and generating the candidate citation information sequence in the target format based on the second document dataset.
[0010] Optionally, generating the candidate citation information sequence under the target format based on the second document dataset includes: using the large language model to generate the content domain type corresponding to the target creative text; determining the citation format corresponding to the content domain type as the target format; sorting each second document data in the second document dataset according to the text content association information corresponding to the second document data to obtain a second document data sequence; and generating the candidate citation information sequence under the target format based on the second document data sequence.
[0011] Optionally, generating the literature matching information corresponding to the literature data based on the first matching information and the second matching information includes: setting a first weight corresponding to the first matching information and a second weight corresponding to the second matching information based on the creative focus direction information corresponding to the target creative text; and performing weighted processing on the first matching information, the first weight, the second matching information, and the second weight to obtain the literature matching information.
[0012] Optionally, the literature data in the aforementioned literature database includes: text fragments in the literature text and corresponding fragment vectors; and the literature dataset corresponding to the aforementioned literature text is generated through the following steps: performing text block processing on the literature text to obtain a set of text fragments, wherein the character length corresponding to each text fragment is less than or equal to the target length; performing high-dimensional vector transformation on each text fragment in the aforementioned text fragment set to generate a fragment vector, thereby obtaining a set of fragment vectors; and combining the aforementioned text fragment set and the aforementioned fragment vector set accordingly to obtain the literature dataset.
[0013] Secondly, some embodiments of this disclosure provide a document addition apparatus, including: a generation unit configured to, in response to receiving a document citation request for a target created text during text creation, generate a sequence of candidate document citation information in a target format based on the target created text; a first display unit configured to display the candidate document sequence on a target page corresponding to the target created text for selecting candidate documents; an addition unit configured to, in response to receiving selection information for a target candidate document in the candidate document sequence, add the target candidate document to the target created text to obtain document-added text; and a second display unit configured to display the document-added text on the target page.
[0014] Optionally, the generation unit can be configured to: use a large language model to generate a keyword set and text summary corresponding to the target creative text; for each document data in the document database, perform the following generation steps: determine the first matching information between the document data and the keyword set, and determine the second matching information between the document data and the text summary; generate document matching information corresponding to the document data based on the first matching information and the second matching information; select a first document dataset from the document database whose corresponding document matching information satisfies the first matching condition; and generate a sequence of candidate document citation information in the target format based on the first document dataset.
[0015] Optionally, the generation unit can be configured to: for each first document data in the first document dataset, perform a second generation step: input the text summary and the target creative text included in the first document data into a pre-trained text content association information generation model to obtain text content association information; generate text content association information based on the text content association information and the document matching information corresponding to the first document data; select a subset of document data whose corresponding text content association information satisfies the second matching condition from the first document dataset as the second document dataset; and generate a sequence of candidate document citation information in the target format based on the second document dataset.
[0016] Optionally, the generation unit can be configured to: use the above-mentioned large language model to generate the content domain type corresponding to the above-mentioned target creative text; determine the literature citation format corresponding to the above-mentioned content domain type as the above-mentioned target format; sort each of the second literature data in the above-mentioned second literature dataset according to the text content association information corresponding to the second literature data to obtain the second literature data sequence; and generate the candidate literature citation information sequence under the above-mentioned target format according to the above-mentioned second literature data sequence.
[0017] Optionally, the generation unit can be configured to: set a first weight corresponding to the first matching information and a second weight corresponding to the second matching information based on the creation focus information corresponding to the target creative text; and perform weighted processing on the first matching information, the first weight, the second matching information and the second weight to obtain the document matching information.
[0018] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, such that when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.
[0019] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method as described in any implementation of the first aspect.
[0020] Fifthly, some embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the implementations of the first aspect above.
[0021] The above-described embodiments of this disclosure have the following beneficial effects: The document addition method of some embodiments of this disclosure allows users to efficiently cite accurate documents related to the content of their text creation. Specifically, the reason why citing relevant documents is cumbersome and inefficient is that it requires interrupting the creation process to search for further information on a document retrieval page, requiring frequent page switching, resulting in low document retrieval efficiency and affecting the normal execution of the creation process. Based on this, the document addition method of some embodiments of this disclosure firstly, in response to receiving a document citation request for the target text during the text creation process, accurately generates a sequence of candidate document citation information in the target format based on the target text. Here, during the creation of the target text, real-time document citation is supported, enabling the display of multiple selectable candidate document citation information to the creator without interrupting the creation process, thus improving the efficiency of the creator in obtaining candidate document citation information. Then, the sequence of candidate document citation information in the target format is displayed on the target page corresponding to the target text, allowing the creator to select candidate document citation information that better matches the semantic content. In addition, by directly displaying a sequence of matching candidate citation information on the target page corresponding to the target creative text, the author can directly select the relevant citation without switching pages, significantly improving the efficiency of citation. Furthermore, in response to receiving selection information for the target candidate citation information from the aforementioned sequence, the candidate citation information is added to the target creative text, resulting in the citation-added text. Finally, the citation-added text is displayed on the target page for the author to view. In summary, during the creation process and upon receiving a request from the author to cite citations, accurate citation information can be added to the creative text without interrupting the creation process or switching search pages, greatly improving the efficiency of citation. Attached Figure Description
[0022] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0023] Figures 1-2 This is a schematic diagram illustrating an application scenario of a document addition method according to some embodiments of this disclosure; Figure 3 This is a flowchart of some embodiments of the method of adding according to the present disclosure; Figure 4 This is a flowchart of some other embodiments of the method of adding according to the present disclosure; Figure 5 This is a schematic diagram of the structure of some embodiments of the adding device according to the present disclosure; Figure 6 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0025] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0026] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0027] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0028] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0029] Before performing any of the operations involving the collection, storage, or use of user personal information (such as target-created text) as disclosed in this disclosure, the relevant organizations or individuals shall fulfill their obligations, including conducting personal information security impact assessments, informing personal information subjects, and obtaining prior authorization and consent from personal information subjects.
[0030] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0031] Figures 1-2 This is a schematic diagram illustrating an application scenario of a document addition method according to some embodiments of this disclosure.
[0032] exist Figures 1-2In this application scenario, firstly, in response to receiving a citation request for the target text during the text creation process, the electronic device 101 can generate a sequence of candidate citation information in the target format based on the target text. In this application scenario, the target text could be "Opinions on the Implementation of the Comprehensive Promotion of the Rural Revitalization Strategy, in order to thoroughly implement..." In accordance with the spirit and deployment of the Central Rural Work Conference, to comprehensively promote the implementation of the rural revitalization strategy and accelerate the process of agricultural and rural modernization, relevant opinions are put forward in light of the actual situation in this region. A document citation request can be triggered by clicking the document citation control in operation 1. The candidate document citation information sequence can be {"According to Article XXX of the XXX Law", "According to Article YYY of the YYY Law", "According to Article HHH of the HHH Law"}. Then, the electronic device 101 can display the above candidate document citation information sequence on the target page corresponding to the above target creation text to select the candidate document citation information. Next, in response to receiving the selection information for the target candidate document citation information in the above candidate document citation information sequence, the electronic device 101 can add the above candidate document citation information to the above target creation text to obtain the document added text. In this application scenario, the selected target candidate document citation information is "According to Article XXX of the XXX Law". The selection information can be triggered by operation 2. Finally, the electronic device 101 can display the above document added text on the above target page. In this application scenario, the document added text can be "Opinions on the Implementation of the Comprehensive Promotion of the Rural Revitalization Strategy, in order to thoroughly implement..." In accordance with the spirit and arrangements of the Central Rural Work Conference, to comprehensively promote the implementation of the rural revitalization strategy and accelerate the modernization of agriculture and rural areas, and in light of the actual conditions of this region, relevant opinions are put forward {based on Article XXX of the XXX Law}.
[0033] It should be noted that the aforementioned electronic device 101 can be either hardware or software. When the electronic device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. When the electronic device is software, it can be installed in the hardware devices listed above. It can be implemented as, for example, multiple software programs or software modules used to provide distributed services, or as a single software program or software module. No specific limitations are made here.
[0034] It should be understood that Figures 1-2 The number of electronic devices shown is merely illustrative. Any number of electronic devices can be used depending on the implementation requirements.
[0035] Continue to refer to Figure 3 The diagram illustrates a flow 300 of some embodiments of a document addition method according to the present disclosure. This document addition method includes the following steps: Step 301: In response to receiving a citation request for the target text during the text creation process, generate a sequence of candidate citation information in the target format based on the target text.
[0036] In some embodiments, in response to receiving citation information for the target text during the text creation process, the execution entity of the aforementioned citation addition method (e.g., Figure 1The electronic device 101 shown can generate a sequence of candidate citation information in a target format based on the aforementioned target creative text. The text creation process can be the process of the creation object creating the text. That is, the target creative text is text created in real-time during the creation process that has a need for citation. For example, in the process of a doctor writing a paper, the paper has a need for citation. Another example is in the process of writing legal documents, there is a need to cite legal provisions. In the process of writing public relations documents, there is a need to cite relevant clauses. The creation object can be the object creating the aforementioned target creative text. For example, the creation object can be a manuscript writer or a document writer. A citation request can be a request to add citation information to the target creative text. In practice, the target creative text is the text that the creation object is creating on the target page. The target page is a page that supports functions such as text creation, citation information display, and citation navigation. A citation request can be a request triggered on the target page for the target creative text. For example, actively clicking the citation button for the aforementioned target creative text on the target page will trigger the automatic generation of citation request information. For example, actively entering a citation command for the target creative text on the target page will trigger the automatic generation of citation request information. The target format can be one that facilitates embedding the citation content into the target creative text. For example, the target format could be "Citing Article XX of the XX Law: XX content". The sequence order of the candidate citation information sequence can represent the order of similarity between the semantic content of the candidate citation information and the text content of the target creative text. Candidate citation information can be citation information that is yet to be determined as the citation content of the target creative text. For example, the candidate citation information sequence could be {first candidate citation information, second candidate citation information, third candidate citation information, fourth candidate citation information}. The semantic fit between the first candidate citation information and the target creative text is higher than that between the second candidate citation information and the target creative text. The semantic fit between the second candidate citation information and the target creative text is higher than that between the third candidate citation information and the target creative text. Further details will not be elaborated further. Semantic fit can represent semantic relationships. For example, the citation information for candidate references could be "in accordance with Article XX of the XX Law" or "based on Article XX of the XX Law".
[0037] As an example, firstly, the aforementioned execution entity can generate prompts representing the generation of a sequence of candidate citation information based on the target text. Then, the prompts are input into a Large Language Model (LLM) to obtain the sequence of candidate citation information.
[0038] Step 302: Display the sequence of candidate literature citation information on the target page corresponding to the target creative text to select candidate literature citation information.
[0039] In some embodiments, the aforementioned executing entity may display the sequence of candidate citation information on the target page corresponding to the target authored text, so as to select candidate citation information. In practice, candidate citation information can be selected by clicking on the target page.
[0040] Step 303: In response to receiving selection information for the target candidate literature citation information in the above candidate literature citation information sequence, the above candidate literature citation information is added to the above target creative text to obtain the literature added text.
[0041] In some embodiments, in response to receiving selection information for target candidate citation information in the aforementioned candidate citation information sequence, the executing entity can add the candidate citation information to the target creative text to obtain citation-added text. The target candidate citation information can be the candidate citation information selected from the candidate citation information sequence displayed on the target page. In practice, the selection of candidate citation information can be achieved through touch. The selection information can be the operation information corresponding to the selection operation.
[0042] As an example, the aforementioned implementing entity can add candidate citation information to the last position of the target creative text to obtain the citation-added text.
[0043] Step 304: Display the added text of the above-mentioned document on the target page.
[0044] In some embodiments, the aforementioned implementing entity may display the aforementioned document with added text on the aforementioned target page.
[0045] The above-described embodiments of this disclosure have the following beneficial effects: The document addition method of some embodiments of this disclosure allows users to efficiently cite accurate documents related to the content of their text creation. Specifically, the reason why citing relevant documents is cumbersome and inefficient is that it requires interrupting the creation process to search for further information on a document retrieval page, requiring frequent page switching, resulting in low document retrieval efficiency and affecting the normal execution of the creation process. Based on this, the document addition method of some embodiments of this disclosure firstly, in response to receiving a document citation request for the target text during the text creation process, accurately generates a sequence of candidate document citation information in the target format based on the target text. Here, during the creation of the target text, real-time document citation is supported, enabling the display of multiple selectable candidate document citation information to the creator without interrupting the creation process, thus improving the efficiency of the creator in obtaining candidate document citation information. Then, the sequence of candidate document citation information in the target format is displayed on the target page corresponding to the target text, allowing the creator to select candidate document citation information that better matches the semantic content. In addition, by directly displaying a sequence of matching candidate citation information on the target page corresponding to the target creative text, the author can directly select the relevant citation without switching pages, significantly improving the efficiency of citation. Furthermore, in response to receiving selection information for the target candidate citation information from the aforementioned sequence, the candidate citation information is added to the target creative text, resulting in the citation-added text. Finally, the citation-added text is displayed on the target page for the author to view. In summary, during the creation process and upon receiving a request from the author to cite citations, accurate citation information can be added to the creative text without interrupting the creation process or switching search pages, greatly improving the efficiency of citation.
[0046] Further reference Figure 4 The following diagram illustrates a flow 400 of another embodiment of the document addition method according to the present disclosure. The document addition method includes the following steps: Step 401: In response to receiving a citation request for the target text during the text creation process, the big language model is used to generate a keyword set and text abstract corresponding to the target text.
[0047] In some embodiments, in response to receiving a citation request for the target text during the text creation process, the executing entity (e.g. Figure 1The electronic device 101 shown can utilize a large language model to generate a keyword set and a text summary corresponding to the target creative text. The keywords can be key terms in the target creative text. The text summary can be an overview of the overall text content corresponding to the target creative text.
[0048] As an example, firstly, the aforementioned execution entity can generate model prompts for the keyword set and text summary corresponding to the target creative text. Then, the model prompts are input into a large language model to obtain the keyword set and text summary corresponding to the target creative text.
[0049] Step 402: For each document in the literature database, perform the generation steps: Step 4021: Determine the first matching information between the above-mentioned document data and the above-mentioned keyword set, and determine the second matching information between the above-mentioned document data and the above-mentioned text abstract.
[0050] In some embodiments, the executing entity may determine first matching information between the aforementioned document data and the aforementioned keyword set, and second matching information between the aforementioned document data and the aforementioned text abstract. The document database may be a database storing various document data. Document data may be a citationable document. The document database may be pre-constructed. The first matching information may be in numerical form or in tag form. The first matching information can characterize the semantic association between the text content corresponding to the document data and the keyword set. When the first matching information is in numerical form, it can be a value between 0 and 1; the higher the value, the closer the semantics between the text content corresponding to the document data and the keyword set. The first matching information can characterize the semantic similarity between the text content corresponding to the document data and the text content corresponding to the text abstract. For a detailed explanation, please refer to the relevant explanation of the first matching information.
[0051] As an example, using a high-dimensional vector transformation model, the aforementioned execution entity can determine the document data vector corresponding to the document data, and the keyword vector corresponding to the keyword set. Then, the cosine similarity between the document data vector and the keyword vector is determined as the first matching information. The high-dimensional vector transformation model can be a deep learning model that converts text into high-dimensional vectors. Similarly, cosine similarity can be used to determine the second matching information between the document data and the text summary.
[0052] Here, by determining the matching information between document data and keyword sets (keywords are generally short texts), similarity matching of key texts within short texts can be achieved. By determining the matching information between text data and text abstracts (text abstracts are the core content corresponding to the target text), similarity matching of the core semantics corresponding to the core content and document data can be achieved. Therefore, through similarity matching within short texts and similarity matching within core semantics, accurate retrieval of document data can be achieved from multiple dimensions.
[0053] In addition, considering that the target creative text has a large number of characters, that is, the target creative text is generally long, direct vectorization will introduce noise and have a large computational load. Therefore, we summarize the text abstract corresponding to the target creative text, extract the core semantics, make the retrieval of literature data more focused, and greatly reduce the computational load of similarity calculation.
[0054] In some optional implementations of certain embodiments, the literature data in the aforementioned literature database includes: text fragments within the literature text and corresponding fragment vectors. The literature text can be the complete text corresponding to a specific document. For example, the literature text might be an entire paper. A text fragment can be a section of text content within the literature text. For example, a text fragment might be a paragraph within an entire paper. The fragment vector can represent the semantic information of the content corresponding to the text fragment. The fragment vector can be obtained by inputting the text fragment into a high-dimensional vector transformation model. The literature text can be divided into multiple text fragments, each with corresponding literature data. Therefore, the literature text corresponds to the literature dataset.
[0055] Optionally, the document dataset corresponding to the above document text is generated through the following steps: The first step is to segment the document text into text fragments, resulting in a set of text segments. Each text fragment has a character length less than or equal to a target length. The target length can be preset and represents the maximum character length. For example, the target length could be "512".
[0056] As an example, the aforementioned implementing entity can segment the document text into blocks based on semantic segmentation rules defined by the content domain type, resulting in a set of text fragments. For instance, if the content domain type corresponds to a policy document, the semantic segmentation rule could be to segment the text using "chapter-section-point" as the segmentation identifier. As another example, if the content domain type corresponds to a legal provision, the semantic segmentation rule could be to segment the text according to "Article X, Clause X," etc. If the content domain type corresponds to an academic paper, the semantic segmentation rule could be to segment the text according to "abstract, framework, method, experiment," etc. In practice, semantic segmentation rules are not fixed and can be customized according to the different content within the content domain type.
[0057] The second step is to perform high-dimensional vector transformation on each text fragment in the above text fragment set to generate fragment vectors, thus obtaining a fragment vector set.
[0058] As an example, the aforementioned execution entity can input text fragments into a high-dimensional vector transformation model to obtain fragment vectors.
[0059] The third step is to combine the above text fragment set and the above fragment vector set accordingly to obtain the literature dataset.
[0060] Here, by segmenting the document text, highly structured documents such as policies and laws can be stored in blocks, allowing for precise matching of retrieval granularity down to the clause or chapter level, such as Article XX of Law XX. Furthermore, semantic representation can be optimized. Excessively long texts introduce more noise; segmented documents are shorter, semantically more focused, and easier for subsequent retrieval.
[0061] Step 4022: Generate literature matching information corresponding to the above-mentioned literature data based on the first matching information and the second matching information.
[0062] In some embodiments, the executing entity can generate document matching information corresponding to the document data based on the first matching information and the second matching information. The document matching information can characterize the degree of semantic matching between the document data and the target creative text. The document matching information can be in numerical form or in tag form. For example, the document matching information can be a value between 0 and 1. The higher the value, the higher the semantic similarity between the document data and the target creative text.
[0063] As an example, the aforementioned executing entity can add the first matching information and the second matching information to obtain the document matching information.
[0064] In some optional implementations of certain embodiments, the execution entity may generate document matching information corresponding to the document data based on the first matching information and the second matching information, including the following steps: The first step is to set the first weight corresponding to the first matching information and the second weight corresponding to the second matching information based on the creation focus information corresponding to the target creative text. The creation focus information can be the content focus direction of the target creative text during the creation process. Here, the content focus direction can be the semantic refinement focus direction of the creation object's content creation. In practice, the creation focus direction information can be one of the following: a semantic refinement focus direction emphasizing semantic understanding, or a focus direction emphasizing precise keyword matching.
[0065] The second step involves weighting the first matching information, the first weight, the second matching information, and the second weight to obtain the document matching information.
[0066] Step 403: Select the first literature dataset from the above literature database that meets the first matching condition for the corresponding literature matching information.
[0067] In some embodiments, the aforementioned executing entity can filter out a first dataset of documents from the aforementioned document database whose corresponding document matching information meets a first matching condition. The first matching condition can be document data whose corresponding numerical value is among the top first target number. The first target number can be a value set during the document data recall phase. For example, if the number of candidate document citations for a candidate document citation information sequence is K, the corresponding first target number can be 2K. By recalling the first target number of first document data, both recall rate and redundancy control are considered, providing optimization space for the subsequent reordering process. The reordering process here refers to the process of reordering the document data. The recall process can be the process of recalling document data.
[0068] Step 404: Based on the first literature dataset mentioned above, generate a sequence of candidate literature citation information in the target format.
[0069] In some embodiments, the execution entity may generate a sequence of candidate citation information in the target format based on the first document dataset.
[0070] As an example, firstly, the aforementioned execution entity can filter out the first document dataset from the first document dataset, selecting those corresponding to the second matching information that rank within the top second target number, as a subset of the first document data. The second target number can be the number of first document data to be filtered out during the rearrangement process. For example, the second target number can be the number of candidate document citations corresponding to the candidate document citation information sequence. For example, the first target number could be 2K, and the second target number could be K. Then, sequence generation prompts are generated to generate candidate document citation information sequences in the target format based on the subset of first document data. Finally, the sequence generation prompts are input into the large language model to generate the candidate document citation information sequence.
[0071] In some optional implementations of certain embodiments, the execution entity can generate a sequence of candidate citation information in the target format based on the first citation dataset, including the following steps: The first step involves inputting the text summary and the target creative text included in each of the aforementioned first-document datasets into a pre-trained text content association information generation model to obtain text content association information. This model can be a neural network model used to generate text content association information. The text content association information can represent the semantic similarity between text contents. For example, the model could be a ReRanker model. The ReRanker model has stronger feature extraction capabilities and can better measure the semantic association between the text summary and the target creative text. The text content association information can be a numerical value between 0 and 1; a higher value indicates a stronger semantic association.
[0072] The second step involves selecting a subset of document data from the first document dataset whose corresponding text content association information satisfies the second matching condition, thus forming the second document dataset. The second matching condition can be document data whose corresponding text content association information values rank among the top two target numbers.
[0073] The third step is to generate a sequence of candidate citation information in the target format based on the second literature dataset mentioned above.
[0074] As an example, the aforementioned execution entity can generate information generation prompts based on the second literature dataset to generate a sequence of candidate literature citation information. These prompts are then input into a large language model to obtain the sequence of candidate literature citation information.
[0075] In some optional implementations of certain embodiments, the execution entity can generate a sequence of candidate citation information in the target format based on the second document dataset, including the following steps: The first step is to use the aforementioned large language model to generate the content domain type corresponding to the target creative text. The content domain type can be the content domain of the text content corresponding to the target creative text. For example, the content domain type could be a policy content type, or it could be a type corresponding to legal provisions.
[0076] As an example, firstly, the aforementioned execution entity can generate type hint information to determine the content domain type corresponding to the target creative text. Then, the type hint information is input into the aforementioned large language model to obtain the content domain type.
[0077] The second step is to determine the corresponding citation format for the content domain types mentioned above, which will be used as the target format. Different content domain types have corresponding citation formats. The citation format can be the way the text quotes the content. For example, for the content domain type of academic paper, the corresponding citation format is "APA format or MLA format".
[0078] The third step is to sort the various second-document data in the above-mentioned second-document dataset according to the text content association information corresponding to the second-document data, so as to obtain the second-document data sequence.
[0079] As an example, the aforementioned execution entity can sort the various second-document data in the aforementioned second-document dataset according to the order of the corresponding text content association information from largest to smallest, thereby obtaining a second-document data sequence.
[0080] The fourth step is to generate a sequence of candidate citation information in the target format based on the second literature data sequence mentioned above.
[0081] As an example, the aforementioned executing entity can convert each second document data in the second document data sequence into a conversion result in the target format, which serves as candidate document citation information, thus obtaining a candidate document citation information sequence.
[0082] Step 405: Display the sequence of candidate literature citation information on the target page corresponding to the target text to select candidate literature citation information.
[0083] Step 406: In response to receiving selection information for the target candidate literature citation information in the above candidate literature citation information sequence, the above candidate literature citation information is added to the above target creative text to obtain the literature added text.
[0084] Step 407: Display the added text of the above-mentioned document on the target page.
[0085] In some embodiments, the specific implementation of steps 405-407 and the resulting technical effects can be found in [reference needed]. Figure 3 Steps 302-304 in the corresponding embodiments will not be repeated here.
[0086] from Figure 4 It can be seen from this that, with Figure 3 Compared to the description of some corresponding embodiments, Figure 4 In some corresponding embodiments, the document addition method process 400 uses multiple perspectives, including keyword sets and text summaries, to select more precise recall data as the first candidate document dataset. Based on this, the citation information sequence of candidate documents can be accurately determined.
[0087] Further reference Figure 5 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a document adding device, which are similar to... Figure 2 Corresponding to the method embodiments shown, this document adds an apparatus that can be specifically applied to various electronic devices.
[0088] like Figure 5 As shown, a document addition device 500 includes: a generation unit 501, a first display unit 502, an addition unit 503, and a second display unit 504. The generation unit 501 is configured to, in response to receiving a document citation request for a target created text during the text creation process, generate a sequence of candidate document citation information in a target format based on the target created text; the first display unit 502 is configured to display the sequence of candidate document citation information on a target page corresponding to the target created text for selecting candidate document citation information; the addition unit 503 is configured to, in response to receiving selection information for target candidate documents in the candidate document sequence, add the target candidate documents to the target created text to obtain document-added text; and the second display unit 504 is configured to display the document-added text on the target page.
[0089] In some optional implementations of certain embodiments, the generation unit 501 may be further configured to: in response to receiving a citation request for a target text during text creation, generate a keyword set and text summary corresponding to the target text using a large language model; for each document data in the document database, perform the following generation steps: determine a first matching information between the document data and the keyword set, and determine a second matching information between the document data and the text summary; generate document matching information corresponding to the document data based on the first matching information and the second matching information; select a first document dataset from the document database whose corresponding document matching information satisfies the first matching condition; and generate a sequence of candidate citation information in the target format based on the first document dataset.
[0090] In some optional implementations of certain embodiments, the generation unit 501 may be further configured to: for each first document data in the first document dataset, perform a second generation step: input the text summary and the target creative text included in the first document data into a pre-trained text content association information generation model to obtain text content association information; generate text content association information based on the text content association information and the document matching information corresponding to the first document data; select a subset of document data whose corresponding text content association information satisfies the second matching condition from the first document dataset as a second document dataset; and generate a sequence of candidate document citation information in the target format based on the second document dataset.
[0091] In some optional implementations of certain embodiments, the generation unit 501 may be further configured to: generate the content domain type corresponding to the target creative text using the large language model; determine the literature citation format corresponding to the content domain type as the target format; sort each second literature data in the second literature dataset according to the text content association information corresponding to the second literature data to obtain a second literature data sequence; and generate a candidate literature citation information sequence under the target format according to the second literature data sequence.
[0092] In some optional implementations of some embodiments, the generation unit 501 may be further configured to: set a first weight corresponding to the first matching information and a second weight corresponding to the second matching information based on the creation focus direction information corresponding to the target creative text; and perform weighted processing on the first matching information, the first weight, the second matching information and the second weight to obtain document matching information.
[0093] It is understandable that the units described in the document addition device 500 are related to the reference. Figure 3 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the document addition device 500 and the units contained therein, and will not be repeated here.
[0094] The following is for reference. Figure 6 It illustrates electronic devices suitable for implementing some embodiments of this disclosure (e.g., Figure 1 A schematic diagram of the structure of electronic device 101)600. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0095] like Figure 6 As shown, the electronic device 600 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory 602 or a program loaded from a storage device 608 into a random access memory 603. The random access memory 603 also stores various programs and data required for the operation of the electronic device 600. The processing unit 601, the read-only memory 602, and the random access memory 603 are interconnected via a bus 604. An input / output interface 605 is also connected to the bus 604.
[0096] Typically, the following devices can be connected to the input / output interface 605: input devices 606 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 608 including, for example, magnetic tape, hard disk, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 6 Each box shown can represent a device or multiple devices as needed.
[0097] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a read-only memory 602. When the computer program is executed by the processing device 601, it performs the functions defined above in the methods of some embodiments of this disclosure.
[0098] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0099] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0100] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: in response to receiving a citation request for a target created text during the text creation process, generate a sequence of candidate citation information in a target format based on the target created text; display the sequence of candidate citation information on a target page corresponding to the target created text for selection of candidate citation information; in response to receiving selection information for target candidate citation information in the sequence of candidate citation information, add the candidate citation information to the target created text to obtain citation-added text; and display the citation-added text on the target page.
[0101] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0102] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0103] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a generation unit, a first display unit, an addition unit, and a second display unit. The names of these units do not necessarily limit the specific unit; for example, the second display unit may also be described as "a unit that displays the added text of the aforementioned document on the target page."
[0104] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0105] Some embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements any of the document addition methods described above.
[0106] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A method for adding references, comprising: In response to receiving a citation request for a target text during the text creation process, a sequence of candidate citation information in the target format is generated based on the target text. The candidate literature citation information sequence is displayed on the target page corresponding to the target creative text, so as to select the candidate literature citation information; In response to receiving selection information for target candidate literature citation information in the candidate literature citation information sequence, the candidate literature citation information is added to the target creative text to obtain the literature added text; The document will be displayed with added text on the target page.
2. The method according to claim 1, wherein, The response to receiving a citation request for a target text during the text creation process, and generating a sequence of candidate citation information in the target format based on the target text, includes: In response to receiving a citation request for a target text during the text creation process, a keyword set and text summary corresponding to the target text are generated using a large language model. For each piece of literature data in the literature database, perform the following generation steps: Determine the first matching information between the document data and the keyword set, and determine the second matching information between the document data and the text abstract; Based on the first matching information and the second matching information, generate document matching information corresponding to the document data; A first dataset of documents whose corresponding document matching information meets the first matching condition is selected from the document database; Based on the first literature dataset, a sequence of candidate literature citation information in the target format is generated.
3. The method according to claim 2, wherein, The step of generating a sequence of candidate citation information in the target format based on the first document dataset includes: For each piece of first document data in the first document dataset, the text summary and the target creative text included in the first document data are input into a pre-trained text content association information generation model to obtain text content association information; A subset of document data whose corresponding text content association information meets the second matching condition is selected from the first document dataset and used as the second document dataset; Based on the second literature dataset, a sequence of candidate literature citation information in the target format is generated.
4. The method according to claim 3, wherein, The step of generating a sequence of candidate citation information in the target format based on the second document dataset includes: Using the large language model, the content domain type corresponding to the target creative text is generated; Determine the literature citation format corresponding to the content domain type as the target format; Based on the text content association information corresponding to the second document data, sort each second document data in the second document dataset to obtain the second document data sequence; Based on the second document data sequence, a candidate document citation information sequence in the target format is generated.
5. The method according to claim 2, wherein, The step of generating document matching information corresponding to the document data based on the first matching information and the second matching information includes: Based on the creative focus information corresponding to the target creative text, set a first weight corresponding to the first matching information and a second weight corresponding to the second matching information; The first matching information, the first weight, the second matching information, and the second weight are weighted to obtain the document matching information.
6. The method according to claim 2, wherein, The literature data in the literature database includes: text fragments from the literature text and corresponding fragment vectors; and The document dataset corresponding to the document text was generated through the following steps: The document text is processed into text blocks to obtain a set of text fragments, wherein the character length of each text fragment is less than or equal to the target length; Each text fragment in the text fragment set is transformed into a high-dimensional vector to generate a fragment vector, thus obtaining a fragment vector set; The text fragment set and the fragment vector set are combined accordingly to obtain the document dataset.
7. A document adding device, comprising: The generation unit is configured to, in response to receiving a citation request for a target text during the text creation process, generate a sequence of candidate citation information in a target format based on the target text. The first display unit is configured to display the sequence of candidate literature citation information on the target page corresponding to the target creative text, so as to select candidate literature citation information; The adding unit is configured to add the target candidate document to the target creative text in response to receiving selection information for a target candidate document in the candidate document sequence, thereby obtaining the document added text. The second display unit is configured as follows.
8. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.
9. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.
10. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-6.