A text steganography method, related methods, and apparatus incorporating a large language model

By combining large language models and knowledge graphs, the problem of poor text quality in traditional text steganography techniques is solved, achieving high-quality, high-capacity hiding of secret information and ensuring the concealment and reliability of information transmission.

CN120163135BActive Publication Date: 2025-11-14HUBEI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510103304.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-11-14
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Traditional generative text steganography techniques are limited by the limitations of early models, resulting in low-quality generated text. As the amount of embedded information increases, the naturalness and fluency decrease, making it difficult to reach the level of human natural language.

Method used

By combining large language models and knowledge graphs, the secret information bitstream is segmented and then encoded using knowledge graphs in candidate and supplementary databases to generate steganographic text. The database is updated during the generation process to ensure that each generated text corresponds to a high-quality knowledge graph entity relationship.

Benefits of technology

It significantly improves the quality and naturalness of the generated text, making the embedded secret information more concealed and less likely to be detected, maintaining the naturalness of the text, achieving a higher capacity of secret information hiding, and enhancing the practicality and reliability of text steganography.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163135B_ABST
    Figure CN120163135B_ABST
Patent Text Reader

Abstract

This invention discloses a text steganography method, related methods, and apparatus that incorporates a large language model. The method includes: dividing a secret information bitstream into multiple segmented bitstreams according to a preset segment length; collecting knowledge graphs based on the secret information bitstreams and the preset segment length, constructing a candidate database and a supplementary database; encoding all knowledge graphs in the candidate database to determine the knowledge graph corresponding to the first segmented bitstream, which is then selected as the chosen knowledge graph; inputting the selected knowledge graph into a preset knowledge graph text generation model to obtain the steganographic text corresponding to the first segmented bitstream; deleting the selected knowledge graph from the candidate database and selecting a knowledge graph from the supplementary database to add to the candidate database; and re-executing the knowledge graph encoding and steganographic text process until all segmented bitstreams have been processed. This method improves the quality and naturalness of text generation, achieving high-capacity, covert secret information hiding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a text steganography method, related methods, and apparatus that incorporates a large language model. Background Technology

[0002] Language steganography embeds secret information into communication through language structure, grammar, or context, and is widely used in military, intelligence, and commercial fields. Text steganography is a common form, divided into generative and modification types. Generative text steganography uses generative models to embed information into text, making it both concealed and secure. With the development of large language models, generative text steganography has become mainstream. Traditional generative steganography uses deep learning models such as Recurrent Neural Networks (RNNs) and Variational Autoencoders (VAEs), generating text with relatively low naturalness. However, with the advent of large language models, text quality has significantly improved, thanks to advancements in Transformer architecture and deep learning.

[0003] The Transformer architecture, proposed in 2017, improved the generative model's ability to model long-range dependencies, laying the foundation for large language models. The GPT model, proposed in 2018, ushered in the era of pre-training. In 2019, the emergence of large-scale models such as BERT marked a breakthrough in parameter scale, significantly improving the performance of natural language processing tasks.

[0004] A knowledge graph is a graphical structure used to organize and represent knowledge, typically displaying the relationships between entities, relations, and attributes in the form of a graph. Its goal is to build rich, structured knowledge networks to help computers better understand and process information. Applications of knowledge graphs span multiple fields, including natural language processing, search engines, and the Semantic Web. Triples are the basic elements in a knowledge graph, used to represent connections within the graph. A triple consists of a subject, a predicate, and an object, representing an entity, a relation, and a related entity, respectively. Through this structured representation, knowledge graphs can more clearly express and store rich knowledge. Summary of the Invention

[0005] To obtain more covert steganographic text, this invention provides a text steganography method, related methods, and apparatus that incorporate a large language model.

[0006] In a first aspect, embodiments of the present invention provide a text steganography method incorporating a large language model, which may include:

[0007] The acquired secret information bit stream is divided into multiple sequentially arranged bit streams according to the preset segment length;

[0008] Based on the secret information bitstream and the preset segment length, the candidate data volume and the supplementary data volume are calculated respectively, and the corresponding number of knowledge graphs are collected to obtain the candidate database and the supplementary database.

[0009] Encode all knowledge graphs in the candidate database according to the order in which they are entered into the database, and determine the encoding string corresponding to each knowledge graph in the candidate database;

[0010] Based on the encoded string corresponding to each knowledge graph in the candidate database, the knowledge graph corresponding to the first segmented bit stream is determined as the selected knowledge graph.

[0011] The selected knowledge graph is input into a preset knowledge graph text generation model constructed using a large language model to obtain the steganographic text corresponding to the first segmented bitstream.

[0012] The selected knowledge graph is deleted from the candidate database, and a knowledge graph is selected from the supplementary database and added to the candidate database to obtain an updated candidate database;

[0013] The process of re-encoding all knowledge graphs in the updated candidate database to obtain the steganographic text corresponding to the next segment bitstream is repeated until all the segment bitstreams have been processed.

[0014] In one or more optional embodiments of this application, the step of calculating the candidate data volume and the supplementary data volume based on the secret information bitstream and the preset segment length, and collecting a corresponding number of knowledge graphs to obtain the candidate database and the supplementary database, includes:

[0015] Based on the preset segment length, calculate the power of 2 to obtain the candidate data volume;

[0016] The supplementary data amount is obtained by calculating the quotient of the length of the secret information bit stream and the preset segment length;

[0017] Based on the amount of candidate data and the amount of supplementary data, a corresponding number of knowledge graphs are collected to obtain a candidate database and a supplementary database.

[0018] In one or more optional embodiments of this application, the step of collecting a corresponding number of knowledge graphs based on the candidate data volume and the supplementary data volume to obtain a candidate database and a supplementary database includes:

[0019] Obtain a knowledge graph;

[0020] Check whether the knowledge graph already exists in the candidate database and the supplementary database;

[0021] If so, then discard the knowledge graph.

[0022] If not, the knowledge graph is input into the preset knowledge graph text generation model to obtain the text to be generated;

[0023] Entity extraction is performed on the generated text to be detected to obtain multiple entities;

[0024] Detect whether the multiple entities match the knowledge graph:

[0025] If so, the knowledge graph is added to the candidate database or supplementary database;

[0026] If not, then discard the knowledge graph.

[0027] A new knowledge graph is acquired, and the above-mentioned knowledge graph uniqueness and suitability detection processes are performed until the number of knowledge graphs in the candidate database and the supplementary database meets the requirements of the candidate data volume and the supplementary data volume.

[0028] In one or more optional embodiments of this application, the selected knowledge graph is input into a preset knowledge graph text generation model constructed using a large language model to obtain the steganographic text corresponding to the first segmented bitstream, including:

[0029] The selected knowledge graph is decomposed into multiple triples;

[0030] The elements in each triplet are concatenated in subject-verb-object order, and adjacent elements are separated by a first separator to obtain the triplet text corresponding to the triplet.

[0031] The text corresponding to all the triples is concatenated, and adjacent triple texts are separated by a second separator to obtain the initial text corresponding to the selected knowledge graph.

[0032] The initial text is input into the preset knowledge graph text generation model to obtain the steganographic text corresponding to the first segmented bitstream.

[0033] In one or more optional embodiments of this application, the preset knowledge graph text generation model is obtained in the following manner:

[0034] Obtain the initial knowledge graph text generation model;

[0035] The initial knowledge graph text generation model is fine-tuned and trained using the P-Tuning v2 method to obtain a preset knowledge graph text generation model.

[0036] In one or more optional embodiments of this application, after dividing the acquired secret information bitstream into multiple sequentially arranged segmented bitstreams according to a preset segment length, the method further includes:

[0037] If the length of the last segment bitstream is not equal to the preset segment length, then padding with zeros is used to extend the length of the last segment bitstream to the preset segment length.

[0038] In one or more optional embodiments of this application, before encoding all knowledge graphs in the candidate database according to the order in which the knowledge graphs are added to the database, and determining the encoding string corresponding to each knowledge graph in the candidate database, the method further includes:

[0039] Convert the length of the secret information bit stream into binary to obtain the total length encoded string;

[0040] The total length encoded string is concatenated with the secret information bit stream to form a new secret information bit stream.

[0041] Secondly, embodiments of the present invention provide a method for extracting steganographic text by incorporating a large language model, which may include:

[0042] Obtain steganographic text, a candidate database, and a supplementary database; the knowledge graphs in the candidate database are arranged in the order of entry into the database.

[0043] Entity extraction is performed on the steganographic text to obtain multiple entities to be matched;

[0044] Based on the multiple entities to be matched, the candidate database is matched to obtain the knowledge graph corresponding to the steganographic text;

[0045] Encode all knowledge graphs in the candidate database according to the order in which they are entered into the database, and determine the encoding string corresponding to each knowledge graph in the candidate database;

[0046] Based on the encoding string corresponding to each knowledge graph in the candidate database, the encoding string corresponding to the knowledge graph of the steganographic text is determined as the bit stream corresponding to the steganographic text.

[0047] The knowledge graph corresponding to the steganographic text is deleted from the candidate database, and a knowledge graph is selected from the supplementary database and added to the candidate database to obtain an updated candidate database.

[0048] The next steganographic text is retrieved again, and the above process of entity extraction is performed on the next steganographic text to obtain the bit stream corresponding to the next steganographic text. This process continues until no more steganographic text is received. The bit streams corresponding to all steganographic texts are then concatenated to obtain the secret information bit stream.

[0049] Thirdly, embodiments of the present invention provide a text steganography device that incorporates a large language model, which may include:

[0050] The first segmentation module is used to segment the acquired secret information bit stream into multiple sequentially arranged segmented bit streams according to a preset segment length;

[0051] The database construction module is used to calculate the amount of candidate data and the amount of supplementary data based on the secret information bit stream and the preset segment length, and to collect the corresponding number of knowledge graphs to obtain the candidate database and the supplementary database.

[0052] The first encoding module is used to encode all knowledge graphs in the candidate database according to the order in which the knowledge graphs are entered into the database, and to determine the encoding string corresponding to each knowledge graph in the candidate database.

[0053] The first selection module is used to determine the first knowledge graph corresponding to the segmented bit stream based on the encoded string corresponding to each knowledge graph in the candidate database, and to select the knowledge graph.

[0054] The text generation module is used to input the selected knowledge graph into a preset knowledge graph text generation model constructed using a large language model, and obtain the steganographic text corresponding to the first segmented bitstream;

[0055] The first update module is used to delete the selected knowledge graph from the candidate database and select a knowledge graph from the supplementary database to add to the candidate database, thereby obtaining an updated candidate database;

[0056] The first judgment module is used to determine whether all the segmented bit streams have been processed; if not, the first encoding module described above is executed again.

[0057] Fourthly, embodiments of the present invention provide a steganography extraction device incorporating a large language model, which may include:

[0058] The first acquisition module is used to acquire steganographic text, candidate database, and supplementary database; the knowledge graph in the candidate database is arranged in the order of entry into the database.

[0059] The entity extraction module is used to extract entities from the steganographic text to obtain multiple entities to be matched.

[0060] The first matching module is used to match the multiple entities to be matched with the candidate database to obtain the knowledge graph corresponding to the steganographic text;

[0061] The second encoding module is used to encode all knowledge graphs in the candidate database according to the order in which the knowledge graphs are entered into the database, and to determine the encoding string corresponding to each knowledge graph in the candidate database.

[0062] The first determining module is used to determine the encoding string corresponding to the knowledge graph corresponding to the steganographic text based on the encoding string corresponding to each knowledge graph in the candidate database, and use it as the bit stream corresponding to the steganographic text.

[0063] The second update module is used to delete the knowledge graph corresponding to the steganographic text from the candidate database and select a knowledge graph from the supplementary database to add to the candidate database, thereby obtaining an updated candidate database.

[0064] The second judgment module is used to determine whether the number of received steganographic texts is equal to the number of knowledge graphs in the supplementary database: if so, the bit streams corresponding to all steganographic texts are concatenated to obtain the secret information bit stream; if not, the next steganographic text is obtained and the above entity extraction module is executed again.

[0065] Fifthly, embodiments of the present invention provide a computer-readable storage medium having a computer program / instruction stored thereon, which, when executed by a processor, implements the text steganography method incorporating a large language model as described above, and / or the steganography extraction method incorporating a large language model.

[0066] Sixthly, embodiments of the present invention provide a computer program product, including a computer program / instruction that, when executed by a processor, implements the text steganography method incorporating a large language model as described above, and / or the steganography extraction method incorporating a large language model.

[0067] In a seventh aspect, embodiments of the present invention provide a computer device, including a memory, a processor, and a computer program stored in the memory. When the processor executes the computer program, it implements the text steganography method incorporating a large language model as described above, and / or the steganography extraction method incorporating a large language model.

[0068] The beneficial effects of the above-described technical solutions provided in the embodiments of the present invention include at least the following:

[0069] This invention provides a text steganography method incorporating a large language model. This method divides the secret information bitstream into multiple segmented bitstreams of preset segment length. Then, it calculates the candidate database and supplementary data volume, collecting a corresponding number of knowledge graphs to form the candidate database and supplementary database. Next, it encodes the knowledge graphs in the candidate database, selects the knowledge graph corresponding to the first segmented bitstream, generates steganographic text, deletes the selected graph, and adds a new graph to the supplementary database, thereby updating the candidate database. Subsequently, the process of encoding the knowledge graphs in the candidate database and updating the candidate database is repeated until all segmented bitstreams have been processed. This method, by incorporating a large language model and selecting based on knowledge graphs, significantly improves the quality and naturalness of the generated text, making the embedded secret information more concealed and less easily detected. Compared with traditional generative text steganography techniques, this method not only overcomes the problem of poor text quality caused by model limitations but also effectively controls the amount of information embedded without affecting the naturalness of the text by constructing and utilizing a knowledge graph database. This ensures that each generated text corresponds to a high-quality knowledge graph entity relationship, thereby achieving higher capacity secret information hiding. Meanwhile, a pre-defined knowledge graph text generation model is constructed using a large-scale language model, ensuring high fluency and semantic accuracy in text generation, enhancing the practicality and reliability of text steganography, and achieving effective transmission of secret information while maintaining excellent concealment.

[0070] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.

[0071] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0072] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0073] Figure 1 A flowchart illustrating the text steganography method incorporating a large language model provided in an embodiment of the present invention;

[0074] Figure 2 Example diagram of the adaptation detection steps provided in the embodiments of the present invention;

[0075] Figure 3 A schematic diagram illustrating the framework for establishing a candidate database and a supplementary database provided in an embodiment of the present invention;

[0076] Figure 4 This is an example diagram of perfect binary tree encoding provided in an embodiment of the present invention;

[0077] Figure 5 A schematic diagram illustrating the conversion steps of the initial text corresponding to the knowledge graph provided in this embodiment of the invention;

[0078] Figure 6 A schematic diagram illustrating the framework of the text steganography and steganography extraction method incorporating a large language model provided in this embodiment of the invention;

[0079] Figure 7 A schematic diagram illustrating the process of steganography extraction using a large language model, provided for embodiments of this application;

[0080] Figure 8 A schematic diagram of the structure of a text steganography device incorporating a large language model, provided in an embodiment of this application;

[0081] Figure 9 This is a schematic diagram of the structure of the stegtext extraction device that incorporates a large language model, as provided in an embodiment of this application. Detailed Implementation

[0082] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0083] The inventors discovered that, while traditional generative text steganography techniques can embed secret information, the quality of the generated text is often low due to limitations of early models. Furthermore, as the amount of embedded information increases, the naturalness and fluency of the text gradually decrease, making it difficult to reach the level of human natural language. Based on this, the inventors conducted further research and development, resulting in this invention, which provides a text steganography method, related methods, and apparatus that incorporates a large language model.

[0084] Example 1

[0085] Embodiment 1 of this invention provides a text steganography method incorporating a large language model, referring to... Figure 1 As shown, the method may include the following steps S101-S108:

[0086] S101: Divide the acquired secret information bit stream into multiple sequentially arranged segment bit streams according to the preset segment length.

[0087] S102: Based on the secret information bit stream and the preset segment length, calculate the amount of candidate data and the amount of supplementary data respectively, and collect the corresponding number of knowledge graphs to obtain the candidate database and the supplementary database.

[0088] S103: Encode all knowledge graphs in the candidate database according to the order in which they were entered into the database, and determine the encoding string corresponding to each knowledge graph in the candidate database.

[0089] S104: Based on the encoded string corresponding to each knowledge graph in the candidate database, determine the knowledge graph corresponding to the first segmented bit stream, and use it as the selected knowledge graph.

[0090] S105: Input the selected knowledge graph into the preset knowledge graph text generation model constructed using a large language model to obtain the steganographic text corresponding to the first segmented bitstream.

[0091] S106: Delete the selected knowledge graph from the candidate database and select a knowledge graph from the supplementary database to add to the candidate database, thus obtaining an updated candidate database.

[0092] S107: Determine whether all segmented bit streams have been processed;

[0093] If so, proceed to step S108;

[0094] If not, repeat steps S103-S106 above to obtain the steganographic text corresponding to the next segment bitstream.

[0095] S108: Complete text steganography.

[0096] This invention provides a text steganography method incorporating a large language model. This method divides the secret information bitstream into multiple segmented bitstreams according to a preset segment length. Then, it calculates the candidate database and supplementary data volume, collecting a corresponding number of knowledge graphs to form the candidate database and supplementary database. Next, it encodes the knowledge graphs in the candidate database, selects the knowledge graph corresponding to the first segmented bitstream, generates steganographic text, deletes the selected graph, and adds a new graph to the supplementary database, thereby updating the candidate database. Subsequently, the process of encoding the knowledge graphs in the candidate database and updating the candidate database is repeated until all segmented bitstreams have been processed. This method, by introducing a large language model and selecting based on knowledge graphs, significantly improves the quality and naturalness of the generated text, making the embedded secret information more concealed and less likely to be detected. Compared with traditional generative text steganography techniques, this method not only overcomes the problem of poor text quality caused by model limitations but also effectively controls the amount of information embedded without affecting the naturalness of the text by constructing and utilizing a knowledge graph database. This ensures that each generated text corresponds to a high-quality knowledge graph entity relationship, thereby achieving a higher capacity for hiding secret information. Meanwhile, a pre-defined knowledge graph text generation model is constructed using a large-scale language model, ensuring high fluency and semantic accuracy in text generation, enhancing the practicality and reliability of text steganography, and achieving effective transmission of secret information while maintaining excellent concealment.

[0097] In step S101 above, the acquired secret information bit stream is divided into multiple sequentially arranged segmented bit streams according to the preset segment length.

[0098] Specifically, it can be done by dividing the secret information bit stream into multiple segmented bit streams from the beginning according to a preset segment length, and arranging these segmented bit streams in order of their position in the secret information bit stream.

[0099] When the length of the secret information bit stream is equal to a multiple of the preset segment length, the length of each segment bit stream is equal to the preset segment length. When the length of the secret information bit stream is not equal to a multiple of the preset segment length, the length of the other segment bit streams, except for the last segment bit stream, is equal to the preset segment length, and the length of the last segment bit stream is less than the preset segment length.

[0100] For example, if the length of the secret information bit stream is 2^10 = 1024 and the preset segment length is 2^4 = 16, then the length of the secret information bit stream is equal to a multiple of the preset segment length, and it can be divided into 64 segmented bit streams, with each segmented bit stream having a length of 16.

[0101] For example, if the length of the secret information bit stream is 1020 and the preset segment length is 2^4 = 16, then the length of the secret information bit stream is not equal to a multiple of the preset segment length. It can be divided into 64 segmented bit streams. The length of the first 63 segmented bit streams is equal to 16, and the length of the last segmented bit stream is equal to 12.

[0102] In this embodiment of the application, when the length of the secret information bit stream is not equal to a multiple of the preset segment length, that is, when the length of the last segment bit stream is not equal to the preset segment length, this situation is referred to as an overflow situation in the art. In order to ensure that subsequent encoding can proceed normally, it is necessary to use the method of padding with zeros to supplement the length of the last segment bit stream to the preset segment length.

[0103] Furthermore, to inform the receiver responsible for extracting the steganographic text whether an overflow has occurred, and if so, to inform them of the number of zeros to pad with, the length of the secret information bitstream can be converted to binary and concatenated with it to obtain a secret information bitstream containing the length information before proceeding with the subsequent steganography process. This ensures that the receiver can correctly decode and recover the original secret information bitstream, thereby enhancing the reliability and security of the text steganography process.

[0104] In step S102 above, based on the secret information bitstream and the preset segment length, the candidate data volume and the supplementary data volume are calculated respectively, and the corresponding number of knowledge graphs are collected to obtain the candidate database and the supplementary database. Specifically, this includes the following steps S1021-S1023:

[0105] S1021: Based on the preset segment length, calculate the power of 2 to obtain the candidate data volume.

[0106] Specifically, the candidate data quantity can be obtained by calculating the corresponding power of 2 based on the preset segment length. For example, if the preset segment length is 16, then the candidate data quantity is 2^16 = 65536.

[0107] S1022: Calculate the quotient of the length of the secret information bit stream and the preset segment length to obtain the supplementary data amount.

[0108] For example, if the length of the secret information bitstream is 2^10 = 1024 and the preset segment length is 2^4 = 16, then the amount of supplementary data is 1024 / 16 = 64.

[0109] S1023: Based on the amount of candidate data and supplementary data, collect the corresponding number of knowledge graphs to obtain the candidate database and supplementary database. This specifically includes the following steps S10231-S10239:

[0110] S10231: Obtain a knowledge graph.

[0111] Specifically, knowledge graphs can be obtained in various ways, including but not limited to: randomly generating knowledge graphs using large language models (such as chatgpt); using open-source knowledge graph databases (such as Neo4j); and using entity relation extraction tools (such as Spacy and OpenNLP) to convert arbitrary natural text into knowledge graphs.

[0112] S10232: Check whether the knowledge graph already exists in the candidate database and the supplementary database: if yes, proceed to step S10233; if no, proceed to step S10234.

[0113] In this embodiment, the proposed encoding approach for knowledge graph selection is primarily based on the fact that if encoding of a certain type of element (such as a knowledge graph) is required, each element must be unique. Therefore, before adding the knowledge graph to the candidate database or supplementary database, uniqueness detection in steps S10232-S10233 is necessary.

[0114] S10233: Discard the knowledge graph and return to step S10231.

[0115] In this embodiment, after uniqueness detection is completed, it is necessary to further detect the compatibility between the knowledge graph and the preset knowledge graph text generation model to ensure that the generated text is consistent with the knowledge graph. Specifically, this includes the following steps S10234-S10238.

[0116] S10234: Input the knowledge graph into the preset knowledge graph text generation model to obtain the generated text to be detected.

[0117] The preset knowledge graph text generation model can be obtained in the following ways:

[0118] A large open-source language model is obtained as the initial knowledge graph text generation model. The WebNatural Language Generation (WebNLG) corpus is used to fine-tune the initial knowledge graph text generation model based on the P-Tuning v2 method to obtain the preset knowledge graph text generation model.

[0119] One open-source large-scale language model is ChatGLM2-6B. ChatGLM-6B is an open-source, bilingual (Chinese and English) dialogue language model released by the Knowledge Engineering Group (KEG) & Data Mining at Tsinghua University. Based on the GLM (General Language Model) architecture, it has 6.2 billion parameters. ChatGLM2-6B is the second-generation version of ChatGLM-6B. While retaining many excellent features of the first-generation model, such as smooth dialogue and low deployment threshold, ChatGLM2-6B boasts even more powerful performance.

[0120] The WebNLG corpus consists of multiple triples and is designed to represent entities and their relationships through natural language text. Each sample in the WebNLG corpus contains up to seven triples and is accompanied by one or more corresponding reference texts for each sample.

[0121] In this embodiment, the method for obtaining a preset knowledge graph text generation model, through fine-tuning using the WebNLG corpus, successfully transforms the initial knowledge graph text generation model into an efficient knowledge graph text generator, namely the preset knowledge graph text generation model. The triple set in the WebNLG corpus and the corresponding reference text enable the model to learn how to transform structured entity relationship information into fluent natural language text. Simultaneously, by applying the P-Tuning v2 method, this method efficiently fine-tunes the initial knowledge graph text generation model under small sample conditions, greatly reducing the need for a large number of parameters and computational resources while maintaining good generation quality. While optimizing storage and memory usage, it also improves the fine-tuning effect of medium-sized models in complex tasks, thereby achieving efficient and resource-saving knowledge graph text generation.

[0122] S10235: Extract entities from the generated text to be detected to obtain multiple entities.

[0123] Specifically, one could use entity relation extraction tools (such as Spacy or OpenNLP) to extract all entities from the generated text to be detected.

[0124] S10236: Detect whether multiple entities match the knowledge graph: if yes, proceed to step S10237; if no, proceed to step S10238.

[0125] S10237: Add the knowledge graph to the candidate database or supplementary database.

[0126] S10238: Discard the knowledge graph and return to step S10231.

[0127] S10239: Obtain a new knowledge graph and perform the knowledge graph uniqueness detection in steps S10232-S10233 and the adaptability detection in steps S10234-S10238 until the number of knowledge graphs in the candidate database and the supplementary database meets the requirements of the candidate data volume and the supplementary data volume.

[0128] To facilitate understanding of this solution by those skilled in the art, the specific implementation process of steps S10234-S10238, the compatibility detection, provided in the embodiments of this application is described more clearly and completely below: (Refer to...) Figure 2 As shown in the diagram, there are three steps: Step 1, text generation; Step 2, entity extraction; and Step 3, entity comparison, corresponding to steps S10234-S10236 above. In Step 1, the red dashed box represents the example knowledge graph. For ease of explanation, this knowledge graph contains only one triplet, containing two entities, "spider man" and "stan lee," and the relation "creator." This knowledge graph undergoes a simple text conversion, represented as text A as "spider man|creator|stan lee." Then, it is input into the LLM (the preset knowledge graph text generation model mentioned above), resulting in text B as "spider man's creator isstan lee." Next, Step 2 extracts entities from text B, obtaining the entities "spider man" and "stanlee." Finally, in Step 3, it determines whether the entities obtained in Step 2 match the entities in text A. If they match, the knowledge graph is stored in the candidate database or supplementary database; otherwise, it is discarded.

[0129] In this embodiment of the application, the method for establishing the candidate database and supplementary database in step S102 above refers to... Figure 3 As shown in the figure, Graphs represents the knowledge graph not yet included in the database, and test represents the uniqueness detection of the knowledge graph in steps S10232-S10233 and the suitability detection in steps S10234-S10238. If the detection passes, it can be added to the CP Database or Supplement Database, where the CP Database is the candidate database mentioned above, and the Supplement Database is the supplementary database mentioned above. Please note that at this time, it is not necessary to focus on... Figure 3The arrow pointing from the supplementary database to the candidate database indicates that there is no step in step S102 where data from the supplementary database is added to the candidate database; the arrow indicates the subsequent step S106.

[0130] The above step S102, through systematic uniqueness and adaptability detection, ensures the uniqueness and encodeability of each knowledge graph in the database, thereby improving the quality and accuracy of the knowledge graphs and ensuring that the knowledge graphs in the database can accurately support subsequent encoding and text generation tasks, providing a strong guarantee for the practical application of knowledge graphs.

[0131] In step S103 above, all knowledge graphs in the candidate database are encoded according to the order in which they are entered into the database, and the encoding string corresponding to each knowledge graph in the candidate database is determined.

[0132] The encoding method can be either fixed-length coding (FLC) or variable-length coding (VLC). FLC is an encoding method based on a perfect binary tree, while VLC is an encoding method based on a Huffman tree. Both are existing technologies and will not be elaborated on here.

[0133] Taking fixed-length encoding as an example, step S103 above can specifically be as follows: construct a perfect binary tree from all knowledge graphs in the candidate database according to the order of their entry into the database, and encode the entire perfect binary tree in the form of left subtree 0 and right subtree 1, so that the leaf nodes of the perfect binary tree correspond to a knowledge graph from left to right. Then the encoding of the leaf node corresponding to each knowledge graph is the encoding string corresponding to that knowledge graph.

[0134] For example, if the encoding method is fixed-length encoding and the preset segment length is 3, then according to step S102 above, it can be determined that the candidate database includes a total of 8 knowledge graphs, numbered ①-⑧ according to the entry order. A perfect binary tree is then constructed based on the candidate database, and the entire perfect binary tree is encoded with the left subtree being 0 and the right subtree being 1. An example diagram of the resulting perfect binary tree is shown below. Figure 4 As shown, the perfect binary tree has 8 leaf nodes. By mapping these 8 leaf nodes from left to right to the 8 sequentially arranged knowledge graphs, we can obtain the following encoding strings: Knowledge graph ① is 000, Knowledge graph ② is 001, Knowledge graph ③ is 010, Knowledge graph ④ is 011, Knowledge graph ⑤ is 100, Knowledge graph ⑥ is 101, Knowledge graph ⑦ is 110, and Knowledge graph ⑧ is 111.

[0135] In step S104 above, the knowledge graph corresponding to the first segmented bit stream is determined based on the encoded string corresponding to each knowledge graph in the candidate database, and is used as the selected knowledge graph.

[0136] Specifically, it can be done by finding the knowledge graph whose encoding string is equal to the first segment bit stream based on the encoding string corresponding to each knowledge graph in the candidate database, and using it as the selected knowledge graph.

[0137] In step S105 above, the selected knowledge graph is input into the preset knowledge graph text generation model to obtain the steganographic text corresponding to the first segmented bitstream. Specifically, this includes the following steps S1051-S1054:

[0138] S1051: Decompose the selected knowledge graph into multiple triples.

[0139] S1052: Concatenate the elements in each triple in subject-verb-object order, and separate adjacent elements with the first separator to obtain the triple text corresponding to the triple.

[0140] S1053: Concatenate the text of all triples corresponding to the triples, and separate adjacent triple texts with the second separator to obtain the initial text corresponding to the selected knowledge graph.

[0141] To facilitate understanding of this solution by those skilled in the art, the specific implementation process of S1051-S1053 provided in the embodiments of this application is described more clearly and completely below: Refer to Figure 5 As shown in the figure, the process of converting the selected knowledge graph into initial text is represented by two steps: step 1 decomposes the knowledge graph to obtain individual triples, and step 2 represents the knowledge graph as text. Step 1 corresponds to step S1051, and step 2 corresponds to steps S1052 and S1053.

[0142] exist Figure 5 In step 1, Graph B in the red box in the upper left corner is the selected knowledge graph. The selected knowledge graph is decomposed into two triples, namely the triples in the two green boxes on the left side of the figure. The triple on the left contains two entities "spiderman" and "Peter park" and the relation "real name". The triple on the right contains two entities "spider man" and "stan lee" and the relation "creator".

[0143] exist Figure 5In step 2, the two triples obtained in step 1 are subjected to simple text transformation. Using the symbol "|" as the first separator, the corresponding triple texts "spider man|real name|Peterpark" and "spider man|creator|stan lee" are obtained. Then, the two triple texts are concatenated, and adjacent triple texts are separated by the second separator "*", to obtain the initial text corresponding to the selected knowledge graph "spider man|real name|Peter park*spider man|creator|stan lee".

[0144] S1054: Input the initial text into the preset knowledge graph text generation model to obtain the steganographic text corresponding to the first segmented bitstream.

[0145] The method for obtaining the preset knowledge graph text generation model has been explained in step S10234 above, and will not be repeated here.

[0146] In this embodiment, step S105 utilizes a preset knowledge graph text generation model to convert the selected knowledge graph into corresponding steganographic text, ensuring that the secret information bitstream can be steganized using natural language. This process achieves the conversion from structured knowledge graph to natural language text, ensuring that the information hiding process not only meets the requirements of steganography but also maintains the traceability and semantic consistency of the information, making the steganographic information more difficult for external observers to detect and improving the security and concealment of the steganography process.

[0147] In step S106 above, the selected knowledge graph is deleted from the candidate database, and a knowledge graph is selected from the supplementary database and added to the candidate database to obtain the updated candidate database.

[0148] Specifically, this could involve deleting the selected knowledge graph from the candidate database and adding a knowledge graph from the supplementary database in a preset order to the candidate database, thus obtaining an updated candidate database. The preset order could be the order in which the knowledge graphs were added to the supplementary database.

[0149] For example, if the current loop obtains the steganographic text corresponding to the first segment bitstream, then a knowledge graph with a preset order of 1 should be selected from the supplementary database and added to the candidate database to obtain an updated candidate database.

[0150] If the current loop obtains the steganographic text corresponding to the i-th segment bitstream, then the knowledge graph with the preset order i should be selected from the supplementary database and added to the candidate database to obtain the updated candidate database.

[0151] In this embodiment, the method defines the size of the candidate database by customizing the preset segment length, thereby determining the amount of secret information to be embedded each time. Therefore, even when there is a large amount of secret information to be embedded, it is only necessary to set a sufficiently large candidate database to ensure that the number of knowledge graphs in the candidate database is sufficient to satisfy the embedding of a segment of secret information.

[0152] After completing the text steganography method for knowledge graph selection described in S101-S107 above, the secret information bitstream can be extracted based on the steganographic text. The extraction method specifically includes the following steps S201-S209:

[0153] For ease of explanation, the party that executes the text steganography method incorporating a large language model will be referred to as the sender, and the party that executes the text steganography extraction method incorporating a large language model will be referred to as the receiver.

[0154] S201: Retrieve steganographic text, candidate database, and supplementary database. The knowledge graphs in the candidate database are arranged in the order they were entered into the database.

[0155] S202: Extract entities from the steganographic text to obtain multiple entities to be matched.

[0156] S203: Based on multiple entities to be matched, match them with the candidate database to obtain the knowledge graph corresponding to the steganographic text.

[0157] S204: Encode all knowledge graphs in the candidate database according to the order in which they were entered into the database, and determine the encoding string corresponding to each knowledge graph in the candidate database.

[0158] S205: Based on the encoding string corresponding to each knowledge graph in the candidate database, determine the encoding string corresponding to the knowledge graph of the steganographic text, and use it as the bit stream corresponding to the steganographic text.

[0159] S206: Delete the knowledge graph corresponding to the steganographic text from the candidate database, and select a knowledge graph from the supplementary database to add to the candidate database, thus obtaining an updated candidate database.

[0160] Specifically, this could involve deleting the knowledge graph corresponding to the steganographic text from the candidate database, and then selecting a knowledge graph from the supplementary database to add to the candidate database in a preset order, thus obtaining an updated candidate database. Note that this preset order must be consistent with the preset order in the steganographic text selection method based on knowledge graphs to ensure that the recipient can update the candidate database in the same way, maintaining consistency with the recipient.

[0161] S207: Determine whether the number of received steganographic texts is equal to the number of knowledge graphs in the supplementary database: if yes, proceed to S209; if no, proceed to S208.

[0162] S208: Obtain the next steganographic text and execute the above steps S202-S207.

[0163] S209: Concatenate the bitstreams corresponding to all the steganographic texts to obtain the secret information bitstream.

[0164] To facilitate understanding of this solution by those skilled in the art, the specific implementation process of the text steganography method incorporating a large language model provided in the embodiments of this application will be described more clearly and completely below:

[0165] Reference Figure 6 As shown, the left side represents the framework of the text steganography method incorporating a large language model, and the right side represents the framework of the stegtext extraction method incorporating a large language model. The leftmost vertical row in the figure represents the secret information bitstream. The blue, gray, and purple boxes each represent a segmented bitstream. The Sum in the red box represents the length of the secret information bitstream. After being converted to binary, it is used in conjunction with other segmented bitstreams to obtain the corresponding knowledge graph G0-Gn based on the candidate database. The corresponding stegtext Text 0-Text n is then obtained through a preset knowledge graph text generation model.

[0166] Figure 6 The leftmost Text 0-Text n in the right-hand box are multiple steganographic texts. Through entity extraction and matching, the corresponding knowledge graphs G0-Gn are obtained. The encoded string corresponding to G0 can be converted into decimal to represent the length of the secret information bit stream. At the same time, the encoded strings corresponding to G1-Gn are concatenated in order to obtain the secret information bit stream.

[0167] Example 2

[0168] Based on the same inventive concept, embodiments of the present invention also provide a method for extracting steganographic text by incorporating a large language model, referring to... Figure 7 As shown, the method includes:

[0169] For ease of explanation, the party that executes the text steganography method incorporating a large language model will be referred to as the sender, and the party that executes the text steganography extraction method incorporating a large language model will be referred to as the receiver.

[0170] S201: Retrieve steganographic text, candidate database, and supplementary database. The knowledge graphs in the candidate database are arranged in the order they were entered into the database.

[0171] S202: Extract entities from the steganographic text to obtain multiple entities to be matched.

[0172] S203: Based on multiple entities to be matched, match them with the candidate database to obtain the knowledge graph corresponding to the steganographic text.

[0173] S204: Encode all knowledge graphs in the candidate database according to the order in which they were entered into the database, and determine the encoding string corresponding to each knowledge graph in the candidate database.

[0174] S205: Based on the encoding string corresponding to each knowledge graph in the candidate database, determine the encoding string corresponding to the knowledge graph of the steganographic text, and use it as the bit stream corresponding to the steganographic text.

[0175] S206: Delete the knowledge graph corresponding to the steganographic text from the candidate database, and select a knowledge graph from the supplementary database to add to the candidate database, thus obtaining an updated candidate database.

[0176] S207: Determine whether the number of received steganographic texts is equal to the number of knowledge graphs in the supplementary database: if yes, proceed to S209; if no, proceed to S208.

[0177] S208: Obtain the next steganographic text and execute the above steps S202-S207.

[0178] S209: Concatenate the bitstreams corresponding to all the steganographic texts to obtain the secret information bitstream.

[0179] Example 3

[0180] Based on the same inventive concept, embodiments of the present invention also provide a text steganography device that incorporates a large language model, referring to... Figure 8 As shown, the device includes:

[0181] The first segmentation module 101 is used to segment the acquired secret information bit stream into multiple sequentially arranged segmented bit streams according to a preset segment length;

[0182] The database construction module 102 is used to calculate the amount of candidate data and the amount of supplementary data based on the secret information bit stream and the preset segment length, and to collect the corresponding number of knowledge graphs to obtain the candidate database and the supplementary database.

[0183] The first encoding module 103 is used to encode all knowledge graphs in the candidate database according to the order in which the knowledge graphs are entered into the database, and to determine the encoding string corresponding to each knowledge graph in the candidate database.

[0184] The first selection module 104 is used to determine the first knowledge graph corresponding to the segmented bit stream based on the encoded string corresponding to each knowledge graph in the candidate database, and to select the knowledge graph.

[0185] The text generation module 105 is used to input the selected knowledge graph into a preset knowledge graph text generation model constructed using a large language model to obtain the steganographic text corresponding to the first segmented bit stream.

[0186] The first update module 106 is used to delete the selected knowledge graph from the candidate database and select a knowledge graph from the supplementary database to add to the candidate database, thereby obtaining an updated candidate database;

[0187] The first judgment module 107 is used to determine whether all the segmented bit streams have been processed: if not, the first encoding module described above is executed again.

[0188] Example 4

[0189] Based on the same inventive concept, embodiments of the present invention also provide a steganography extraction device incorporating a large language model, referring to... Figure 9 As shown, the device includes:

[0190] The first acquisition module 201 is used to acquire steganographic text, a candidate database, and a supplementary database; the knowledge graphs in the candidate database are arranged in the order of entry into the database.

[0191] Entity extraction module 202 is used to extract entities from the steganographic text to obtain multiple entities to be matched;

[0192] The first matching module 203 is used to match the multiple entities to be matched with the candidate database to obtain the knowledge graph corresponding to the steganographic text;

[0193] The second encoding module 204 is used to encode all knowledge graphs in the candidate database according to the order in which the knowledge graphs are entered into the database, and to determine the encoding string corresponding to each knowledge graph in the candidate database.

[0194] The first determining module 205 is used to determine the encoding string corresponding to the knowledge graph corresponding to the steganographic text based on the encoding string corresponding to each knowledge graph in the candidate database, and use it as the bit stream corresponding to the steganographic text.

[0195] The second update module 206 is used to delete the knowledge graph corresponding to the steganographic text from the candidate database and select a knowledge graph from the supplementary database to add to the candidate database, thereby obtaining an updated candidate database.

[0196] The second judgment module 207 is used to determine whether the number of received steganographic texts is equal to the number of knowledge graphs in the supplementary database: if yes, then all the bit streams corresponding to the steganographic texts are concatenated to obtain the secret information bit stream; if no, then the next steganographic text is obtained and the above entity extraction module is re-executed.

[0197] Example 5

[0198] Based on the same inventive concept, embodiments of the present invention also provide a computer-readable storage medium storing a computer program / instruction thereon, which, when executed by a processor, implements the text steganography method incorporating a large language model as described in Embodiment 1 above, and / or the steganography extraction method incorporating a large language model as described in Embodiment 2 above.

[0199] Example 6

[0200] Based on the same inventive concept, embodiments of the present invention also provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the text steganography method incorporating a large language model as described in Embodiment 1 above, and / or the steganography extraction method incorporating a large language model as described in Embodiment 2 above.

[0201] Example 7

[0202] Based on the same inventive concept, embodiments of the present invention also provide a computer device, including a memory, a processor, and a computer program stored in the memory. When the processor executes the computer program, it implements the text steganography method incorporating a large language model as described in Embodiment 1 above, and / or the steganography extraction method incorporating a large language model as described in Embodiment 2 above.

[0203] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0204] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0205] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0206] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0207] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A text steganography method incorporating a large language model, characterized in that, include: The acquired secret information bit stream is divided into multiple sequentially arranged bit streams according to the preset segment length; Based on the secret information bitstream and the preset segment length, the candidate data volume and the supplementary data volume are calculated respectively, and the corresponding number of knowledge graphs are collected to obtain the candidate database and the supplementary database. Encode all knowledge graphs in the candidate database according to the order in which they were entered into the database, and determine the encoding string corresponding to each knowledge graph in the candidate database; Based on the encoded string corresponding to each knowledge graph in the candidate database, the knowledge graph corresponding to the first segmented bit stream is determined as the selected knowledge graph. The selected knowledge graph is input into a preset knowledge graph text generation model constructed using a large language model to obtain the steganographic text corresponding to the first segmented bitstream. The selected knowledge graph is deleted from the candidate database, and a knowledge graph is selected from the supplementary database and added to the candidate database to obtain an updated candidate database; The process of re-encoding all knowledge graphs in the updated candidate database to obtain the steganographic text corresponding to the next segment bitstream is repeated until all the segment bitstreams have been processed.

2. The method according to claim 1, characterized in that, Based on the secret information bitstream and the preset segment length, the candidate data volume and supplementary data volume are calculated respectively, and a corresponding number of knowledge graphs are collected to obtain a candidate database and a supplementary database, including: Based on the preset segment length, calculate the power of 2 to obtain the candidate data volume; The supplementary data amount is obtained by calculating the quotient of the length of the secret information bit stream and the preset segment length; Based on the amount of candidate data and the amount of supplementary data, a corresponding number of knowledge graphs are collected to obtain a candidate database and a supplementary database.

3. The method according to claim 2, characterized in that, The step of collecting a corresponding number of knowledge graphs based on the candidate data volume and the supplementary data volume to obtain a candidate database and a supplementary database includes: Obtain a knowledge graph; Check whether the knowledge graph already exists in the candidate database and the supplementary database; If so, then discard the knowledge graph. If not, the knowledge graph is input into the preset knowledge graph text generation model to obtain the text to be generated; Entity extraction is performed on the generated text to be detected to obtain multiple entities; Detect whether the multiple entities match the knowledge graph: If so, the knowledge graph is added to the candidate database or supplementary database; If not, then discard the knowledge graph. A new knowledge graph is acquired, and the above-mentioned knowledge graph uniqueness and suitability detection processes are performed until the number of knowledge graphs in the candidate database and the supplementary database meets the requirements of the candidate data volume and the supplementary data volume.

4. The method according to claim 1, characterized in that, The selected knowledge graph is input into a preset knowledge graph text generation model constructed using a large language model to obtain the steganographic text corresponding to the first segmented bitstream, including: The selected knowledge graph is decomposed into multiple triples; The elements in each triplet are concatenated in subject-verb-object order, and adjacent elements are separated by a first separator to obtain the triplet text corresponding to the triplet. The text corresponding to all the triples is concatenated, and adjacent triple texts are separated by a second separator to obtain the initial text corresponding to the selected knowledge graph. The initial text is input into the preset knowledge graph text generation model to obtain the steganographic text corresponding to the first segmented bitstream.

5. The method according to claim 1, characterized in that, The preset knowledge graph text generation model is obtained through the following method: Obtain the initial knowledge graph text generation model; The initial knowledge graph text generation model is fine-tuned and trained using the P-Tuning v2 method to obtain a preset knowledge graph text generation model.

6. A method for extracting steganographic text by incorporating a large language model, characterized in that, include: Retrieve steganographic text, candidate database, and supplementary database; The knowledge graphs in the candidate database are arranged in the order they were entered into the database; Entity extraction is performed on the steganographic text to obtain multiple entities to be matched; Based on the multiple entities to be matched, the candidate database is matched to obtain the knowledge graph corresponding to the steganographic text; Encode all knowledge graphs in the candidate database according to the order in which they are entered into the database, and determine the encoding string corresponding to each knowledge graph in the candidate database; Based on the encoding string corresponding to each knowledge graph in the candidate database, the encoding string corresponding to the knowledge graph of the steganographic text is determined as the bit stream corresponding to the steganographic text. The knowledge graph corresponding to the steganographic text is deleted from the candidate database, and a knowledge graph is selected from the supplementary database and added to the candidate database to obtain an updated candidate database. The next steganographic text is retrieved again, and the above process of entity extraction is performed on the next steganographic text to obtain the bit stream corresponding to the next steganographic text. This process continues until no more steganographic text is received. The bit streams corresponding to all steganographic texts are then concatenated to obtain the secret information bit stream.

7. A text steganography device incorporating a large language model, characterized in that, include: The first segmentation module is used to segment the acquired secret information bit stream into multiple sequentially arranged segmented bit streams according to a preset segment length; The database construction module is used to calculate the amount of candidate data and the amount of supplementary data based on the secret information bit stream and the preset segment length, and to collect the corresponding number of knowledge graphs to obtain the candidate database and the supplementary database. The first encoding module is used to encode all knowledge graphs in the candidate database according to the order in which the knowledge graphs are entered into the database, and to determine the encoding string corresponding to each knowledge graph in the candidate database. The first selection module is used to determine the first knowledge graph corresponding to the segmented bit stream based on the encoded string corresponding to each knowledge graph in the candidate database, and to select the knowledge graph. The text generation module is used to input the selected knowledge graph into a preset knowledge graph text generation model constructed using a large language model, and obtain the steganographic text corresponding to the first segmented bitstream; The first update module is used to delete the selected knowledge graph from the candidate database and select a knowledge graph from the supplementary database to add to the candidate database, thereby obtaining an updated candidate database. The first judgment module is used to determine whether all the segmented bit streams have been processed; if not, the first encoding module described above is executed again.

8. A stegtext extraction device incorporating a large language model, characterized in that, include: The first acquisition module is used to acquire steganographic text, candidate database, and supplementary database; the knowledge graph in the candidate database is arranged in the order of entry into the database. The entity extraction module is used to extract entities from the steganographic text to obtain multiple entities to be matched. The first matching module is used to match the multiple entities to be matched with the candidate database to obtain the knowledge graph corresponding to the steganographic text; The second encoding module is used to encode all knowledge graphs in the candidate database according to the order in which the knowledge graphs are entered into the database, and to determine the encoding string corresponding to each knowledge graph in the candidate database. The first determining module is used to determine the encoding string corresponding to the knowledge graph corresponding to the steganographic text based on the encoding string corresponding to each knowledge graph in the candidate database, and use it as the bit stream corresponding to the steganographic text. The second update module is used to delete the knowledge graph corresponding to the steganographic text from the candidate database and select a knowledge graph from the supplementary database to add to the candidate database, thereby obtaining an updated candidate database. The second judgment module is used to determine whether the number of received steganographic texts is equal to the number of knowledge graphs in the supplementary database: if so, the bit streams corresponding to all steganographic texts are concatenated to obtain the secret information bit stream; if not, the next steganographic text is obtained and the above entity extraction module is executed again.

9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the text steganography method incorporating a large language model as described in any one of claims 1-5, and / or the steganography extraction method incorporating a large language model as described in claim 6.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the text steganography method incorporating a large language model as described in any one of claims 1-5, and / or the steganography extraction method incorporating a large language model as described in claim 6.

Citation Information

Patent Citations

  • Knowledge graph-based text generation steganography method, related method and device

    CN117131203A

  • Method for realizing main-standby communication of single-link equipment

    CN118118325A