Text steganography method introducing large language model, related method and device

By introducing large language models and knowledge graph-based selection methods, the problem of poor text quality in traditional text steganography technology is solved, and the generation of steganography text with higher quality and naturalness is achieved, ensuring high concealment and effective transmission of secret information.

CN120163135AActive Publication Date: 2025-06-17HUBEI UNIV +1

Patent Information

Application Number
CN202510103304.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-17
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The traditional generative text steganography technology is limited by the limitations of early models, and the generated text quality is not high. As the amount of embedded information increases, the naturalness and fluency of the text decrease, making it difficult to reach the level of human natural language.

Method used

A large language model is introduced and based on the method of knowledge graph selection, the secret information bitstream is divided into multiple segments. The bitstream and preset segment length calculate candidates and supplementary data volumes, collect knowledge graphs to form a database, generate steganographic text by encoding and selecting knowledge graphs, and update the database.

Benefits of technology

It significantly improves the quality and naturalness of generated text, makes the embedded secret information more hidden and less likely to be noticed, overcomes the problem of poor text quality in traditional technology, and effectively controls the amount of information embedding through the knowledge graph database to achieve higher capacity secret information hiding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163135A_ABST
    Figure CN120163135A_ABST
Patent Text Reader

Abstract

The invention discloses a text steganography method introducing a large language model, a related method and a related device. The method comprises the following steps: segmenting a secret information bit stream into a plurality of segmented bit streams according to a preset segment length; based on the secret information bit stream and a preset segment length, collecting a knowledge graph, and constructing a candidate database and a supplementary database; coding is carried out according to all knowledge maps in the candidate database, so that the knowledge map corresponding to the first segmented bit stream is determined to serve as a selected knowledge map; inputting the selected knowledge graph into a preset knowledge graph text generation model to obtain a steganographic text corresponding to the first segmented bit stream; deleting the selected mapping knowledge domain from the candidate database, selecting one mapping knowledge domain from the supplementary database, and adding the selected mapping knowledge domain into the candidate database; and the knowledge graph coding and text steganography processes are executed again until all the segmented bit streams are processed. According to the method, the quality and naturalness of text generation are improved, and high-capacity and hidden secret information hiding is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a text steganography method introducing a large language model, and related methods and devices. Background Art

[0002] Language steganography embeds secret information in communication through language structure, grammar or context, and is widely used in military, intelligence and business fields. Text steganography is a common form of this, which is divided into two categories: generative and modification. Generative text steganography uses generative models to embed information into text, making it hidden and secure. With the development of large language models, generative text steganography has become mainstream. Traditional generative steganography uses deep learning models such as recurrent neural networks (RNN) and variational autoencoders (VAE), and the generated text is less natural. However, with the emergence of large language models, the quality of text has been significantly improved, thanks to the Transformer architecture and advances in deep learning.

[0003] The Transformer architecture proposed in 2017 improved the generative model's ability to model long-distance dependencies, laying the foundation for large language models. The GPT model proposed in 2018 ushered in the era of pre-training. In 2019, the emergence of large-scale models such as BERT marked a breakthrough in parameter scale, significantly improving the performance of natural language processing tasks.

[0004] Knowledge Graph is a graphical structure used to organize and represent knowledge, usually showing the connections between entities, relationships and attributes in the form of a graph. Its goal is to establish a rich, structured knowledge network to help computers better understand and process information. The application of knowledge graphs covers many fields, including natural language processing, search engines, semantic web, etc. Triples are the basic elements in knowledge graphs and are used to represent the connection relationships in the graph. A triple consists of three parts: subject, predicate and object, which represent entities, relationships and related entities respectively. Through this structured representation, knowledge graphs can express and store rich knowledge more clearly. Summary of the invention

[0005] In order to obtain more concealed steganographic text, an embodiment of the present invention provides a text steganographic method, related methods and devices that introduce a large language model.

[0006] In a first aspect, an embodiment of the present invention provides a text steganography method introducing a large language model, which may include:

[0007] The obtained secret information bit stream is divided into a plurality of segmented bit streams arranged in sequence according to a preset segment length;

[0008] Based on the secret information bit stream and the preset segmentation length, calculate the candidate data volume and the supplementary data volume respectively, and collect the corresponding number of knowledge graphs to obtain a candidate database and a supplementary database;

[0009] Encode all the knowledge graphs in the candidate database according to the order of knowledge graph storage, and determine the encoding string corresponding to each knowledge graph in the candidate database;

[0010] According to the encoding string corresponding to each knowledge graph in the candidate database, determine the knowledge graph corresponding to the first segmentation bit stream as the selected knowledge graph;

[0011] Input the selected knowledge graph into a preset knowledge graph text generation model constructed using a large language model to obtain the steganographic text corresponding to the first segmentation bit stream;

[0012] Delete the selected knowledge graph from the candidate database, and select a knowledge graph from the supplementary database and add it to the candidate database to obtain an updated candidate database;

[0013] Re - execute the above process of encoding all the knowledge graphs in the updated candidate database to obtain the steganographic text corresponding to the next segmentation bit stream until all the segmentation bit streams are processed.

[0014] In one or some alternative embodiments of the embodiments of the present application, the step of calculating the candidate data volume and the supplementary data volume respectively based on the secret information bit stream and the preset segmentation length, and collecting the corresponding number of knowledge graphs to obtain a candidate database and a supplementary database includes:

[0015] Based on the preset segmentation length, calculate the power of 2 to obtain the candidate data volume;

[0016] Calculate the quotient of the length of the secret information bit stream and the preset segmentation length to obtain the supplementary data volume;

[0017] According to the candidate data volume and the supplementary data volume, collect the corresponding number of knowledge graphs respectively to obtain a candidate database and a supplementary database.

[0018] In one or some alternative embodiments of the embodiments of the present application, the step of collecting the corresponding number of knowledge graphs respectively according to the candidate data volume and the supplementary data volume to obtain a candidate database and a supplementary database includes:

[0019] Obtain a knowledge graph;

[0020] Check whether the knowledge graph already exists in the candidate database and the supplementary database;

[0021] If so, discard the knowledge graph;

[0022] If not, input the knowledge graph into the preset knowledge graph text generation model to obtain the text to be detected and generated;

[0023] Perform entity extraction on the text to be detected and generated to obtain a plurality of entities;

[0024] Detect whether the plurality of entities match the knowledge graph:

[0025] If so, add the knowledge graph to the candidate database or the supplementary database;

[0026] If not, discard the knowledge graph;

[0027] Re-obtain a new knowledge graph, and perform the above-mentioned knowledge graph uniqueness detection and adaptability detection process until the number of knowledge graphs in the candidate database and the supplementary database meets the candidate data volume and the supplementary data volume.

[0028] In one or some optional implementation manners of the embodiments of the present application, inputting the selected knowledge graph into a preset knowledge graph text generation model constructed using a large language model to obtain the steganographic text corresponding to the first segmented bitstream includes:

[0029] Decompose the selected knowledge graph into a plurality of triples;

[0030] Concatenate the elements in each triple in the order of subject, predicate, and object, and separate adjacent elements with a first separator symbol to obtain the triple text corresponding to the triple;

[0031] Concatenate the triple texts corresponding to all the triples, and separate adjacent triple texts with a second separator symbol to obtain the initial text corresponding to the selected knowledge graph;

[0032] Input the initial text into the preset knowledge graph text generation model to obtain the steganographic text corresponding to the first segmented bitstream.

[0033] In one or some optional implementation manners of the embodiments of the present application, the preset knowledge graph text generation model is obtained by the following method:

[0034] Obtain an initial knowledge graph text generation model;

[0035] Fine-tune and train the initial knowledge graph text generation model based on the P-Tuning v2 method to obtain a preset knowledge graph text generation model.

[0036] In one or some alternative embodiments of the embodiments of the present application, after splitting the obtained secret information bitstream into a plurality of sequentially arranged segmented bitstreams according to a preset segmentation length, the method further includes:

[0037] If the length of the last segmented bitstream is not equal to the preset segmentation length, then use the method of padding with 0s to supplement the length of the last segmented bitstream to the preset segmentation length.

[0038] In one or some alternative embodiments of the embodiments of the present application, before encoding all the knowledge graphs in the candidate database according to the order of knowledge graph storage in the knowledge graph and determining the encoding string corresponding to each knowledge graph in the candidate database, the method further includes:

[0039] Convert the length of the secret information bitstream into binary to obtain a total length encoding string;

[0040] Concatenate the total length encoding string and the secret information bitstream to form a new secret information bitstream.

[0041] In a second aspect, an embodiment of the present invention provides a steganographic text extraction method incorporating a large language model, which may include:

[0042] Obtain a steganographic text, a candidate database, and a supplementary database; the knowledge graphs in the candidate database are arranged in the order of storage;

[0043] Perform entity extraction on the steganographic text to obtain a plurality of entities to be matched;

[0044] Match the plurality of entities to be matched with the candidate database to obtain the knowledge graph corresponding to the steganographic text;

[0045] Encode all the knowledge graphs in the candidate database according to the order of knowledge graph storage in the knowledge graph and determine the encoding string corresponding to each knowledge graph in the candidate database;

[0046] Determine the encoding string corresponding to the knowledge graph corresponding to the steganographic text according to the encoding string corresponding to each knowledge graph in the candidate database, and use it as the bitstream corresponding to the steganographic text;

[0047] Delete the knowledge graph corresponding to the steganographic text from the candidate database, and select a knowledge graph from the supplementary database and add it to the candidate database to obtain an updated candidate database;

[0048] Obtain the next steganographic text again, and execute the above process of performing entity extraction on the next steganographic text to obtain the bitstream corresponding to the next steganographic text, until no more steganographic texts are received, and concatenate the bitstreams corresponding to all the steganographic texts to obtain a secret information bitstream.

[0049] In a third aspect, an embodiment of the present invention provides a text steganography device incorporating a large language model, which may include:

[0050] A first segmentation module, configured to segment the obtained secret information bitstream into a plurality of segmented bitstreams arranged in sequence according to a preset segmentation length;

[0051] A database construction module, configured to respectively calculate a candidate data volume and a supplementary data volume based on the secret information bitstream and the preset segmentation length, and collect knowledge graphs corresponding to the respective quantities to obtain a candidate database and a supplementary database;

[0052] A first encoding module, configured to encode all the knowledge graphs in the candidate database according to the order of knowledge graph warehousing, and determine the encoding string corresponding to each knowledge graph in the candidate database;

[0053] A first selection module, configured to determine the knowledge graph corresponding to the first segmented bitstream as the selected knowledge graph according to the encoding string corresponding to each knowledge graph in the candidate database;

[0054] A text generation module, configured to input the selected knowledge graph into a preset knowledge graph text generation model constructed using a large language model to obtain a steganographic text corresponding to the first segmented bitstream;

[0055] A first update module, configured to delete the selected knowledge graph from the candidate database and select a knowledge graph from the supplementary database to add to the candidate database to obtain an updated candidate database;

[0056] A first judgment module, configured to judge whether all the segmented bitstreams have been processed: if not, the above first encoding module is re-executed.

[0057] In a fourth aspect, an embodiment of the present invention provides a steganographic text extraction device incorporating a large language model, which may include:

[0058] A first acquisition module, configured to acquire a steganographic text, a candidate database, and a supplementary database; the knowledge graphs in the candidate database are arranged in the warehousing order;

[0059] An entity extraction module, configured to perform entity extraction on the steganographic text to obtain a plurality of entities to be matched;

[0060] A first matching module, configured to match the plurality of entities to be matched with the candidate database to obtain the knowledge graph corresponding to the steganographic text;

[0061] A second encoding module, configured to encode all the knowledge graphs in the candidate database according to the order of knowledge graph storage in the knowledge graph, and determine an encoding string corresponding to each knowledge graph in the candidate database;

[0062] A first determination module, configured to determine an encoding string corresponding to the knowledge graph corresponding to the stego text according to the encoding string corresponding to each knowledge graph in the candidate database, as the bitstream corresponding to the stego text;

[0063] A second update module, configured to delete the knowledge graph corresponding to the stego text from the candidate database, and select a knowledge graph from the supplementary database to add to the candidate database to obtain an updated candidate database;

[0064] A second determination module, configured to determine whether the number of received stego texts is equal to the number of knowledge graphs in the supplementary database: if so, splice the bitstreams corresponding to all the stego texts to obtain a secret information bitstream; if not, obtain the next stego text and re-execute the above-mentioned entity extraction module.

[0065] In a fifth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the text steganography method introducing a large language model and / or the stego text extraction method introducing a large language model as described above are implemented.

[0066] In a sixth aspect, an embodiment of the present invention provides a computer program product, including computer program / instructions, and when the computer program / instructions are executed by a processor, the text steganography method introducing a large language model and / or the stego text extraction method introducing a large language model as described above are implemented.

[0067] In a seventh aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory, and when the processor executes the computer program, the text steganography method introducing a large language model and / or the stego text extraction method introducing a large language model as described above are implemented.

[0068] The beneficial effects of the above technical solutions provided by the embodiments of the present invention at least include:

[0069] An embodiment of the present invention provides a text steganography method introducing a large language model. The method divides a secret information bitstream into multiple segmented bitstreams according to a preset segmentation length. Then, it calculates the candidate database and the supplementary data volume, collects the corresponding number of knowledge graphs to form the candidate database and the supplementary database. Next, it encodes the knowledge graphs in the candidate database, selects the knowledge graph corresponding to the first segmented bitstream, generates the steganographic text, deletes the selected graph, and adds a new graph from the supplementary database to update the candidate database. Subsequently, it repeats the process of encoding the knowledge graphs in the candidate database to updating the candidate database until all segmented bitstreams are processed. By introducing a large language model and selecting based on knowledge graphs, this method significantly improves the quality and naturalness of the generated text, making the embedded secret information more concealed and less noticeable. Compared with traditional generative text steganography techniques, this method not only overcomes the problem of poor text quality caused by model limitations but also can effectively control the information embedding volume by constructing and utilizing the knowledge graph database without affecting the naturalness of the text, ensuring that each generated text corresponds to high-quality knowledge graph entity relationships, thereby achieving higher-capacity secret information hiding. At the same time, using a large language model to construct a preset knowledge graph text generation model ensures the high fluency and semantic accuracy of text generation, enhances the practicability and reliability of text steganography, and realizes the effective transmission of secret information while maintaining excellent concealment.

[0070] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained by the structures specifically pointed out in the written specification and the drawings.

[0071] The technical solutions of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings

[0072] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0073] Figure 1 It is a schematic flowchart of the text steganography method introducing a large language model provided by the embodiment of the present invention;

[0074] Figure 2 It is an example diagram of the steps of adaptability detection provided by the embodiment of the present invention;

[0075] Figure 3 It is a schematic framework diagram of establishing a candidate database and a supplementary database provided by the embodiment of the present invention;

[0076] Figure 4 This is an example diagram of perfect binary tree encoding provided by an embodiment of the present invention;

[0077] Figure 5 This is a schematic diagram of the conversion steps of the initial text corresponding to the knowledge graph provided by an embodiment of the present invention;

[0078] Figure 6 This is a schematic framework diagram of a text steganography and stego-text extraction method introducing a large language model provided by an embodiment of the present invention;

[0079] Figure 7 This is a schematic process diagram of stego-text extraction introducing a large language model provided by an embodiment of the present application;

[0080] Figure 8 This is a schematic structural diagram of a text steganography device introducing a large language model provided by an embodiment of the present application;

[0081] Figure 9 This is a schematic structural diagram of a stego-text extraction device introducing a large language model provided by an embodiment of the present application. Detailed implementation manners

[0082] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0083] The inventors found that in the prior art, although traditional generative text steganography techniques can embed secret information, due to the limitations of early models, the quality of the generated text is often not high, and as the amount of embedded information increases, the naturalness and fluency of the text will gradually decline, making it difficult to reach the level of human natural language. Based on this, the inventors made further research and developed the present invention to provide a text steganography method, related methods and devices introducing a large language model.

[0084] Embodiment 1

[0085] An embodiment of the present invention provides a text steganography method introducing a large language model. Referring to Figure 1 as shown, the method may include the following steps S101-S108:

[0086] S101: Divide the obtained secret information bit stream into a plurality of segmented bit streams arranged in sequence according to a preset segmentation length.

[0087] S102: Based on the secret information bitstream and the preset segmentation length, calculate the candidate data volume and the supplementary data volume respectively, and collect the corresponding number of knowledge graphs to obtain a candidate database and a supplementary database.

[0088] S103: Encode all the knowledge graphs in the candidate database according to the order of knowledge graph storage, and determine the encoding string corresponding to each knowledge graph in the candidate database.

[0089] S104: Based on the encoding string corresponding to each knowledge graph in the candidate database, determine the knowledge graph corresponding to the first segmented bitstream as the selected knowledge graph.

[0090] S105: Input the selected knowledge graph into a preset knowledge graph text generation model constructed using a large language model to obtain the steganographic text corresponding to the first segmented bitstream.

[0091] S106: Delete the selected knowledge graph from the candidate database, and select a knowledge graph from the supplementary database and add it to the candidate database to obtain an updated candidate database.

[0092] S107: Determine whether all segmented bitstreams have been processed;

[0093] If so, execute step S108;

[0094] If not, re-execute the above steps S103 - S106 to obtain the steganographic text corresponding to the next segmented bitstream.

[0095] S108: Complete text steganography.

[0096] An embodiment of the present invention provides a text steganography method introducing a large language model. This method divides the secret information bit stream into multiple segmented bit streams according to a preset segmentation length. Then, it calculates the candidate database and the supplementary data volume, collects the corresponding number of knowledge graphs to form the candidate database and the supplementary database. Next, it encodes the knowledge graphs in the candidate database, selects the knowledge graph corresponding to the first segmented bit stream, generates the steganographic text, deletes the selected graph and adds a new graph from the supplementary database to update the candidate database. Subsequently, it repeats the process of encoding the knowledge graphs in the candidate database to updating the candidate database until all segmented bit streams are processed. By introducing a large language model and selecting based on knowledge graphs, this method significantly improves the quality and naturalness of the generated text, making the embedded secret information more concealed and less noticeable. Compared with traditional generative text steganography techniques, this method not only overcomes the problem of poor text quality caused by model limitations, but also can effectively control the amount of information embedding by constructing and utilizing the knowledge graph database without affecting the naturalness of the text, ensuring that each generated text corresponds to a high-quality knowledge graph entity relationship, thereby achieving a higher capacity of secret information hiding. At the same time, using a large language model to construct a preset knowledge graph text generation model ensures the high fluency and semantic accuracy of text generation, enhances the practicality and reliability of text steganography, and realizes the effective transmission of secret information while maintaining excellent concealment.

[0097] In the above step S101, the obtained secret information bit stream is segmented into multiple sequentially arranged segmented bit streams according to the preset segmentation length.

[0098] Specifically, it can be that, according to the preset segmentation length, the secret information bit stream is segmented into multiple segmented bit streams starting from the beginning, and these segmented bit streams are arranged in the order of their positions in the secret information bit stream.

[0099] When the length of the secret information bit stream is a multiple of the preset segmentation length, the length of each segmented bit stream is equal to the preset segmentation length. When the length of the secret information bit stream is not a multiple of the preset segmentation length, then except for the last segmented bit stream, the lengths of other segmented bit streams are equal to the preset segmentation length, and the length of the last segmented bit stream is less than the preset segmentation length.

[0100] For example, the length of the secret information bit stream is 2^10 = 1024, and the preset segmentation length is 2^4 = 16. At this time, the length of the secret information bit stream is a multiple of the preset segmentation length, and 64 segmented bit streams can be segmented, and the length of each segmented bit stream is equal to 16.

[0101] For another example, the length of the secret information bitstream is 1020, and the preset segmentation length is 2^4 = 16. At this time, the length of the secret information bitstream is not an integer multiple of the preset segmentation length. 64 segmented bitstreams can be obtained. The lengths of the first 63 segmented bitstreams are equal to 16, and the length of the last segmented bitstream is equal to 12.

[0102] In the embodiments of the present application, when the length of the secret information bitstream is not an integer multiple of the preset segmentation length, that is, when the length of the last segmented bitstream is not equal to the preset segmentation length, this situation is referred to as an overflow situation in the art. At this time, in order to ensure normal subsequent encoding, the method of padding with 0s is required to supplement the length of the last segmented bitstream to the preset segmentation length.

[0103] In addition, in order to let the receiving party responsible for steganographic text extraction know whether an overflow situation occurs, if an overflow situation occurs, the number of 0s padded needs to be informed to the receiving party. Therefore, the length of the secret information bitstream can be converted into binary and concatenated with the secret information bitstream to obtain a secret information bitstream containing length information, and then the subsequent steganographic process is performed. Ensure that the receiving party can correctly decode and restore the original secret information bitstream, thereby enhancing the reliability and security of the text steganographic process.

[0104] In the above step S102, based on the secret information bitstream and the preset segmentation length, the candidate data volume and the supplementary data volume are calculated respectively, and the corresponding number of knowledge graphs are collected to obtain the candidate database and the supplementary database. Specifically, it includes the following steps S1021 - S1023:

[0105] S1021: Calculate the power of 2 based on the preset segmentation length to obtain the candidate data volume.

[0106] Specifically, it can be that, according to the preset segmentation length, the corresponding power of 2 is calculated to obtain the candidate data volume. For example, if the preset segmentation length is 16, then the candidate data volume is 2^16 = 65536.

[0107] S1022: Calculate the quotient of the length of the secret information bitstream and the preset segmentation length to obtain the supplementary data volume.

[0108] For example, if the length of the secret information bitstream is 2^10 = 1024 and the preset segmentation length is 2^4 = 16, then the supplementary data volume is 1024 / 16 = 64.

[0109] S1023: According to the candidate data volume and the supplementary data volume, collect the corresponding number of knowledge graphs respectively to obtain the candidate database and the supplementary database. Specifically, it includes the following steps S10231 - S10239:

[0110] S10231: Obtain a knowledge graph.

[0111] Specifically, it can be that various methods can be adopted to obtain a knowledge graph, including but not limited to: randomly generating a knowledge graph using a large language model (such as ChatGPT); using an open-source knowledge graph database (such as Neo4j); using entity relationship extraction tools (such as Spacy, OpenNLP) to convert any natural text into a knowledge graph.

[0112] S10232: Check whether the knowledge graph already exists in the candidate database and the supplementary database: If so, execute step S10233; if not, execute step S10234.

[0113] In the embodiments of the present application, the proposed encoding idea for knowledge graph selection is mainly based on the fact that if an encoding operation needs to be performed on a certain type of element (such as a knowledge graph), the types of each element should be unique. Therefore, before adding the knowledge graph to the candidate database or the supplementary database, it is necessary to first perform the uniqueness detection in steps S10232 - S10233.

[0114] S10233: Discard the knowledge graph and return to step S10231.

[0115] In the embodiments of the present application, after completing the uniqueness detection, it is necessary to continue to detect the adaptability between the knowledge graph and the preset knowledge graph text generation model to ensure that the generated text is consistent with the knowledge graph. Specifically, it includes the following steps S10234 - S10238.

[0116] S10234: Input the knowledge graph into the preset knowledge graph text generation model to obtain the generated text to be detected.

[0117] Among them, the preset knowledge graph text generation model can be obtained through the following method:

[0118] Obtain an open-source large language model as the initial knowledge graph text generation model, and use the WebNLG (Web Natural Language Generation) corpus to fine-tune and train the initial knowledge graph text generation model based on the P-Tuning v2 method to obtain the preset knowledge graph text generation model.

[0119] Among them, the open-source large language model can be ChatGLM2-6B. ChatGLM-6B is an open-source, bilingual (Chinese-English) dialogue language model released by the Knowledge Engineering Group (KEG) & Data Mining at Tsinghua University. Based on the GLM (General Language Model) architecture, it has 6.2 billion parameters. ChatGLM2-6B is the second-generation version of ChatGLM-6B. On the basis of retaining many excellent features of the original model such as smooth dialogue and low deployment threshold, ChatGLM2-6B has more powerful performance.

[0120] The WebNLG corpus consists of multiple triples and aims to express entities and their relationships through natural language text. Each sample in the WebNLG corpus contains at most 7 triples, and one or more corresponding reference texts are provided for each sample.

[0121] In the embodiments of this application, the above method for obtaining the preset knowledge graph text generation model successfully transforms the initial knowledge graph text generation model into an efficient knowledge graph text generator, that is, the preset knowledge graph text generation model by fine-tuning using the WebNLG corpus. The triple set and the corresponding reference texts in the WebNLG corpus enable the model to learn how to transform structured entity relationship information into fluent natural language text. At the same time, by applying the P-Tuning v2 method, this method efficiently fine-tunes the initial knowledge graph text generation model under the few-shot condition, greatly reducing the requirements for a large number of parameters and computing resources, and being able to maintain good generation quality. While optimizing storage and memory usage, it also improves the fine-tuning effect of medium-scale models in complex tasks, thus achieving efficient and resource-saving knowledge graph text generation.

[0122] S10235: Extract entities from the text to be detected for generation to obtain multiple entities.

[0123] Specifically, it can be to use an entity relationship extraction tool (such as Spacy, OpenNLP) to extract all entities in the text to be detected for generation.

[0124] S10236: Detect whether the multiple entities match the knowledge graph: If so, execute step S10237; if not, execute step S10238.

[0125] S10237: Add the knowledge graph to the candidate database or the supplementary database.

[0126] S10238: Discard the knowledge graph and return to step S10231.

[0127] S10239: Retrieve a new knowledge graph and perform the uniqueness detection of the knowledge graph in steps S10232 - S10233 and the adaptability detection in steps S10234 - S10238 until the number of knowledge graphs in the candidate database and the supplementary database meets the candidate data volume and the supplementary data volume.

[0128] To facilitate the understanding of this solution by those skilled in the art, the following provides a clearer and more complete description of the specific implementation process of the adaptability detection in steps S10234 - S10238 provided in the embodiments of this application: Refer to Figure 2 As shown, the figure includes 3 steps: Step1 text generation, Step2 entity extraction, and Step3 entity comparison, which respectively correspond to steps S10234 - S10236 above. Among them, the knowledge graph in the red dashed box in Step1 is an example knowledge graph. For the convenience of explanation, this knowledge graph only contains one triple, including 2 entities "spider man", "stan lee", and the relationship "creator". Convert this knowledge graph into simple text, and the text text A is "spider man|creator|stan lee", and then input it into the LLM, that is, the preset knowledge graph text generation model described above, and get the input text B as "spider man's creator is stan lee.", and then enter Step2 to extract entities "spider man" and "stan lee" from the text text B. Finally, in Step3, determine whether the entities obtained in Step2 match the entities in the text text A. If so, store the knowledge graph in the candidate database or the supplementary database; otherwise, discard the knowledge graph.

[0129] In the embodiments of this application, for the establishment methods of the candidate database and the supplementary database described in step S102 above, refer to Figure 3 As shown, Graphs in the figure are knowledge graphs not yet stored in the database, and test is the uniqueness detection of the knowledge graph in steps S10232 - S10233 and the adaptability detection in steps S10234 - S10238. If the detection is passed, it can be added to the CP Database or the Supplement Database, where CP Database is the candidate database described above, and Supplement Database is the supplementary database described above. Note that at this time, there is no need to pay attention to Figure 3An arrow pointing from the supplementary database to the candidate database is added. In step S102, there is no step of supplementing the data in the supplementary database to the candidate database. This arrow indicates the subsequent step S106.

[0130] The above step S102 ensures the uniqueness and encodability of each knowledge graph in the database through systematic uniqueness detection and adaptability detection, thereby improving the quality and accuracy of the knowledge graph, ensuring that the knowledge graphs in the database can accurately support subsequent coding and text generation tasks, and providing strong guarantee for the practical application of the knowledge graph.

[0131] In the above step S103, all the knowledge graphs in the candidate database are encoded according to the order of their entry into the knowledge graph database, and the encoding string corresponding to each knowledge graph in the candidate database is determined.

[0132] Among them, the coding method can choose fixed-length coding (Fixed-Length Coding, FLC) and variable-length coding (Variable-Length Coding, VLC). FLC is a coding method based on a perfect binary tree, and VLC is a coding method based on a Huffman tree. Both are existing technologies and will not be elaborated here.

[0133] Taking the coding method as fixed-length coding as an example, the above step S103 specifically can be to construct a perfect binary tree for all the knowledge graphs in the candidate database according to the entry order, and code the entire perfect binary tree in the form of left subtree 0 and right subtree 1, so that the leaf nodes of the perfect binary tree correspond to a knowledge graph in turn from left to right. Then the coding of the leaf node corresponding to each knowledge graph is the coding string corresponding to the knowledge graph.

[0134] For example, the coding method is fixed-length coding, and the preset segment length is 3. Then according to the above step S102, it can be determined that there are 8 knowledge graphs in the candidate database, numbered ①-⑧ in the entry order. Then a perfect binary tree is constructed based on the candidate database and coded in the form of left subtree 0 and right subtree 1 for the entire perfect binary tree. The example diagram of the obtained perfect binary tree is as Figure 4 shown. There are 8 leaf nodes in the perfect binary tree. By corresponding these 8 leaf nodes to the 8 knowledge graphs arranged in order from left to right, it can be obtained that the coding string of knowledge graph ① is 000, the coding string of knowledge graph ② is 001, the coding string of knowledge graph ③ is 010, the coding string of knowledge graph ④ is 011, the coding string of knowledge graph ⑤ is 100, the coding string of knowledge graph ⑥ is 101, the coding string of knowledge graph ⑦ is 110, and the coding string of knowledge graph ⑧ is 111.

[0135] In the above step S104, according to the encoding strings corresponding to each knowledge graph in the candidate database, determine the knowledge graph corresponding to the first segmented bitstream as the selected knowledge graph.

[0136] Specifically, it can be to find the knowledge graph whose encoding string is equal to the first segmented bitstream according to the encoding strings corresponding to each knowledge graph in the candidate database as the selected knowledge graph.

[0137] In the above step S105, input the selected knowledge graph into a preset knowledge graph text generation model to obtain the steganographic text corresponding to the first segmented bitstream. Specifically, it includes the following steps S1051 - S1054:

[0138] S1051: Decompose the selected knowledge graph into multiple triples.

[0139] S1052: Concatenate the elements in each triple in the order of subject, predicate, and object, and separate adjacent elements with a first separator symbol to obtain the triple text corresponding to the triple.

[0140] S1053: Concatenate the triple texts corresponding to all triples, and separate adjacent triple texts with a second separator symbol to obtain the initial text corresponding to the selected knowledge graph.

[0141] To facilitate those skilled in the art to understand this solution, the following provides a clearer and more complete description of the specific implementation processes of S1051 - S1053 provided by the embodiments of this application: Refer to Figure 5 As shown in the figure, the process of converting the selected knowledge graph into the initial text in the figure is represented by 2 steps. Step1 decomposes the knowledge graph to obtain individual triples, and step2 represents the knowledge graph text. Among them, step1 corresponds to step S1051, and step2 corresponds to steps S1052 and S1053.

[0142] In Figure 5 In step1 of, Graph B in the red box in the upper left corner is the selected knowledge graph. The selected knowledge graph is decomposed into two triples, that is, the two triples in the green boxes on the left side of the figure. Among them, the left triple contains 2 entities "spiderman", "Peter park", and the relationship "real name", and the right triple contains 2 entities "spider man", "stan lee", and the relationship "creator".

[0143] In Figure 5In step 2, simple text conversions are performed on the two triples obtained in step 1 respectively. Using the symbol "|" as the first delimiter, the triple texts corresponding to the two triples are obtained as "spider man|real name|Peter park" and "spider man|creator|stan lee". Then, the two triple texts are concatenated, and adjacent triple texts are separated by the second delimiter "*", obtaining the initial text corresponding to the selected knowledge graph as "spider man|real name|Peter park*spider man|creator|stan lee".

[0144] S1054: Input the initial text into a preset knowledge graph text generation model to obtain the steganographic text corresponding to the first segmented bitstream.

[0145] Among them, the obtaining method of the preset knowledge graph text generation model has been described in the above step S10234 and will not be elaborated here.

[0146] In the embodiment of the present application, the above step S105 uses a preset knowledge graph text generation model to convert the selected knowledge graph into the corresponding steganographic text, ensuring that the secret information bitstream can be steganographically hidden through the way of natural language expression. This process realizes the conversion from a structured knowledge graph to a natural language text, ensuring that the information hiding process not only meets the steganography requirements but also maintains the traceability and semantic consistency of the information, making the steganographic information more difficult to detect by external observers and improving the security and concealment of the steganography process.

[0147] In the above step S106, the selected knowledge graph is deleted from the candidate database, and a knowledge graph is selected from the supplementary database and added to the candidate database to obtain an updated candidate database.

[0148] Specifically, it can be that the selected knowledge graph is deleted from the candidate database, and a knowledge graph is selected from the supplementary database in a preset order and added to the candidate database to obtain an updated candidate database. Among them, the preset order can be the storage order of the knowledge graphs in the supplementary database.

[0149] For example, if the steganographic text corresponding to the first segmented bitstream is obtained in the current loop, then at this time, the knowledge graph with a preset order of 1 should be selected from the supplementary database and supplemented to the candidate database to obtain an updated candidate database.

[0150] If the steganographic text corresponding to the i-th segmented bitstream is obtained in the current loop, then at this time, the knowledge graph with a preset order of i should be selected from the supplementary database and supplemented to the candidate database to obtain an updated candidate database.

[0151] In the embodiments of the present application, the method defines the size of the candidate database by customizing the preset segmentation length, so as to determine the number of secret information embedded each time. Therefore, even when there is a large amount of secret information to be embedded, only a sufficiently large candidate database needs to be set to ensure that the number of knowledge graphs in the candidate database is sufficient to meet the embedding of a piece of secret information.

[0152] After completing the text steganography method for knowledge graph selection described in S101 - S107 above, correspondingly, a secret information bitstream can be extracted based on the steganographic text. The extraction method specifically includes the following steps S201 - S209:

[0153] For ease of explanation, the party that executes the text steganography method introducing the large language model will be referred to as the sender, and the party that executes the text steganography extraction method introducing the large language model will be referred to as the receiver.

[0154] S201: Obtain the steganographic text, the candidate database, and the supplementary database. The knowledge graphs in the candidate database are arranged in the order of storage.

[0155] S202: Perform entity extraction on the steganographic text to obtain multiple entities to be matched.

[0156] S203: Match according to the multiple entities to be matched with the candidate database to obtain the knowledge graph corresponding to the steganographic text.

[0157] S204: Encode all the knowledge graphs in the candidate database according to the storage order of the knowledge graphs to determine the encoding string corresponding to each knowledge graph in the candidate database.

[0158] S205: Determine the encoding string corresponding to the knowledge graph corresponding to the steganographic text according to the encoding string corresponding to each knowledge graph in the candidate database, as the bitstream corresponding to the steganographic text.

[0159] S206: Delete the knowledge graph corresponding to the steganographic text from the candidate database, and select a knowledge graph from the supplementary database and add it to the candidate database to obtain an updated candidate database.

[0160] Specifically, it can be to delete the knowledge graph corresponding to the steganographic text in the candidate database, and select a knowledge graph from the supplementary database and supplement it to the candidate database according to the preset order to obtain an updated candidate database. Note that the preset order here needs to be consistent with the preset order in the steganographic text method based on knowledge graph selection to ensure that the receiver can update the candidate database in the same way and maintain consistency with the receiver.

[0161] S207: Determine whether the number of received steganographic texts is equal to the number of knowledge graphs in the supplementary database: if so, execute S209; if not, execute S208;

[0162] S208: Obtain the next steganographic text and execute the above steps S202 - S207.

[0163] S209: Concatenate the bitstreams corresponding to all steganographic texts to obtain the secret information bitstream.

[0164] To facilitate the understanding of this solution by those skilled in the art, the following provides a clearer and more complete description of the specific implementation process of the text steganography method introducing a large language model provided in the embodiments of this application:

[0165] Refer to Figure 6 As shown, on the left is the framework of the text steganography method introducing a large language model, and on the right is the framework of the steganographic text extraction method introducing a large language model. In the figure, the secret information bitstream is arranged vertically and side by side on the far left. The boxes in blue, gray, and purple all represent a segmented bitstream. The Sum in the red box represents the length of the secret information bitstream. After being converted to binary, corresponding knowledge graphs G0 - Gn are obtained from the candidate database in sequence with other segmented bitstreams, and corresponding steganographic texts Text 0 - Text n are obtained through a preset knowledge graph text generation model.

[0166] Figure 6 Text 0 - Text n at the far left in the right box are multiple steganographic texts. Through entity extraction and matching, corresponding knowledge graphs G0 - Gn are obtained. Among them, the encoded string corresponding to G0 can be converted to decimal, representing the length of the secret information bitstream. At the same time, the encoded strings corresponding to G1 - Gn are concatenated in order to obtain the secret information bitstream.

[0167] Embodiment 2

[0168] Based on the same inventive concept, the embodiments of the present invention also provide a steganographic text extraction method introducing a large language model. Refer to Figure 7 As shown, this method includes:

[0169] For the convenience of description, the party executing the text steganography method introducing a large language model will be referred to as the sender later, and the party executing the steganographic text extraction method introducing a large language model will be referred to as the receiver.

[0170] S201: Obtain the steganographic text, the candidate database, and the supplementary database. The knowledge graphs in the candidate database are arranged in the storage order.

[0171] S202: Perform entity extraction on the steganographic text to obtain multiple entities to be matched.

[0172] S203: Match according to the multiple entities to be matched with the candidate database to obtain the knowledge graph corresponding to the steganographic text.

[0173] S204: Encode all the knowledge graphs in the candidate database according to the order of knowledge graph storage in the knowledge graph database, and determine the encoding string corresponding to each knowledge graph in the candidate database.

[0174] S205: According to the encoding string corresponding to each knowledge graph in the candidate database, determine the encoding string corresponding to the knowledge graph corresponding to the stego text as the bitstream corresponding to the stego text.

[0175] S206: Delete the knowledge graph corresponding to the stego text from the candidate database, and select a knowledge graph from the supplementary database to add to the candidate database to obtain an updated candidate database.

[0176] S207: Determine whether the number of received stego texts is equal to the number of knowledge graphs in the supplementary database: if so, execute S209; if not, execute S208;

[0177] S208: Obtain the next stego text and execute the above steps S202 - S207.

[0178] S209: Concatenate the bitstreams corresponding to all stego texts to obtain the secret information bitstream.

[0179] Embodiment 3

[0180] Based on the same inventive concept, an embodiment of the present invention further provides a text steganography device incorporating a large language model. Referring to Figure 8 as shown, the device includes:

[0181] The first segmentation module 101 is configured to segment the obtained secret information bitstream into a plurality of sequentially arranged segmented bitstreams according to a preset segmentation length;

[0182] The database construction module 102 is configured to calculate a candidate data volume and a supplementary data volume respectively based on the secret information bitstream and the preset segmentation length, and collect the corresponding number of knowledge graphs to obtain a candidate database and a supplementary database;

[0183] The first encoding module 103 is configured to encode all the knowledge graphs in the candidate database according to the order of knowledge graph storage in the knowledge graph database, and determine the encoding string corresponding to each knowledge graph in the candidate database;

[0184] The first selection module 104 is configured to determine the knowledge graph corresponding to the first segmented bitstream according to the encoding string corresponding to each knowledge graph in the candidate database as the selected knowledge graph;

[0185] The text generation module 105 is configured to input the selected knowledge graph into a preset knowledge graph text generation model constructed using a large language model to obtain the stego text corresponding to the first segmented bitstream;

[0186] The first update module 106 is configured to delete the selected knowledge graph from the candidate database and select a knowledge graph from the supplementary database to add to the candidate database, so as to obtain an updated candidate database;

[0187] The first judgment module 107 is configured to judge whether all the segmented bitstreams have been processed: if not, the above-mentioned first encoding module is re-executed.

[0188] Embodiment 4

[0189] Based on the same inventive concept, an embodiment of the present invention further provides a steganographic text extraction device introducing a large language model. Referring to Figure 9 as shown, the device includes:

[0190] The first acquisition module 201 is configured to acquire steganographic text, a candidate database, and a supplementary database; the knowledge graphs in the candidate database are arranged in the storage order;

[0191] The entity extraction module 202 is configured to perform entity extraction on the steganographic text to obtain a plurality of entities to be matched;

[0192] The first matching module 203 is configured to match according to the plurality of entities to be matched with the candidate database to obtain the knowledge graph corresponding to the steganographic text;

[0193] The second encoding module 204 is configured to encode all the knowledge graphs in the candidate database according to the storage order of the knowledge graphs to determine the encoding string corresponding to each knowledge graph in the candidate database;

[0194] The first determination module 205 is configured to determine the encoding string corresponding to the knowledge graph corresponding to the steganographic text according to the encoding string corresponding to each knowledge graph in the candidate database as the bitstream corresponding to the steganographic text;

[0195] The second update module 206 is configured to delete the knowledge graph corresponding to the steganographic text from the candidate database and select a knowledge graph from the supplementary database to add to the candidate database to obtain an updated candidate database;

[0196] The second judgment module 207 is configured to judge whether the number of received steganographic texts is equal to the number of knowledge graphs in the supplementary database: if so, splice the bitstreams corresponding to all the steganographic texts to obtain a secret information bitstream; if not, obtain the next steganographic text and re-execute the above-mentioned entity extraction module.

[0197] Embodiment 5

[0198] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the text steganography method introducing a large language model described in the first embodiment above, and / or the steganographic text extraction method introducing a large language model described in the second embodiment above are implemented.

[0199] Embodiment Six

[0200] Based on the same inventive concept, an embodiment of the present invention further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the text steganography method introducing a large language model described in the first embodiment above, and / or the steganographic text extraction method introducing a large language model described in the second embodiment above are implemented.

[0201] Embodiment Seven

[0202] Based on the same inventive concept, an embodiment of the present invention further provides a computer device, including a memory, a processor, and a computer program stored on the memory, and when the processor executes the computer program, the text steganography method introducing a large language model described in the first embodiment above, and / or the steganographic text extraction method introducing a large language model described in the second embodiment above are implemented.

[0203] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0204] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0205] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means embodying the functionality specified in the flowchart(s) Figure 1 of one or more flowcharts and / or block diagram(s) Figure 1 of one or more block diagrams.

[0206] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functionality specified in the flowchart(s) Figure 1 of one or more flowcharts and / or block diagram(s) Figure 1 of one or more block diagrams.

[0207] It will be apparent to those skilled in the art that various modifications and variations can be made to the present invention without departing from the spirit and scope of the invention. Thus, if these modifications and variations of the present invention come within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A text steganography method using a large language model, characterized in that: include: The obtained secret information bit stream is divided into a plurality of segmented bit streams arranged in sequence according to a preset segment length; Based on the secret information bit stream and the preset segment length, respectively calculating the amount of candidate data and the amount of supplementary data, and collecting a corresponding number of knowledge graphs to obtain a candidate database and a supplementary database; Encode all knowledge graphs in the candidate database according to the order in which the knowledge graphs are stored, and determine the encoding string corresponding to each knowledge graph in the candidate database; According to the encoding string corresponding to each knowledge graph in the candidate database, determine the knowledge graph corresponding to the first segmented bit stream as the selected knowledge graph; Inputting the selected knowledge graph into a preset knowledge graph text generation model constructed using a large language model to obtain a steganographic text corresponding to the first segmented bit stream; Deleting the selected knowledge graph from the candidate database, and selecting a knowledge graph from the supplementary database to add to the candidate database, to obtain an updated candidate database; The above process of encoding all knowledge graphs in the update candidate database to obtain the steganographic text corresponding to the next segmented bit stream is re-executed until all the segmented bit streams are processed.

2. The method according to claim 1, characterized in that The method of calculating the candidate data amount and the supplementary data amount based on the secret information bit stream and the preset segment length, and collecting a corresponding number of knowledge graphs to obtain a candidate database and a supplementary database includes: Based on the preset segment length, calculating the power of 2 to obtain the candidate data amount; Calculating the quotient of the length of the secret information bit stream and the preset segment length to obtain the amount of supplementary data; According to the amount of candidate data and the amount of supplementary data, corresponding quantities of knowledge graphs are collected to obtain a candidate database and a supplementary database.

3. The method according to claim 2, characterized in that According to the amount of candidate data and the amount of supplementary data, a corresponding number of knowledge graphs are collected respectively to obtain a candidate database and a supplementary database, including: Get a knowledge graph; Check whether the knowledge graph already exists in the candidate database and the supplementary database; If yes, the knowledge graph is discarded; If not, the knowledge graph is input into the preset knowledge graph text generation model to obtain the generated text to be tested; Performing entity extraction on the generated text to be detected to obtain multiple entities; Detect whether the multiple entities match the knowledge graph: If yes, adding the knowledge graph to the candidate database or supplementary database; If not, the knowledge graph is discarded; Reacquire a new knowledge graph and perform the above-mentioned knowledge graph uniqueness detection and adaptability detection process until the number of knowledge graphs in the candidate database and the supplementary database meets the candidate data amount and the supplementary data amount.

4. The method according to claim 1, characterized in that Inputting the selected knowledge graph into a preset knowledge graph text generation model constructed using a large language model to obtain the steganographic text corresponding to the first segmented bit stream, including: Decomposing the selected knowledge graph into multiple triples; The elements in each triple are concatenated in the order of subject, predicate and object, and adjacent elements are separated by the first spacing symbol to obtain the triple text corresponding to the triple; The triple texts corresponding to all the triples are concatenated, and adjacent triple texts are separated by a second spacing symbol, to obtain the initial text corresponding to the selected knowledge graph; The initial text is input into the preset knowledge graph text generation model to obtain the steganographic text corresponding to the first segmented bit stream.

5. The method according to claim 1, characterized in that: The preset knowledge graph text generation model is obtained in the following way: Obtain the initial knowledge graph text generation model; Based on the P-Tuning v2 method, the initial knowledge graph text generation model is fine-tuned and trained to obtain a preset knowledge graph text generation model.

6. A method for extracting steganographic text by introducing a large language model, characterized in that: include: Obtaining stegographic text, candidate database and supplementary database; The knowledge graphs in the candidate database are arranged in the order of entry; Performing entity extraction on the stegoscopy text to obtain a plurality of entities to be matched; According to the multiple entities to be matched, matching is performed with the candidate database to obtain a knowledge graph corresponding to the steganographic text; Encode all knowledge graphs in the candidate database according to the order in which the knowledge graphs are stored, and determine the encoding string corresponding to each knowledge graph in the candidate database; According to the coding string corresponding to each knowledge graph in the candidate database, determine the coding string corresponding to the knowledge graph corresponding to the steganographic text as the bit stream corresponding to the steganographic text; Deleting the knowledge graph corresponding to the stegotext from the candidate database, and selecting a knowledge graph from the supplementary database to add to the candidate database, to obtain an updated candidate database; Re-acquire the next stegotext, perform the above-mentioned entity extraction on the next stegotext, and obtain the bit stream corresponding to the next stegotext, until no more stegotext is received, and splice the bit streams corresponding to all stegotexts to obtain the secret information bit stream.

7. A text steganography device introducing a large language model, characterized in that: include: A first segmentation module, used to segment the acquired secret information bit stream into a plurality of segmented bit streams arranged in sequence according to a preset segment length; A database construction module, used to calculate the amount of candidate data and the amount of supplementary data based on the secret information bit stream and the preset segment length, and collect a corresponding number of knowledge graphs to obtain a candidate database and a supplementary database; A first encoding module is used to encode all knowledge graphs in the candidate database according to the order in which the knowledge graphs are stored, and determine the encoding string corresponding to each knowledge graph in the candidate database; A first selection module is used to determine the knowledge graph corresponding to the first segmented bit stream as the selected knowledge graph according to the encoding string corresponding to each knowledge graph in the candidate database; A text generation module, used for inputting the selected knowledge graph into a preset knowledge graph text generation model constructed using a large language model to obtain a steganographic text corresponding to the first segmented bit stream; A first updating module is used to delete the selected knowledge graph from the candidate database, and select a knowledge graph from the supplementary database to add to the candidate database, so as to obtain an updated candidate database; The first judging module is used to judge whether all the segmented bit streams have been processed: if not, the first encoding module is re-executed.

8. A device for extracting steganographic text by introducing a large language model, characterized in that: include: The first acquisition module is used to acquire the stegotext, the candidate database and the supplementary database; the knowledge graphs in the candidate database are arranged in the order of entry; An entity extraction module, used to extract entities from the steganographic text to obtain a plurality of entities to be matched; A first matching module, configured to match the candidate database according to the multiple entities to be matched, to obtain a knowledge graph corresponding to the steganographic text; The second encoding module is used to encode all the knowledge graphs in the candidate database according to the order in which the knowledge graphs are stored, and determine the encoding string corresponding to each knowledge graph in the candidate database; A first determination module is used to determine the encoding string corresponding to the knowledge graph corresponding to the steganographic text according to the encoding string corresponding to each knowledge graph in the candidate database, as the bit stream corresponding to the steganographic text; A second updating module is used to delete the knowledge graph corresponding to the stegotext from the candidate database, and select a knowledge graph from the supplementary database to add to the candidate database, so as to obtain an updated candidate database; The second judgment module is used to determine whether the number of received steganographic texts is equal to the number of knowledge graphs in the supplementary database: if so, the bit streams corresponding to all steganographic texts are spliced ​​to obtain the secret information bit stream; if not, the next steganographic text is obtained and the above-mentioned entity extraction module is re-executed.

9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the text steganography method introducing a large language model as described in any one of claims 1 to 5 and / or the steganographic text extraction method introducing a large language model as described in claim 6 are implemented.

10. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the text steganography method introducing a large language model as described in any one of claims 1 to 5, and / or the steganalysis text extraction method introducing a large language model as described in claim 6.

Citation Information

Patent Citations

  • Knowledge graph-based text generation steganography method, related method and device

    CN117131203A

  • Method for realizing main-standby communication of single-link equipment

    CN118118325A

  • Firewall attacked surface carding and security reinforcement method

    CN119276632A

Cited By

  • Unmanned aerial vehicle safety communication text steganography and transmission method based on reinforcement learning

    CN122137572A